模型与实验室 4.0 · 优秀 2026-05-29 · 文章

9 demos of Gemini Omni and Gemini 3.5 in action

Google 发布 9 个 Gemini Omni 与 Gemini 3.5 的视频演示,展示新一代模型在多模态推理长上下文理解Agent 编排上的实际能力

打开原文回到归档

9 demos of Gemini Omni and Gemini 3.5 in action

English (Original)

9 demos of Gemini Omni and Gemini 3.5 in action

11 min read

By Zahra Thompson, Contributor, The Keyword

At Google I/O 2026, we announced our latest models: Gemini Omni and the Gemini 3.5 family of models.

Gemini Omni is our new model that can create anything from any input, starting with video. With Omni, you can combine images, audio, video and text as input and generate high-quality videos grounded in Gemini's real-world knowledge. You can also easily edit your videos through conversation.

Then there's Gemini 3.5, our latest family of models combining frontier intelligence with action. This represents a major leap forward in building more capable, intelligent agents. We're kicking off the series by releasing 3.5 Flash. It delivers frontier performance for agents and coding, excelling at complex long-horizon tasks that deliver real-world utility.

To give you a clearer understanding of Gemini Omni and Gemini 3.5 Flash, here are 9 demos of what they can help you do.

Gemini Omni

Edit your videos through conversation. One capability that makes Omni special is that it gives you an easier way to edit video — with natural language. Every instruction builds on the last. Your characters stay consistent, the physics hold up and the scene remembers what came before. That means you can transform the world around you. Change specific things, or change everything. Your video becomes the starting point for something you never could have filmed yourself.

Prompt: Make the sculpture out of bubbles.

Reimagine the action. Take a video you shot and just ask Omni to change what's happening. Edit the action, add in new characters or objects or transform a moment into something unexpected.

Prompt: Dim the lights in the room. Put a black and white checkerboard room inside a glass sphere that floats tracking above the hand, inside it contains a recursive representation of the same hand holding the sphere, creating an infinite recursive of rooms. Camera slowly gets closer into the sphere, creating a video loop.

Refine your videos across multiple turns. Change the environment, angle, style or even specific details, without ever losing the thread of your original scene. Multi-turn prompts: 1) "A video of a violinist playing a song." 2) "Transport the violinist to the image environment." 3) "Make the violin invisible." 4) "Change the camera angle to be over the violinist's shoulder."

Gemini 3.5 Flash

Take on agentic tasks at scale. 3.5 Flash delivers intelligence that rivals large flagship models on multiple dimensions, at the speeds you have come to expect from the Flash series. This balance of speed and performance makes 3.5 Flash ideal for tackling long-horizon agentic tasks. Here, powered by Antigravity, 3.5 Flash executes multi-step workflows to automatically rename and categorize unstructured assets based on dynamic criteria.

When coupled with the updated Antigravity harness, 3.5 Flash becomes a powerful engine for deploying collaborative subagents to tackle problems at scale for the most demanding use cases. Under supervision, it can reliably execute multi-step workflows and coding tasks while sustaining frontier performance.

Create richer, more interactive web UIs and graphics with 3.5 Flash. 3.5 Flash builds on the strong multimodal foundation of Gemini 3. Watch as 3.5 Flash generates different UX approaches for a checkout flow in just 60 seconds on AI Studio.

Try personal AI agents and new intelligent experiences. 3.5 Flash is now the default model for the Gemini app and AI Mode in Search globally. Its agentic capabilities are powering new features to bring frontier-level intelligence to your daily life.

The enhanced agentic coding capabilities of 3.5 Flash are delivering even more intelligent experiences in Search, like our new information agents. Operating in the background, 24/7, these agents intelligently reason across information to find exactly what you need at exactly the right moment. They will send a comprehensive update along with links to the web to dive deeper, so you can take action. Information agents will launch first for Google AI Pro & Ultra subscribers this summer.

Now that we're bringing the power of Google Antigravity and agentic coding capabilities of Gemini 3.5 Flash right into Search, Search can build the ideal response, in the right format for your question — completely on the fly. So you can get custom generative UI, including visual tools and simulations, tailored precisely to your needs. These generative UI capabilities will be available for everyone in Search this summer, free of charge.

For your ongoing tasks like planning a wedding or establishing a new fitness routine, Search will also build you custom experiences – like dashboards, trackers or mini apps – that you can keep coming back to. You'll be able to create your own custom experiences with Antigravity right in Search in the coming months, starting first for Google AI Pro and Ultra subscribers in the U.S.

Then there's the new Gemini Spark, your personal AI agent, which runs on Gemini 3.5 and uses the Antigravity harness. It runs 24/7, helping you navigate your digital life, taking action on your behalf while under your direction. It's deeply integrated with the Workspace tools you rely on daily, like Gmail, Docs, Slides and more. Gemini Spark is now available to all Google AI Ultra subscribers in the U.S.

Availability

Gemini Omni Flash is rolling out to all Google AI Plus, Pro and Ultra subscribers globally through the Gemini app and Google Flow. It's also rolling out at no cost to users on YouTube Shorts and YouTube Create App. In the coming weeks, we'll also be rolling it out to developers and enterprise customers via APIs.

Gemini 3.5 Flash is generally available via Google Antigravity, the Gemini API in Google AI Studio and Android Studio, Gemini Enterprise Agent Platform, and Gemini Enterprise. It's also available for everyone in AI Mode in Search and now rolling out to everyone globally in the Gemini app.

中文 (Translation)

9 个 Gemini Omni 与 Gemini 3.5 实机演示

阅读时长 11 分钟

作者:Zahra Thompson,The Keyword 撰稿人

Google I/O 2026 上,我们正式发布了最新一代模型:Gemini OmniGemini 3.5 系列。

Gemini Omni 是我们新发布的"任意输入、任意输出"模型,先从视频开始落地。借助 Omni,你可以把图像、音频、视频与文本组合作为输入,让模型基于对真实世界的理解生成高质量视频,也可以通过对话对视频进行编辑。

紧接着是 Gemini 3.5——把"前沿智能"与"行动能力"合二为一的全新系列,是构建更强大智能体的重大飞跃。该系列首发的是 3.5 Flash:在智能体与编程场景下提供旗舰级性能,擅长复杂的"长链路"任务并带来真实业务价值。

为了让你更直观地理解 Gemini Omni 与 Gemini 3.5 Flash 能做什么,下面带来 9 个实机演示。

Gemini Omni

用对话的方式编辑视频。 Omni 的一项核心能力,是让你用自然语言就能改写视频。每一条指令都承接前一条——角色保持一致、物理规律保持一致、场景会"记住"之前发生过的事。你可以彻底改写周围的世界,改一个细节、改一个整体,把你原本拍不出的东西变成新视频的起点。

提示词:Make the sculpture out of bubbles.(把雕塑变成泡泡做的。)

重新想象镜头中的动作。 拿一段你拍的视频,让 Omni 改写它正在发生的事:修改动作、加入新角色或新物件,把某一刻彻底改写成你意想不到的样子。

提示词:Dim the lights in the room. Put a black and white checkerboard room inside a glass sphere that floats tracking above the hand, inside it contains a recursive representation of the same hand holding the sphere, creating an infinite recursive of rooms. Camera slowly gets closer into the sphere, creating a video loop.

多轮迭代精修视频。 切换环境、视角、风格乃至具体细节,全程不会丢掉原始场景的"主线"。多轮提示词示例:1)"A video of a violinist playing a song." 2)"Transport the violinist to the image environment." 3)"Make the violin invisible." 4)"Change the camera angle to be over the violinist's shoulder."

Gemini 3.5 Flash

把智能体任务规模化跑起来。 3.5 Flash 在多个维度上提供足以比肩大型旗舰模型的智能,速度却依然是 Flash 系列的看家本领。这种"速度 × 性能"的组合,使 3.5 Flash 成为处理"长链路"智能体任务的理想选择。在 Antigravity 的加持下,3.5 Flash 能够执行多步骤工作流,按照动态条件自动给无结构素材做重命名与归类。

搭配升级后的 Antigravity 框架之后,3.5 Flash 变成了一台能调度多个子智能体协作、解决大规模问题的强劲引擎。在监督机制的约束下,它可以稳定地执行多步工作流与编程任务,并保持旗舰级表现。

用 3.5 Flash 创造更丰富的 Web UI 与图形。 3.5 Flash 建立在 Gemini 3 强大的多模态底座之上。在 AI Studio 里,仅需 60 秒,它就能为一个 checkout 流程生成多种不同的 UX 方案。

体验个人 AI 智能体与全新的智能体验。 3.5 Flash 已成为 Gemini app 与 Search 中 AI Mode 全球范围内的默认模型。它的智能体能力正在驱动一系列新功能,把旗舰级智能带入你的日常生活。

3.5 Flash 增强的"智能体式编程"能力让 Search 体验更聪明——比如全新的信息智能体(information agents)。它们 7×24 小时在后台运行,跨多源信息做推理,精准在你最需要的时刻交付结果,并附上可点击深挖的链接。这个夏天,Google AI Pro 与 Ultra 订阅用户将率先用上信息智能体。

借助 Google Antigravity 与 3.5 Flash 的智能体式编程能力,Search 可以"当场造"出最合适的回答——按问题的形式,定制化生成 UI、视觉工具、模拟器等。这个夏天,生成式 UI 能力将对所有用户免费开放

对于"筹备婚礼""建立健身计划"这种长周期任务,Search 还能为你构建可以反复回来的定制化体验——仪表盘、追踪器、小应用等。几个月后,AI Pro 与 Ultra 订阅用户将率先在美国用上"在 Search 里直接用 Antigravity 搭建定制体验"的能力。

还有一个值得关注的:Gemini Spark——你的个人 AI 智能体。它跑在 Gemini 3.5 之上、调用 Antigravity 框架,7×24 小时待命,替你穿梭在数字生活里、按照你的方向采取行动。它与 Workspace 工具深度集成:Gmail、Docs、Slides 等。Gemini Spark 已向美国 Google AI Ultra 订阅用户开放。

上市与可用性

Gemini Omni Flash 正在向全球 Google AI Plus、Pro 与 Ultra 订阅用户推送,可在 Gemini app 和 Google Flow 中使用;同时面向 YouTube Shorts 与 YouTube Create App 的用户免费推出。未来几周也将通过 API 向开发者与企业客户开放。

Gemini 3.5 Flash 已通过 Google Antigravity、Google AI Studio 中的 Gemini API、Android Studio、Gemini Enterprise Agent Platform 与 Gemini Enterprise 全面可用;同时面向 AI Mode in Search 的全部用户开放,并正在全球范围内逐步推送给 Gemini app 的所有用户。