I/O 2026: Welcome to the agentic Gemini era
- 来源:Google Blog
- 原文链接:https://blog.google/innovation-and-ai/sundar-pichai-io-2026
- 作者:Sundar Pichai (Google & Alphabet CEO)
- 日期:2026-05-19
- 分类:models
- 标签:google, gemini, agentic, io-2026, sundar-pichai
- 抓取时间:2026-06-15
- 抓取方式:opencli web read
English (Original)
I/O 2026: Welcome to the agentic Gemini era
19 min read
By Sundar Pichai, CEO of Google and Alphabet
_Editor's note: Below is an edited transcript of Google CEO Sundar Pichai's remarks at Google I/O 2026, adapted to include more of what was announced on stage._
It's been an extraordinary year since our last I/O, a period of relentless shipping, technology advances and hyper progress. We're now in the part of the AI cycle where people want to see the value in the products they use every day. We've been really focused on that, and you'll see that in the products and features we're announcing today at I/O.
Ten years since we pivoted the company to be AI-first, we still see AI as the most profound way to advance our mission and improve people's lives at scale. That's why we've been taking a differentiated, full-stack approach to AI innovation, from our custom silicon and secure foundation, to our world-class research and models, to our products and platforms that touch billions of people. This approach enables us to iterate and innovate faster in ways that are lighting up every part of the company.
AI momentum across the full stack
These stories of how people are using AI are the best measure of progress. To understand the scale at which people are adopting AI, there is another great proxy — tokens, the fundamental units of data our models process.
Two years ago, we were processing 9.7 trillion tokens a month across our surfaces. Last year at I/O, that grew to roughly 480 trillion tokens. Fast forward to today, that number jumped 7x to over 3.2 quadrillion per month.
It tells an important story about our products and how others are building as well:
- Over 8.5 million developers are now building new apps and experiences with our models monthly.
- Our model APIs are now processing roughly 19 billion tokens per minute.
- Over the past 12 months, over 375 Google Cloud customers each processed more than one trillion tokens.
Momentum with our products
Today we have 13 products with over a billion users each. Five of those have more than 3 billion users. Our Gemini models are a big reason more people are using our products, and why they're using our products more.
It all starts with Search, which is bringing the benefits of generative AI to more people than any other product in the world. AI Overviews now has over 2.5 billion monthly active users. And AI Mode has been a revelation, our biggest upgrade to Search ever. People love it, and in just a year, it's already surpassed 1 billion monthly active users.
The Gemini app had 400 million monthly active users last year at I/O. Today, we've surpassed 900 million, more than doubling in a year. Daily requests have grown over seven times. To date more than 50 billion images have been generated with our Nano Banana image generation models.
Natural, conversational AI in products
Ask YouTube
People come to YouTube everyday to ask a lot of questions. There's a lot of great videos, but sometimes it's hard to know where to start. Ask YouTube entirely reimagines the experience, making information much more digestible and easy to navigate. You'll see videos that best match your interest, and most importantly, it jumps right to the part of the video most relevant to you.
Voice-powered Docs Live
A new feature called Docs Live takes this to another level. To create a doc with Gemini before, you had to type out a precise prompt. With Docs Live, you can just verbally "brain dump" whatever is on your mind, and let Gemini do the rest.
In the future, you'll be able to create new docs _and_ edit them directly, all with your voice. Docs Live is rolling out for subscribers this summer, and powerful voice capabilities will come to Gmail and Keep then too.
Infrastructure supporting innovation at scale
In 2022, we were spending $31 billion annually in capex. This year, we expect that number to be about six times that, approximately $180 to $190 billion. A key part of this investment is our custom silicon.
A decade ago, we announced our very first commercial tensor processing unit, or TPU, on the I/O stage. Since then, we have transformed how the industry builds for AI. We recently announced our 8th generation of TPUs at Cloud Next. For the first time, we've taken a dual chip approach with specialized architectures for training and inference: TPU 8t and 8i.
- TPU 8t is optimized for large-scale pretraining, and it's nearly three times the raw computing power of our previous generation. With JAX and Pathways, training is no longer constrained by the limits of a single, massive data center. We can now seamlessly distribute training across multiple sites, scaling training across more than 1 million TPUs globally.
- TPU 8i is designed for inference. We have dramatically improved speed at every step. Both chips are more energy efficient, delivering up to two times better performance-per-watt.
Gemini Omni
Gemini Omni is our new model that is capable of generating samples in any output modality from any input. We're starting with video outputs, and over time we'll enable image and text. This new model combines Gemini's intelligence with our generative media models — a huge leap forward in world understanding. We're launching the first model in the Omni family: Gemini Omni Flash.
Gemini Omni Flash is available starting today. You will be able to try it on the Gemini app, Google Flow and on YouTube Shorts. We'll also be rolling it out to developers and enterprise customers via APIs in the coming weeks.
New SynthID updates and partners
Three years ago, we launched SynthID, our watermark that is invisible to the naked eye. Since launch, SynthID has now watermarked over one hundred billion images and videos, along with sixty thousand years of audio assets.
We're adding Content Credentials verification across products. This will show you if the origin of the content was AI or a camera, and if it's been edited with generative AI tools. We're expanding both Content Credentials and SynthID verification to Search and Chrome.
Today, we are thrilled to announce that OpenAI, Kakao and Eleven Labs are adopting SynthID, too.
Gemini 3.5 Flash
Today, we're introducing Gemini 3.5 Flash, our first in a series of models combining frontier intelligence with action.
- When compared to 3.1 Pro, 3.5 Flash is better across almost all benchmarks. It's made huge progress in coding — look at the extraordinary jump in GDPVal.
- 3.5 Flash is a very capable model, at the frontier and comparable to the best models, but it's still very fast. When looking at output tokens per second, it is four times faster than other frontier models.
The new model has been a game changer for us internally at Google. In March we were processing half a trillion tokens a day internally across our AI developer tools, and we've been doubling every few weeks. Now, we're processing more than three trillion tokens a day.
Top companies are processing about 1 trillion tokens a day. If they shifted 80% of their workloads from other frontier models to 3.5 Flash, they'd save over $1 billion dollars annually.
Gemini 3.5 Flash is available for everyone today across our products and APIs. We're also excited for Gemini 3.5 Pro. We are using it internally, it's showing great improvements, and it will be coming next month.
Antigravity 2.0
We're also bringing 3.5 Flash to developers in Antigravity. Antigravity is expanding beyond the coding environment, turning it into a platform to develop and manage cohorts of autonomous AI agents. This includes Antigravity 2.0, a new standalone desktop application that acts as a central home for agent interaction, where anyone can orchestrate agents for all sorts of tasks. And we developed an even more optimized version of Flash: not just 4x but 12x faster than other frontier models.
Gemini Spark is your 24/7 agent
Gemini Spark is your personal AI agent in the Gemini app that helps you navigate your digital life, taking action on your behalf and under your direction.
- It runs on dedicated virtual machines on Google Cloud. And it's 24/7 so you don't need to keep your laptop open.
- It's powered by Gemini 3.5 and the Google Antigravity harness, which allows it to perform long-horizon tasks easily in the background.
- Spark will integrate seamlessly with tools, starting with our own, and in the coming weeks with third-party tools through MCP.
- And you can work with Spark however is most convenient: in the Gemini app or soon, through email and chat.
- On Android, you will be able to view live updates and task progress of agents like Spark through a new UI space called Android Halo, coming later this year. Later this summer, Spark will operate directly within Chrome, acting as your agentic browser across the web.
We're starting to roll out Gemini Spark to trusted testers this week and the Beta is coming to Google AI Ultra subscribers in the U.S. next week.
Search in the agentic era
Today, we're introducing information agents in Search. These are personalized AI agents you can set up to work in the background, 24/7, to find what you need at exactly the right moment, and help you take action. Information agents are rolling out this summer starting with Google AI Pro and Ultra subscribers.
Another way we're building a truly agentic Search is by infusing it with agentic coding capabilities. With the power of Gemini 3.5 Flash and Google Antigravity, Search will build custom experiences just for your individual questions, like dynamic layouts and interactive visuals. These generative UI capabilities will be available for everyone in Search this summer, free of charge.
More from our agentic Gemini era
- Daily Brief is another out-of-the-box agent coming to the Gemini app. It gives you a personalized digest and synthesizes information from your inbox, calendar and tasks to find the most important things to be aware of.
- Google Flow is rolling out a new agent today to everyone that can plan and reason through complex tasks with your inputs, under your control.
- Google Pics is our new AI image creation and editing tool, built on our latest Nano Banana model.
- We also shared more about our intelligent eyewear — audio glasses and display glasses, both launching later this fall.
- Gemini for Science brings together a number of AI tools to help accelerate scientific research, including new experiments on Labs and Science Skills to connect agentic platforms like Google Antigravity to over 30 major life science databases and tools.
As we look across the full stack of innovation, from the infrastructure behind TPU 8i to the frontier capabilities of Gemini 3.5 and Antigravity, it's clear we're firmly in our agentic Gemini era.
中文 (Translation)
I/O 2026:欢迎来到 Gemini 智能体时代
阅读时长 19 分钟
作者:Sundar Pichai,Google 与 Alphabet CEO
编辑注:下文为 Sundar Pichai 在 Google I/O 2026 上演讲的编辑整理稿,并补充了现场宣布的更多内容。
距离上一届 I/O 刚过去一年——这是疯狂出货、技术跃迁、飞速前进的一年。我们正处于 AI 周期中那个"用户想要在日常产品里看到价值"的阶段。过去一年我们一直聚焦于此,今天的发布正是这种聚焦的体现。
距离公司全面转向"AI 优先"已经过去十年。我们依然认为,AI 是推进使命、规模化改善人们生活的最深刻方式。我们采取的是差异化、全栈的 AI 创新路径:从自研芯片与安全底座,到世界级研究与模型,再到触达数十亿人的产品与平台。这种路径让公司每个角落都能更快地迭代与创新。
全栈 AI 加速前进
要理解 AI 普及的规模,token 是一个很好的代理指标——它是我们模型处理的"基本数据单元"。
两年前,我们全产品矩阵每月处理 9.7 万亿 token;去年 I/O 是 约 480 万亿;到今年,这个数字又 7 倍跃升,达到每月 3.2 千万亿以上。
这背后是产品被采用与被构建的速度:
- 每月使用我们的模型构建应用与体验的开发者已超过 850 万;
- 模型 API 每分钟处理约 190 亿 token;
- 过去 12 个月里,375+ Google Cloud 客户各自处理了超过 1 万亿 token。
产品动量
我们今天有 13 款产品月活破亿,其中 5 款超过 30 亿。Gemini 模型是更多人用、更多人更频繁地使用我们产品的关键原因。
搜索 是一切的起点。AI Overviews 月活已突破 25 亿;AI Mode 则是搜索史上最大的一次升级,仅一年月活就突破 10 亿。
Gemini app 在去年 I/O 时月活 4 亿;今天已突破 9 亿——一年翻了一倍多,日请求量增长了 7 倍以上。Nano Banana 图像生成模型累计已生成 500 亿+ 张图片。
在产品里讲"人话"的 AI
Ask YouTube
Ask YouTube 彻底重塑了 YouTube 的信息获取体验:信息更易消化、导航更顺畅。它会给出最匹配你兴趣的视频,并直接跳到最相关的片段。
语音驱动的 Docs Live
全新 Docs Live 把这种体验往前推了一步。以前用 Gemini 创建文档要敲出精确的提示词,现在你可以直接"动口"把脑子里的内容倒出来,让 Gemini 接手完成。
未来,你将能用语音创建并直接编辑文档。Docs Live 将在今年夏天向订阅用户推出,强大的语音能力随后也会登陆 Gmail 与 Keep。
支撑规模化创新的基础设施
2022 年我们的资本开支约 310 亿美元;今年这个数字预计是 6 倍——约 1800-1900 亿美元。其中最关键的投资之一就是自研芯片。
十年前我们在 I/O 舞台发布了第一代商用 TPU。如今我们在 Cloud Next 上发布了第八代 TPU——首次采用"双芯片"路线,针对训练与推理做专门架构优化:TPU 8t 与 TPU 8i。
- TPU 8t 面向大规模预训练,算力是上一代的近 3 倍。借助 JAX 与 Pathways,训练不再被单数据中心的天花板锁死,可跨多地调度,全球可调度 TPU 超过 100 万。
- TPU 8i 面向推理。每一步都做了速度优化。每瓦性能也提升到上一代的 2 倍。
Gemini Omni
Gemini Omni 是我们的新模型:可以从任意输入生成任意模态的输出。我们先从视频开始,未来会扩展到图像与文本。它把 Gemini 的"理解世界"能力与我们的生成式媒体模型结合在一起,是世界理解能力的一次巨大飞跃。Omni 系列首发的是 Gemini Omni Flash——今天起即可使用,可在 Gemini app、Google Flow 与 YouTube Shorts 中体验。几周内也会通过 API 向开发者与企业客户开放。
SynthID 新进展与新伙伴
三年前我们推出了肉眼看不可见的水印 SynthID。至今,SynthID 已在 1000 亿+ 张图片和视频上添加水印,音频资产覆盖 6 万年之久。
我们把 Content Credentials 验证集成到更多产品里,告诉你"内容来自 AI 还是相机""是否被生成式工具编辑过",并把这一能力扩展到 Search 和 Chrome。
今天我们非常高兴地宣布:OpenAI、Kakao 与 Eleven Labs 也将采用 SynthID。
Gemini 3.5 Flash
今天我们发布 Gemini 3.5 Flash——"前沿智能 × 行动能力"系列的首款模型。
- 与 3.1 Pro 相比,3.5 Flash 几乎在所有基准测试上都更强,尤其在编程上大幅跃升——GDPVal 表现尤其突出。
- 3.5 Flash 是面向真实业务的旗舰级模型,同时保持了"快"。输出 tokens/秒是其他旗舰模型的 4 倍。
它已经在 Google 内部发挥了关键作用。3 月时我们内部 AI 开发者工具日均处理 0.5 万亿 token;几周翻一倍;今天已超过 3 万亿 token/天。
一家头部公司一天大约处理 1 万亿 token。如果把其中 80% 的工作负载从其他旗舰模型切到 3.5 Flash,一年能省下超过 10 亿美元。
3.5 Flash 已在我们的产品与 API 全面可用。3.5 Pro 已在内部使用,表现亮眼,下月发布。
Antigravity 2.0
我们也把 3.5 Flash 带给 Antigravity 上的开发者。Antigravity 不再只是编码环境,而是演化成"开发与管理大批自主智能体"的平台。这就是 Antigravity 2.0:一款独立的桌面应用,充当智能体交互的中枢,让每个人都能编排智能体完成各类任务。我们为它专门定制了更快版本的 Flash——比其他旗舰模型快 4-12 倍。
Gemini Spark:你的 24/7 智能体
Gemini Spark 是 Gemini app 里的个人 AI 智能体,帮你穿梭在数字生活里、按照你的方向采取行动。
- 它跑在 Google Cloud 的专属虚拟机上,7×24 小时待命;
- 由 Gemini 3.5 驱动、使用 Google Antigravity 框架,能在后台轻松处理"长链路"任务;
- 与工具无缝集成——先从我们自己的工具开始,未来几周通过 MCP 接入第三方工具;
- 你可以在 Gemini app、邮件、聊天里以最方便的方式与它协作;
- Android 上将通过全新的 Android Halo UI 空间查看 Spark 等智能体的实时更新与任务进度(今年晚些时候)。今年夏天,Spark 将直接进入 Chrome,成为你在 Web 上的"智能体浏览器"。
本周开始向可信测试者推送;下周向美国 Google AI Ultra 订阅用户开放 Beta。
智能体时代的 Search
今天我们推出 Search 中的信息智能体(information agents)——你可以设置这类个性化 AI 智能体,让它们 7×24 小时在后台工作,精准在你最需要的时刻找到所需并辅助决策。今年夏天先在 Google AI Pro 与 Ultra 订阅用户中开放。
我们还把"智能体式编程"注入 Search:借助 Gemini 3.5 Flash 与 Google Antigravity,Search 会当场为你的问题构建定制化体验——动态布局、交互可视化等。生成式 UI 能力将在今年夏天向所有人免费开放。
智能体 Gemini 时代的更多发布
- Daily Brief 是 Gemini app 中即将上线的开箱即用智能体——汇总你收件箱、日程与任务里的关键信息,做优先级排序、归纳、给出下一步建议,浓缩成一份"扫一眼就能看完"的早间简报。
- Google Flow 今天向所有用户推出新智能体——能基于你的输入做规划与推理,帮你完成从早期头脑风暴到创作、编辑的复杂任务,并支持"vibe code"创意工具。
- Google Pics 是基于最新 Nano Banana 模型的 AI 图像创作与编辑工具,把每个元素当作独立对象处理,让你能精确地创建、替换或修饰细节。
- 智能眼镜:音频眼镜与显示眼镜,今年秋季推出。
- Gemini for Science 把 Gemini 深度推理/研究能力与 Deep Think、Deep Research 整合到一起,提供 Science Skills 等实验,把 Google Antigravity 这样的智能体平台与 30 多个生命科学数据库/工具相连。
从 TPU 8i 的底层算力,到 Gemini 3.5 的前沿智能,再到 Antigravity 的智能体能力——毫无疑问,我们已经稳稳站在 Gemini 智能体时代。