中文正文
Google 于 2026 年 4 月发布 Gemini 3.1 Flash TTS,这是迄今为止表现力最强、控制粒度最细的语音合成模型。
核心能力
- 70+ 语言支持:覆盖全球主要市场,提供高保真语音和精确控制。
- 200+ 音频标签:通过在文本中嵌入自然语言命令(如
[whispers]、[happy]、[cautious]),可以精细控制语速、语调、停顿和表达方式。 - SynthID 水印:所有 AI 生成的音频都带有 SynthID 水印,可用于识别 AI 生成内容。
- 提示词框架:
[pacing tag] + spoken text + [expressive tag] + spoken text + [pause tag] + spoken text,所有标签必须用英文。
应用场景
- 无障碍与包容性设计:为视力障碍用户生成自然音频导航、读屏内容。
- 创意与娱乐:有声书、游戏配音、广告旁白等。
- 企业应用:银行系统语音导航、客服语音交互等。
可通过 Google AI Studio 和 Vertex AI 体验。
English Original
Google released Gemini 3.1 Flash TTS in April 2026 — the most expressive and granularly controllable text-to-speech model yet.
Core Capabilities
- 70+ language support: High-fidelity speech across major global markets.
- 200+ audio tags: Embed natural language commands directly in text (
[whispers],[happy],[cautious]) for fine-grained control over pace, tone, pauses, and expressiveness. - SynthID watermarking: All AI-generated audio carries SynthID watermark for identification.
- Prompting framework:
[pacing tag] + spoken text + [expressive tag] + spoken text + [pause tag] + spoken text— all tags must be in English.
Use Cases
- Accessibility: Natural audio navigation and screen reading for visually impaired users.
- Creative & Entertainment: Audiobooks, game voiceovers, advertisement narration.
- Enterprise: Banking system voice navigation, customer service voice interaction.
Available via Google AI Studio and Vertex AI.