工具与项目 3.0 · 值得看 2026-06-07 · X

Gemini 3.1 Flash TTS:支持自然语言音频标签控制的语音合成 API

Gemini 3.1 Flash TTS:支持自然语言音频标签控制的语音合成 API

回到归档

Gemini 3.1 Flash TTS:支持自然语言音频标签控制的语音合成 API

English

Today we launched Gemini 3.1 Flash TTS, our most expressive and controllable text-to-speech model yet.

This launch includes audio tags! 🎙🎶 Audio tags are a seamless way to guide vocal style, pace, and delivery using natural language commands embedded directly in your text. Want a different tempo or tone? Just tag the audio to steer the AI-speech output!

The model supports 70+ languages (24 of which are high-quality evaluated languages, including: Japanese, Hindi, and Arabic). Watch the audio tags in action in the demo below ↓

中文

今天我们发布了 Gemini 3.1 Flash TTS,这是我们迄今为止表现力最强、控制最灵活的文字转语音模型。

本次发布包含音频标签功能!🎙🎶 音频标签是一种无缝的方式,可以用直接嵌入文本的自然语言命令来引导语音风格、节奏和表达方式。想要不同的速度或音调?只需为音频添加标签,即可控制 AI 语音输出的方向!

该模型支持 70 多种语言(其中 24 种为高质量评估语言,包括:日语、印地语和阿拉伯语)。在下方演示中观看音频标签的实际效果 ↓