工具与项目 3.0 · 值得看 2026-06-07 · X

Gemini 3.1 Flash TTS:用自然语言精准控制语音风格

Gemini 3.1 Flash TTS:用自然语言精准控制语音风格

回到归档

Gemini 3.1 Flash TTS:用自然语言精准控制语音风格

来源: X/Twitter · @GoogleAI

原文链接: https://x.com/GoogleAI/status/2044447560384102592

English

Today we launched Gemini 3.1 Flash TTS, our most expressive and controllable text-to-speech model yet.

This launch includes audio tags! Audio tags are a seamless way to guide vocal style, pace, and delivery using natural language commands embedded directly in your text. Want a different tempo or tone? Just tag the audio to steer the AI-speech output!

The model supports 70+ languages (24 of which are high-quality evaluated languages, including: Japanese, Hindi, and Arabic). Watch the audio tags in action in the demo below.

中文

今天我们发布了 Gemini 3.1 Flash TTS——我们迄今为止最具表现力、最可控的文本转语音模型。

本次发布包含音频标签功能!音频标签是一种将自然语言命令直接嵌入文本中的方式,可以精准引导语音风格、节奏和表达方式。想要不同的语速或语调?只需在文本中打上标签,就能控制 AI 语音的输出效果!

该模型支持超过 70 种语言(其中 24 种经过高质量评估,包括日语、印地语和阿拉伯语)。点击下方演示查看音频标签的实际效果。

*本文由 AI Field Notes 自动抓取翻译,原始内容归原作者所有。*