工具与项目 3.0 · 值得看 2026-06-07 · X

Google DeepMind:Gemini 3.1 Flash TTS用自然语言控制语音风格

Google DeepMind:Gemini 3.1 Flash TTS用自然语言控制语音风格

回到归档

Google DeepMind:Gemini 3.1 Flash TTS用自然语言控制语音风格

原文链接: https://x.com/GoogleAI/status/2044447560384102592
发布时间: 2026-04-16
作者: @GoogleAI(Google DeepMind)

原文 / Original:

Today we launched Gemini 3.1 Flash TTS, our most expressive and controllable text-to-speech model yet.

This launch [excitement] includes audio tags! 🗣🏷 Audio tags [explanatory] are a seamless way to guide vocal style, pace, and delivery using natural language commands embedded directly in your text. Want a different tempo or tone? [amazement] Just tag the audio to steer the AI-speech output! The model supports 70+ languages (24 of which are high-quality evaluated languages, including: Japanese, Hindi, and Arabic). Watch the audio tags in action in the demo below ↓

译文 / Translation:

今天我们发布了Gemini 3.1 Flash TTS,我们迄今为止最具表现力和最可控的文本转语音模型。

本次发布包含音频标签!🗣🏷 音频标签是一种无缝的方式,可以用直接嵌入文本的自然语言命令来引导语音风格、节奏和表达方式。想要不同的速度或语调?只需给音频打上标签,就能控制AI语音输出的走向!该模型支持70多种语言(其中24种为高质量评估语言,包括:日语、印地语和阿拉伯语)。在下方演示中观看音频标签的实际效果 ↓

来源: Google AI (@GoogleAI)
链接: https://x.com/GoogleAI/status/2044447560384102592