Live dialogue
Fluid and natural live dialogue and translation capabilities, for powerful voice-first applications.
Our most advanced audio models push new frontiers with intuitive inputs, natural expressiveness, and the ability to take action
Best for near real-time voice interfaces with conversational capabilities optimized for both cost-effective, high-volume scale and near real-time reasoning.
Best for speech transcription. Transcribes pre-recorded audio across 85+ languages with high alphanumeric accuracy and word-level timestamps for up to three speakers.
Best for near real-time speech-to-speech translation. Overcomes language barriers across 70+ languages while maintaining the speaker’s natural tone and rhythm.
Best for directing intonation and inflection. Intuitive audio tags give you granular command over style, pace, and tone with unprecedented precision.
Natural and powerful audio models. Helping people communicate, developers build, and enterprises manage business.
Engage in almost real-time conversations. Control with precision. Understand every nuance.
Fluid and natural live dialogue and translation capabilities, for powerful voice-first applications.
Context-aware transcription that identifies who’s talking and keeps pace with natural conversation.
Craft anything from short snippets to long-form narratives, with granular control over style, pace, delivery and performance.
Our audio models generate natural vocals at speed and scale for different developer workflows.
Best for complex reasoning and high-complexity tasks. Thinks and narrates its progress in near real-time to solve difficult, multi-step challenges.
Best for near real-time voice interfaces with conversational capabilities optimized for both cost-effective, high-volume scale and near real-time reasoning.
Best for speech transcription. Transcribes pre-recorded audio across 85+ languages with high alphanumeric accuracy and timestamps for up to three speakers.
Best for near real-time speech-to-speech translation. Overcomes language barriers across 70+ languages while maintaining the speaker’s natural tone and rhythm.
Best for directing intonation and inflection. Intuitive audio tags give you granular command over style, pace, and tone with unprecedented precision.
Explore what you can do with Gemini Audio
Orchestrates multiple agents to solve background tasks, while holding a natural-sounding conversation.
Understands visual input to grasp deeper context during conversations. This allows for richer, more natural, and more relevant responses.
Gemini 3.5 Transcribe handles live language switches and seamless streaming transcription.
Translates multiple languages in a single session, while preserving each speaker’s original intonation, pacing and pitch.
Best for directing intonation and inflection. Intuitive audio tags give you granular command over style, pace, and tone with unprecedented precision.
Building with responsibility at the core
We’ve proactively assessed potential risks during every stage of the development process for these native audio features, using what we’ve learned to inform our mitigation strategies. We validate these measures through rigorous internal and external safety evaluations, including comprehensive red teaming for responsible deployment.
All audio outputs from our models are marked with SynthID, our advanced watermarking technology, allowing you to detect whether speech has been created or edited using Google AI.
The fastest path from prompt to production
AI-powered video creation for work
Get started with cutting-edge AI models
Low-latency, real-time voice and video interactions with Gemini
Build, scale, and govern agents
Deploy specialized agents for product discovery, shopping, and customer service
Understand your world and communicate across languages
Collaborate, create, and communicate all in one place