Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new models for direct voice conversations. Although they initially target companies building voice assistants, they may shape how you use Artificial Intelligence (AI) in apps, at work, and on connected devices. Amazon’s nearly simultaneous expansion of Alexa+ into Hindi also shows that multilingual conversations and local service connections are becoming central features of voice assistants.
Why live voice models matter
A live voice model processes spoken input and produces spoken answers almost immediately, without making you visibly move between voice and text. Gemini 3.8 Live therefore belongs to the same broad model category as OpenAI’s GPT-Live family. According to the main report on Google’s release, the new Gemini models can call tools during a conversation, process live visual input, and switch among 97 languages.
The main difference from traditional voice assistants is continuity. Instead of starting another rigid exchange for every command, you should be able to ask follow-up questions, interrupt, and change languages while the assistant retains the previous context. A browser experiment by Simon Willison allowed users to select a model and voice, enter an optional system instruction, and interrupt the model while it was speaking.
This makes voice control more appealing when your hands or eyes are occupied. Voice still does not become the best interface for every job: long sets of numbers, confidential material, and results you need to compare precisely are often easier to handle on a screen. A more natural-sounding conversation is also not evidence that an answer is correct.
What Gemini 3.8 Live improves
Google distinguishes between Gemini 3.8 Live for quick, fluid conversations and Gemini 3.8 Live Extended Thinking for harder tasks that need additional processing. Both models can use tools and an application programming interface (API) in the background; an API is a connection through which an AI system calls functions in other software. The conversation is meant to continue rather than pause for every individual action.
That can be useful when a voice assistant retrieves information from a connected service or carries out a task with multiple stages. You could refine a spoken request while the system is already working in the background, rather than issuing several separate commands. Google also describes real-time visual context, meaning the model can combine speech with a live visual feed, although the supplied sources do not provide a complete list of supported devices or applications.
For measured performance, Google says Extended Thinking ranks first on the Artificial Analysis Speech-to-Speech Quality Index with a score of 82.6. The main report also lists a score of 97.7% on Big Bench Audio, a test of audio understanding. These figures come from the launch material or tests cited by it, and they reveal only part of how reliably the model will handle your accent, background noise, or a specific workflow.
Google offers both versions through the Gemini API and Google AI Studio. The Google DeepMind announcement also points to the Gemini app, Google Workspace, and Search. However, the supplied information does not fully explain which functions are immediately available in each product, country, or account.
What lower prices change
Google charges $0.005 per minute for audio input and $0.018 per minute for generated audio output, according to the sources. Based on the calculation published by The Decoder, an hour of voice conversation with both input and output would cost about $1.38. These are usage-based model prices for service providers, not a stated consumer subscription.
For comparison, The Decoder lists OpenAI’s GPT-Live-1 at $0.05 per minute and at least $3 for an hour. This is not a fully equivalent comparison because the report says GPT-Live-1 supports full duplex. Full duplex means a system can listen and speak at the same time, which may allow more natural interruptions and changes between speakers.
Lower model costs could make voice features more economical for customer service, internal assistance, or educational products. Whether you see any savings depends on whether a provider passes them on and what it charges for devices, connected services, or subscriptions. The sources also do not quantify data use or the cost of additional background processing.
Google marks all generated audio with SynthID, which the provider describes as an embedded digital watermark for AI-generated media. This may support later detection, but it is not a substitute for visible disclosure and, based on the supplied information, does not guarantee that every recording can be identified under all conditions. The sources do not describe an independent evaluation of this particular implementation.
Pros and Cons of live voice assistants
Pros:
- Multilingual conversations – According to Google, Gemini can switch among 97 languages without requiring you to begin another dialogue.
- Smoother interaction – Interruptions and retained context make the experience feel less like a sequence of isolated voice commands.
- Background tasks – Connected tools can keep working while you continue to explain or refine your request.
- Lower model costs – Google’s published audio prices are substantially below those of GPT-Live-1 in the cited comparison.
Cons:
- Uncertain reliability – Leaderboards and audio tests capture accents, noise, and real-life situations only partially.
- Incomplete availability – The sources mention several Google products but do not define the exact feature set for every country and account.
- Privacy questions – The supplied material does not explain how voice recordings, visual input, or data from connected services are stored.
- Conversation quality – The less expensive Gemini option may offer less natural speaker changes than a system supporting full duplex.
What Alexa+ shows about everyday use
Amazon is taking a more device- and service-oriented approach with Alexa+. The assistant is available in early access in India in Hindi and English, can sustain longer conversations, and can retain context. According to TechCrunch’s report on the Indian launch, users can even switch between English and Hindi in the middle of a sentence.
The concrete examples indicate where voice assistants are heading. Alexa+ can order groceries through Amazon Now and turn connected home devices on or off. It also integrates with delivery services, travel websites, a ticket-booking platform, a restaurant reservation app, and music services. Voice AI is therefore becoming less of a standalone question-and-answer tool and more of a control layer across multiple services.
After the testing period, Alexa+ is expected to be free for Prime customers in India; Amazon lists a monthly price of ₹2,000, or $20.85, for customers without Prime. The assistant was previously introduced in the United States and, according to the report, has since launched with local language context in markets including the United Kingdom, Canada, Brazil, Mexico, Italy, and Germany. Neither Swiss availability nor Swiss pricing is stated.
For beginners: Your most useful first step is a limited test with a task whose result you can easily verify. In an available Gemini product, you could check whether a longer voice conversation retains its context and how well interruptions work. With connected devices, start with low-risk actions such as turning something on or off before using voice commands for purchases or reservations.
For advanced users: You can gain more value by combining recurring, multistep tasks with connected services and by setting clear instructions for the conversation. The sources say Google AI Studio provides access for trying the models; Willison’s interface additionally demonstrates selecting a voice preset and entering a system instruction. Track audio input and output charges separately, and keep sensitive information out of experiments while the handling of that data remains unclear.
For Switzerland, Gemini’s broad language support and Amazon’s expansion of localized language context are the most relevant developments. The sources do not confirm a Swiss Alexa+ launch, specific Swiss language variants, local service connections, or privacy terms. Multilingual voice agents are becoming more attractive technically and financially, but their reliability, transparency, and local fit outside the announced markets remain unresolved.
Sources
- Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents – MarkTechPost, 2026-09-15
- Gemini Live audio – Simon Willison’s Weblog, 2026-09-15
- Google Deepmind stellt Gemini 3.8 Live für Sprach-Agenten vor – The Decoder, 2026-09-15
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking – Google DeepMind, 2026-09-15
- Amazon launches Alexa+ in India with Hindi support – TechCrunch, 2026-09-16


Image: Anete Lusina via Pexels
