• Deutsch
  • English
  • Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are Google’s new models for direct voice conversations with Artificial Intelligence (AI). They matter if you prefer talking to AI over typing or want to explore voice features for professional workflows. The notable combination is continuous dialogue, live visual input, background tool use, support for 97 languages, and a comparatively low per-minute price.

    What Gemini 3.8 Live can do

    The two models process speech and respond with speech. Gemini 3.8 Live is designed for fast, fluid conversation, while Gemini 3.8 Live Extended Thinking is intended to spend more time on complex tasks, according to Google. Both can use tools and make calls through an Application Programming Interface (API), a connection that lets software request data or functions from another service, without bringing the entire conversation to a halt.

    They also accept real-time visual input. You can show the system a live image while talking about it instead of taking a photo, uploading it, and typing a separate question. One practical everyday example would be holding a visible object or document page in front of the camera, then using spoken follow-up questions to explore the explanation. The supplied sources do not show how reliably this works with small text, poor lighting, or ambiguous scenes.

    Google also says the models can switch among 97 languages in the middle of a conversation. This fits the company’s broader effort to understand more than translated text, including tone, pace, emotion, and context. In its overview of multilingual AI, Google says its technologies support everyday interactions in more than 300 languages, while Google Translate now covers more than 250.

    Those broader figures do not automatically apply to Gemini 3.8 Live, which is specifically described as supporting 97 languages. A practical workplace example would be a conversation in which colleagues move between two supported languages without opening a separate translation session each time. The sources do not provide a full list or indicate how well the models handle regional varieties such as Swiss German.

    Why conversations may feel more natural

    Traditional voice systems often break an exchange into several stages: they transcribe audio into text, process that text, and synthesize the result back into audio. Google argues that this pipeline can lose pauses, emphasis, overlapping speech, emotion, and other context. Gemini is designed to process audio more directly, which may help it cope with the way people actually speak: rarely in polished sentences and occasionally at the same time.

    The ability to interrupt the model is another part of that more natural experience. Independent software expert Simon Willison describes a basic browser interface in his hands-on account of Gemini Live. It lets the user select a model and voice, add an optional system instruction, and speak over the model while it is responding. That is more useful than politely waiting for a digital monologue to finish when the answer has already gone off course.

    Tool use is also supposed to continue alongside the conversation. In a professional workflow, you could discuss a task verbally while a connected tool performs the necessary function in the background, although the supplied sources do not name a specific business application. The meaningful test is therefore not only whether parallel work is possible, but whether requests, confirmations, and results remain understandable throughout the exchange.

    Google reports a score of 82.6 for Extended Thinking on the Artificial Analysis Speech-to-Speech Quality Index and 97.7 percent on Big Bench Audio. According to the release summary, Extended Thinking ranks first on the former benchmark. These claims come from the launch material or summaries of it and are not independently tested in the supplied reporting. A benchmark score also says little by itself about accents, interruptions, or difficult room acoustics.

    What the price comparison really shows

    Gemini 3.8 Live is available through the Gemini API and Google AI Studio. The stated rates are $0.005 per minute for incoming audio and $0.018 per minute for generated audio. If one hour includes 60 minutes of input and 60 minutes of output, the total is about $1.38. This is usage-based pricing rather than a stated flat monthly subscription.

    By comparison, The Decoder lists OpenAI’s GPT-Live-1 at a minimum of $0.05 per minute, or at least $3 for one hour. Its comparison of Google and OpenAI pricing therefore puts Gemini well ahead on stated cost. The real bill will depend on how much incoming and outgoing audio you use; pauses, unequal speaking time, and any connected services can make an actual session look different from the neat hourly example.

    Lower cost does not necessarily mean more natural dialogue. The same report notes that GPT-Live-1 may offer smoother conversation through full duplex, meaning that it can listen and speak at the same time. That is presented as a possibility rather than a conclusive finding. Google undercuts OpenAI in the stated hourly calculation, but none of the supplied sources offers a controlled test comparing price, response quality, and conversational flow under identical conditions.

    Google says all audio generated by Gemini carries SynthID, a digital watermark intended to mark AI-generated content. The supplied material does not explain how an ordinary user can detect or verify that mark. Its inclusion nevertheless signals that synthetic voices should not be treated as ordinary human recordings.

    Pros and Cons of Gemini 3.8 Live

    Pros:

    • Fluid dialogue – You can speak, interrupt, and ask follow-up questions without treating every exchange as a separate text prompt.
    • Multiple input types – Speech, live visuals, and connected tools can work within the same conversation.
    • Multilingual use – Switching among 97 languages may help when people do not stay with a single language throughout an exchange.
    • Lower comparison price – The published estimate of about $1.38 per hour is below the minimum $3 cited for OpenAI’s model.

    Cons:

    • Incomplete quality evidence – Leaderboard results do not replace independent testing with dialects, background noise, and difficult visual material.
    • Variable charges – Incoming and generated audio are priced separately, so the final cost depends on the shape of the conversation.
    • Unclear product boundaries – Google mentions the API, AI Studio, Workspace, Search, and the Gemini app without specifying every feature available in each interface.
    • Possible dialogue disadvantage – OpenAI’s full-duplex approach may feel more natural during simultaneous listening and speaking, although no direct test is provided.

    What this means for you

    If you are a beginner, the most sensible first step is a limited voice trial in one of the interfaces named by Google, such as the Gemini app or Google AI Studio. Hold a short conversation, interrupt a long answer, and, if appropriate, switch between two languages you can evaluate yourself. This reveals more about whether the system understands you than a leaderboard position alone.

    For the visual feature, start with a clearly visible item and ask several related questions about it. This tests not only the initial interpretation but also whether Gemini correctly refers back to what it has already seen as the conversation continues. Avoid using confidential documents or sensitive images merely because the camera is convenient; the supplied sources provide no details about storage, data location, or privacy controls.

    Advanced users can focus more closely on the workflow: Which tasks are handled adequately by the faster Live model, and when does Extended Thinking provide a noticeable benefit? Willison’s experiment also shows that an optional system instruction and different voice presets can be compared in a test interface. A useful evaluation repeats the same task and considers conversational quality, interruptions, and audio consumption separately.

    For Switzerland, broad language support could be useful in multilingual teams, customer service, and education. However, the supplied sources do not confirm a specific Swiss rollout or assess the new model’s performance in Swiss German, French, Italian, or Romansh. They also provide no information about Swiss data locations or detailed privacy terms, so technical availability alone does not establish suitability for sensitive workplace or educational data.

    Gemini 3.8 Live combines spoken dialogue, visual information, and tool use in an interface intended to feel less like a sequence of isolated commands. Its published price is attractive next to the cited OpenAI alternative, while Google’s quality scores are not a substitute for independent everyday testing. The main open risk is how reliably the system handles dialects, noise, sensitive data, and overlapping speech outside controlled demonstrations.

    Sources

    AI-FunghiAI-Funghi

    Image: Andrey Matveev via Pexels

    © 2024 - 2026 ai-funghi.com | All Rights Reserved | Impressum | Datenschutz