• Deutsch
  • English
  • Google is expanding Gemini with music generation, voice-controlled tasks, and more selective video analysis. These new Artificial Intelligence (AI) features affect both private users and people who want to automate routine work. At the same time, current market figures and a failed hiking plan show why more capability does not automatically mean greater reliability.

    Gemini’s update push at a glance

    The most visible addition is Lyria 3.5, Google’s model for generating music. According to the report on the Lyria 3.5 integration, it has been released directly in the Gemini app. You can choose a genre and style, select vocals or an instrumental track, and create shorter or longer pieces.

    Templates for background music and personalized birthday songs are intended to make the first attempt easier. One concrete everyday use would be a short instrumental track for a private video presentation; another would be a customized birthday song. Google says the model provides more expressive vocals and richer arrangements than its predecessor, but those quality claims come from the provider and have not been independently verified in the supplied sources.

    Lyria 3.5 is also available in Google Flow Music, Google Vids, and Google AI Studio. Flow Music is said to offer additional features, while Google Vids places generated music inside a video workflow. Google says Lyria was trained exclusively on licensed content, but it does not identify the specific training data. That leaves limited scope for checking how comprehensive the claim is.

    The report provides no prices, subscription requirements, or detailed geographic restrictions. The supplied sources also do not explicitly confirm availability in Switzerland. Swiss users therefore cannot determine from the announcement alone whether Lyria 3.5 has been enabled for their accounts or what language settings and costs apply.

    Music, voice, and video in daily use

    Google is developing Gemini not only as a chatbot but also as an AI agent, meaning a system that can carry out tasks with some degree of autonomy. Gemini Spark is designed to review emails, collect information from them, reserve restaurants, and check calendars. According to an early look at Gemini Spark, the service had launched in the United States by early September 2026 but was not yet generally available in Germany.

    Spoken instructions are meant to play a larger role than typed prompts. A Google product manager reported using the agent to add appointments automatically and draft text based on the instruction “match my writing style.” In a workplace example, the system could collect relevant details from several emails for an upcoming meeting. That also means giving it access to material that may contain private or business information; the example in the report included sharing an address.

    Another update concerns video. Gemini’s Flash models are no longer supposed to process every video uniformly at one frame per second. Instead, they navigate to the sections needed for a particular request. The summary of agentic video understanding reports a reduction of up to 88 percent in video tokens. Tokens are small processing units into which an AI model divides content.

    In practice, Gemini could load the passage related to your question from a long recording rather than process the entire video at the same level of detail. That may make analysis more efficient, but it does not establish whether the selected scene will be interpreted correctly. Because the source provides only a short summary, it also offers no details about testing, usage costs, or specific availability.

    Pros and Cons of Gemini’s new capabilities

    Pros:

    • Direct access – You can create music inside the familiar Gemini app without necessarily opening a separate music tool.
    • More input options – Spoken instructions may simplify tasks when typing a longer request is inconvenient.
    • Practical automation – Email review, calendar entries, and reservations address common repetitive activities.
    • Selective video analysis – Loading relevant segments can reduce processing compared with analyzing a video uniformly.

    Cons:

    • Unclear availability – The sources provide neither a Swiss launch date nor complete information about regions and subscription tiers.
    • Unknown costs – No prices are given for Lyria 3.5, Spark, or the expanded video analysis.
    • Sensitive data – An agent handling emails, calendars, and reservations may need access to private and professional information.
    • Consequential errors – Plausible recommendations can be wrong or incomplete and should not be confused with verified expert information.

    The expansion is also taking place under market pressure. Similarweb figures show ChatGPT’s share of measured AI chatbot website traffic rising from 52.7 to 55.5 percent over three months, while Gemini declined from 27.8 to 25.6 percent. The year-over-year picture is different: ChatGPT fell from 73.3 percent, Gemini doubled its share, and Anthropic’s Claude increased from 1.9 to 9.3 percent.

    The apparent contradiction comes from the different periods being measured: Gemini gained ground over the longer term but slipped slightly in the most recent period. The chatbot traffic figures also cover websites only. Interactions through Android, the Gemini app, mobile ChatGPT services, and desktop applications are not included. The data therefore reflects competition on the web, not a complete ranking of actual usage.

    A sensible way to start

    If you are a beginner, start with a task whose failure would have limited consequences. You could generate a short instrumental track for a private presentation or ask Gemini to summarize non-sensitive material. Then check whether the style, content, and supporting information fit your purpose before reusing the result.

    Advanced users can get more value by defining tightly controlled workflows. You might test voice commands for calendar entries, request several versions of a text in your own style, or ask the system to locate a particular passage in a long recording. Greater automation makes more sense when you have decided which data the system may access and where human review remains mandatory.

    For Switzerland, the picture remains incomplete based on the available sources. Gemini Spark had launched in the United States and was announced for Germany, but no Swiss date is provided. The reports also do not say whether every new feature will work in German, which Swiss language variants will be supported, or how sensitive data will be handled in the specific service.

    In professional and educational settings, it is therefore useful to distinguish harmless drafts from confidential material. An automatically generated track for an internal prototype carries different risks from access to customer emails, addresses, or calendars. The convenience of voice control does not automatically reduce the amount of data being shared.

    Trust remains the limiting factor

    A hike on California’s Mount Shasta illustrates the consequences of unchecked trust. Three young men said they had used Gemini, among other sources, to plan their route and determine the provisions they needed. What was supposed to be an eight-hour ascent turned into a multiday rescue operation; they did not reach the summit until 7 p.m., attempted to descend in darkness, and one of them was injured.

    The sheriff’s office reported that Gemini had recommended far less food and water than the group required. After two unsuccessful helicopter rescue attempts, emergency workers evacuated the men safely the following Monday. The account of the failed hiking plan also notes that the climbers continued even though standard advice was to turn back if they had not reached the summit by noon. AI was therefore a significant risk factor, but not the only decision in the chain of events.

    The case does not make every Gemini use equally hazardous. A flawed birthday song may be irritating; incorrect advice about water, equipment, or a route can be life-threatening. The greater the potential consequences, the less appropriate it is to treat a smoothly worded answer as the sole basis for a decision.

    Gemini is evolving from a text assistant into a broader tool platform for music, workflows, and video. Its new features may shorten creative and administrative tasks, while prices, Swiss availability, and parts of the training-data picture remain unresolved. The central risk is not simply that Gemini sometimes fails, but that persuasive outputs can receive more trust than their reliability warrants.

    Sources

    AI-FunghiAI-Funghi

    Image: Egor Komarov via Pexels

    © 2024 - 2026 ai-funghi.com | All Rights Reserved | Impressum | Datenschutz