Google is expanding Gemini from a chatbot into an assistant for music, study, video, photos, and office work. The new functions affect you if you create content, use Google Workspace, or manage a large media collection. Several announcements from September 2026 also show that development is moving faster than international availability and answers to some privacy questions.
What makes Gemini a broader assistant
Gemini is Google’s product for Artificial Intelligence (AI), meaning software that analyzes material or generates new content in response to instructions. Until now, conversations with a chatbot have been the main experience. Gemini Spark adds an AI agent: a system that can do more than reply by carrying out several steps across connected services.
One example comes from t3n’s look at Gemini Spark. The agent can review emails and compile information from them, reserve restaurants, or check a calendar. A Google product manager used it for professional and personal tasks and shared information including her address with the tool provided by her employer. That illustrates the potential convenience, but it also makes the cost in extensive data access unusually clear.
The way you interact with the software is changing as well. The new voice features for Gmail, Docs, and Keep are intended to understand the context of a request rather than merely convert speech into text. In Gmail Live, for example, you can ask when your next flight leaves or which messages concern your child’s school activities. Gemini then searches the existing inbox and provides an answer.
The common element is integration with Google’s established software. Gemini is no longer meant to produce only a block of text that you must manually move elsewhere. It is supposed to access email, calendars, photos, and other workspaces and take action inside them. This reduces handoffs between applications, but it also increases the number of places where a mistaken interpretation could have practical consequences.
What creative work and study gain
The latest addition is Lyria 3.5, Google’s music-generation model. According to Google, it produces more expressive vocals and richer arrangements than its predecessor; the supplied sources do not include an independent test of that claim. In the Gemini app, you can select a genre and style, choose between vocals and instrumental music, and create shorter or longer pieces. Templates for background music and personalized birthday songs are designed to make the first attempt easier.
In daily use, you could create an instrumental accompaniment for a presentation or draft a personal birthday song without first learning music-production software. The report on Lyria 3.5 says the model is also available in Google Flow Music with additional features and in Google Vids. Google says Lyria was trained exclusively on licensed content but does not identify the training data. That makes the statement impossible to examine more closely from the information provided.
The result for academic work is also favorable, though much narrower. In a blind test run by the Studyarena platform, students anonymously chose between responses from Gemini, Claude, and ChatGPT in August 2026. Among 6,851 recorded votes for writing tasks, Gemini received 39.6 percent, Claude 31.8 percent, and ChatGPT 29.2 percent. Participants did not know which model had produced an answer when they voted.
The outcome suggests that respondents in this test found Gemini’s answers balanced, complete, and clear. It does not prove that Gemini handles every paper more accurately or provides correct citations. The Studyarena comparison measured which response people preferred, not a universal ranking of academic quality. You still need to check statements against original material and avoid replacing your own voice with a polished but generic answer.
What changes for video, photos, and Workspace
For video, Gemini Flash is supposed to search more selectively for relevant moments. This form of agentic video analysis means that the model navigates through a video and loads only the segments needed for a request. According to the supplied summary, the previous approach ingested video at one frame per second. MarkTechPost reports a reduction of up to 88 percent in video tokens, the small units an AI model processes.
For users, the change could make long recordings easier to handle. Instead of processing an entire video at the same level, the system searches for passages that match the task. That could help when examining a lengthy recording of an event. However, the available source is only a short summary and contains no independent tests of accuracy or examples in which Gemini misses a decisive scene.
Gemini Spark can also access Google Photos. After you connect a photo library, the agent can edit images, curate albums, automatically collect selected shots in shared albums, and turn photos of concert flyers into calendar events. The last case shows how this differs from a basic image search: Gemini is intended to recognize information in a picture and then trigger an action in another service.
According to TechCrunch, Gemini Spark’s Google Photos features are initially rolling out over several weeks to eligible Gemini AI Pro and Ultra subscribers in the United States and in English. Google did not say whether or when the feature would reach additional markets. None of the supplied sources gives prices. As a result, no launch date or specific cost can be established for Switzerland.
Workspace extends these media functions into office tasks. In Gmail, Docs, and Keep, Google wants you to complete more work through conversation rather than simple dictation. The practical value lies less in one spectacular feature than in reducing intermediate steps: asking for information from email, speaking an idea into a document, or continuing to work with a note in the appropriate context. Its reliability will depend on whether Gemini understands that context correctly.
Pros and Cons of Gemini’s expansion
Pros:
- Fewer application switches – Gemini connects voice, email, calendars, photos, video, and creative tools within Google’s environment.
- Practical automation – Reviewing email, curating albums, or moving details from a concert flyer into a calendar can reduce repetitive manual work.
- Accessible starting point – Music templates and spoken requests make several functions usable without specialized software.
- More selective media analysis – According to the provider, video analysis loads only the segments needed instead of processing all material uniformly.
Cons:
- Extensive data access – Many agent tasks require access to sensitive material such as email, calendars, addresses, and private photos.
- Limited availability – Several Spark functions begin in the United States and in English, with no specific information for Switzerland.
- Unverified performance claims – Statements about better music and more efficient video analysis largely come from the provider or brief reports.
- Errors can trigger actions – A misunderstood request can do more than produce a poor chat response, potentially creating the wrong album or an inaccurate calendar event.
What this means for you and Switzerland
If you are a beginner, a sensible first step is a tightly limited task with an output you can easily inspect. You could create an instrumental track in a defined style, ask about emails whose contents you already know, or have Gemini suggest how to organize a small album. Check the result before sharing it, adding anything to a calendar, or using it in academic work.
If you are an advanced user, you can gain more by treating the features as connected workflows. A photographed concert flyer can become a calendar event, a large photo library can be curated, and a long video can be searched for relevant segments. At the same time, limiting access to the services and data actually required for each task reduces unnecessary exposure.
The situation remains unclear for Switzerland. The sources report that Gemini Spark has launched in the United States and was expected in Germany in the third quarter of 2026, but they provide no Swiss date. Photo management is initially restricted to eligible Pro and Ultra subscriptions in the United States and to English. The reports do not establish whether German-language requests will be supported in Switzerland or which functions will be available there.
Caution with sensitive information is also appropriate in education and business. The Studyarena test shows a preference for Gemini’s writing style, but it says nothing about the rules a university applies to AI-generated text. In a workplace, the convenience of an agent should not obscure the amount of personal or organizational context exposed through emails, addresses, appointments, and photos. The supplied sources provide no further details about data processing in Switzerland.
Taken together, Gemini’s expansion tells a clearer story than any individual feature: Google wants to place AI directly where you already write, search, plan, and manage media. The value is easy to see in contained routine tasks, while some additions offer modest convenience rather than a fundamental change. The main open risks are how reliably agent actions perform in everyday use and under what conditions these functions will become available outside the United States.
Sources
- Google bringt Musikgenerierung mit Lyria 3.5 direkt in die Gemini-App – The Decoder, 2026-09-06
- Sprechen statt Tippen: So will Googles Gemini Spark deinen Arbeitsalltag verändern – t3n, 2026-09-06
- Beste KI fürs Studium: Warum Gemini bei Hausarbeiten besser als Claude oder ChatGPT ist – t3n, 2026-09-05
- Google Launches Agentic Video Understanding for Gemini Flash Models, Cutting Video Tokens by Up to 88% – MarkTechPost, 2026-09-05
- Google’s Gemini Spark can now manage your Google Photos library – TechCrunch, 2026-09-04
- Google Workspace: Die neuen Sprach-KI-Funktionen für Gmail, Docs und Keep – t3n, 2026-09-04


