New Artificial Intelligence (AI) tools are bringing video production, transcription, research, and travel planning into widely used applications. That matters if you create content, examine large collections of material, or want to organize a trip without moving through several separate websites. The latest announcements also show that convenient interfaces, actual availability, and affordable pricing do not always advance at the same pace.
From rough draft to finished media
Google is expanding AI video generation with Gemini Omni 1.1 Flash. When extending a scene, the model now considers up to ten seconds of the existing footage rather than only its final second, according to the provider. This is intended to keep characters, movement, and visual style more consistent across newly generated sections. Scenes can be extended in ten-second increments to a total length of up to 40 seconds.
You can also upload as much as three seconds of external video as a style reference. Start and end images can serve as keyframes, meaning fixed images between which the model generates a camera movement. This could be useful for a short product video in which the camera moves in a controlled way from a full view to a detail. The claims about Gemini Omni 1.1 Flash come from Google and have not been independently verified.
A 360p mode is available for rapid drafts. Google says it works up to 60 percent faster and costs one-third as much as producing video at 720p. The listed prices are $0.03 per second at 360p, $0.10 at 720p, $0.15 at 1080p, and $0.30 at 4K. A 40-second draft would therefore cost $1.20 if no charges beyond the stated per-second price apply. Completed videos can be upscaled, meaning converted afterward to a higher resolution, at 1080p or 4K.
The description of the model as broadly cheaper still needs qualification. At 720p, Omni 1.1 Flash costs the same as the previous Omni Flash and more than Veo 3.1 Lite, while its 4K rate matches Veo 3.1 Fast. The clearest price advantage is therefore the new draft mode. The model is available through Google AI Studio and developer documentation, rather than as an explicitly described standard feature for every Gemini user.
Adobe is similarly trying to make AI image editing easier to approach. A new optional Photoshop interface, initially launching in beta, places features such as prompt-based editing, background removal, and image extension in one simplified toolbar. A markup feature lets you draw directly on an image, select areas for recoloring, or use arrows and rough shapes to indicate what should move or be added. Editing consequently depends less on your ability to formulate the perfect written instruction.
Adobe’s other examples include opening closed eyes and putting a hat on someone’s head. Its updated masking feature is intended to understand the context of the whole image, so you no longer have to identify every editable area manually. According to the report on Photoshop’s AI-assisted editor, the new interface remains optional. The report does not provide an independent assessment of the results or the contextual accuracy of the underlying Firefly Image 5 model.
Knowledge becomes easier to use
Gemini Notebook can now access supported books that you purchased through Google Play Books. Its “Expert Intelligence” feature allows you to ask questions about their contents and generate plans, infographics, or AI podcasts from the material. More than 100,000 books from several publishers are supported at launch. Eligible titles display a Gemini Notebook label under the tools section in Google Play Books.
Google demonstrated the feature with practical examples. The tool generated a recipe book from information in Michael Pollan’s Food Rules. It also applied ideas from Kim Scott’s management book Radical Candor to a question about handling a difficult conversation with an employee. Uses like these could help you connect a nonfiction book with a specific workplace situation rather than merely searching for isolated keywords.
The feature does not make ownership of the book irrelevant. You can read the full text inside the notebook, but access remains tied to the purchase. If you share a notebook with someone who does not own the included book, that person can see its title as a source but cannot access the full text or information derived from it. The access controls for purchased books therefore place at least one barrier in front of casual sharing with nonowners.
For companies and other organizations, Cohere addresses an earlier stage of the research process. Parse 5 is a vision-language model, meaning a model that processes both visual and textual information, and converts PDFs, slide decks, and images into Markdown, a simply structured text format. It can preserve HTML tables, positional data, and image descriptions. This can make scanned reports or presentation decks easier to search, summarize, and analyze with other AI tools.
Use through an application programming interface (API), a technical connection between services, costs $1.50 per 1,000 pages, according to the provider. Dedicated Model Vault instances start at $2,500 per month and are therefore aimed more at organizations than individual users. Cohere reports a ParseBench score of 79.2 and places the model ahead of several competing services. However, the report on Parse 5 notes that this figure averages only three of the benchmark’s five dimensions and excludes charts and visual grounding.
Speech becomes usable text faster
Gemini 3.5 Transcribe converts speech into text in real time and automatically recognizes more than 85 languages, according to Google. The model removes filler words such as “um,” corrects verbal slips during a conversation, and formats the text on its own. For recorded audio, it also provides speaker identification and timestamps. One practical use would be turning a multilingual interview into a structured working transcript more quickly.
Google reports a word error rate of 4.0 percent for streaming audio and 2.6 percent for recordings. The company also says latency has improved by 70 percent compared with its Chirp 3 predecessor. These provider figures reveal little about performance with dialects, proper names, poor microphones, or several people speaking at once. Swiss German remains an open question because the source refers only to more than 85 languages and does not identify individual dialects.
The model is available through Google AI Studio and the Gemini Enterprise Agent Platform. It is also already used in Gboard for Android and the Gemini app for macOS, and it is expected to come to Chrome soon. This moves transcription closer to ordinary writing and browsing tools than access through technical interfaces alone. The performance claims for Gemini 3.5 Transcribe are again based on information supplied by Google.
Travel planning moves closer to booking
Google’s AI Mode is expanding from conversational search into a tool that handles parts of travel planning and booking. You can describe your destination and dates, compare current options from more than 300 airlines and travel sites, and build an itinerary. If you are not ready to book, the system can track prices and email you when they change. Flight-price tracking is available in more than 180 countries, according to the report, although Switzerland is not explicitly named.
Points and miles can also be included in a search. Google demonstrates the feature with a trip from Atlanta to Miami, showing suitable nonstop flights and the number of miles required. The provider says this option is available globally. The report does not establish whether every Swiss rewards program or local travel company is fully integrated.
For hotels, AI Mode turns a description of your trip and preferences into a list with reviews and relevant factors. After choosing a property, you select “Continue on Google,” move to an integrated hotel or booking site, choose a room, review details such as the cancellation policy, and pay with Google Pay. The hotel or booking platform remains responsible for the transaction and customer service. The report on the new travel features says hotel booking is initially rolling out only in the United States and in English.
Pros and Cons of the new AI tools
Pros:
- Fewer tool changes – Research, editing, and planning increasingly take place inside applications you may already use.
- Faster drafts – Lower-cost video drafts, automatic transcripts, and simplified image markup can shorten early production stages.
- More practical use of knowledge – Purchased books and processed documents can be connected to specific workplace or everyday tasks.
- Multilingual access – Automatic recognition of more than 85 languages could make international conversations and content easier to use.
Cons:
- Unverified provider claims – Most figures for speed, quality, and error rates come from the companies offering the tools.
- Limited availability – Some features are betas, technical services, or initially restricted to the United States and English.
- Additional costs – Per-second video fees, book purchases, and enterprise subscriptions can accumulate with regular use.
- Continued need for review – Transcripts, book answers, document structures, and travel conditions still need comparison with the original material.
What this means for you
If you are a beginner, start with a clearly limited task. You might create a short video draft at 360p, mark one image change in Photoshop, or ask a specific question about one chapter of a book you already own. Compare the result with the original material before you publish it, share it, or use it for a booking.
If you are an advanced user, you can get more value by connecting the tools into a traceable workflow. A recorded conversation could be transcribed, placed in context with purchased specialist literature, and then turned into a visual summary or short video. Structured conversion may also accelerate research across a large document collection, but benchmark results and pricing need to make sense for your own material.
The implications for Switzerland are mixed. Multilingual transcription and the global search for points or miles are relevant, but the sources do not confirm Swiss German support, local rewards programs, or the specific availability of flight-price tracking in Switzerland. AI Mode hotel booking remains limited to the United States and English at first, while some media tools are offered through Google AI Studio rather than standard consumer products.
These releases make AI more practical mainly by moving it closer to existing content, familiar tools, and real transactions. Their value lies less in one spectacular capability than in reducing the distance between drafting, research, and execution. The unresolved risk is how reliably they perform beyond provider demonstrations and how transparent their costs, source use, and booking conditions remain over time.
Sources
- Gemini Omni 1.1 Flash: Google macht KI-Videogenerierung günstiger und flexibler – Unknown, 2026-08-27
- Google’s AI note-taking app now allows you to interact with books – Unknown, 2026-08-27
- Google’s AI Mode can now track flight prices, help book hotels, and more – Unknown, 2026-08-27
- Googles Gemini 3.5 Transcribe erkennt über 85 Sprachen und filtert Füllwörter in Echtzeit heraus – Unknown, 2026-08-27
- Adobe is adding more AI to Photoshop – Unknown, 2026-08-27
- Cohere Releases Parse 5: A Vision Language Model That Turns Enterprise Documents Into Markdown – Unknown, 2026-08-27


