Qwen3.8-Omni-Flash processes text, audio, and video together while promising substantially lower costs than comparable offerings. That matters if you want to analyze videos, transcribe conversations, or review large collections of documents. Several models released around the same time show that competition is shifting beyond raw performance toward price and specialization.
Why the price competition matters
Multimodal Artificial Intelligence (AI) refers to models that process different media, including text, images, audio, and video. These jobs have often become expensive because long recordings consume large numbers of tokens. A token is a small unit into which a model divides its inputs and responses for processing.
According to the report on Qwen’s new omni model, Qwen3.8-Omni-Flash costs $0.15 per million input tokens and $0.47 per million output tokens. The provider estimates that an hour of audio costs less than $0.01. Processing 720p video with audio at one sampled frame per second is estimated at $0.20, excluding the cost of the generated response.
By comparison, Gemini 3.8 Flash is listed at introductory prices of $0.75 for input and $3.75 for output per million tokens. Those prices are scheduled to double on January 1, 2027. These are prices for an Application Programming Interface (API), a connection that lets other services access a model. They do not automatically represent the cost of a finished consumer subscription.
Google positions Gemini 3.8 Flash for coding, AI agents, and complex reasoning, according to the coverage of the new Flash release. An AI agent is a system that plans tasks and can call tools on its own. Gemini 3.8 Flash Cyber is instead tailored to finding and fixing software vulnerabilities. This was Google’s third Flash release within a few weeks, with Gemini 3.7 Flash arriving only three weeks earlier.
What the new models can do
Qwen3.8-Omni-Flash is described by Qwen as the company’s first multimodal model built specifically for AI agents. It can jointly examine audio and video, draw conclusions, and use tools. Qwen’s concrete examples include editing a vlog, translating short videos, and summarizing movies. Its context window, meaning the amount of information the model can consider in one session, holds one million tokens.
Qwen says the model approaches Gemini 3.8 Flash on combined audio-video tasks. A separate summary of the Qwen release also reports 45.7% lower token use on OmniVideoBench. This benchmark is a standardized model test for video tasks. Both performance claims come from the provider and have not been independently verified in the supplied sources.
Access is available through Qwen Studio, Qwen Cloud, and the API. Additional Qwen-MM plugins are meant to give existing agents features such as video editing, speaker recognition, PDF video notes, and reusable workflows. Qwen-Live Harness also supports real-time interaction through a camera and microphone.
Other providers are dividing the market into narrower tasks. Grok Voice Transcribe 2.0 focuses on speech-to-text, which converts spoken language into written text. It handles both stored recordings and live audio streams. According to the provider claims for the transcription model, batch processing costs $0.10 per hour and streaming costs $0.20 per hour.
SpaceXAI claims it doubled accuracy compared with version 1.0 while keeping the same price. The short-phrase word error rate across 19 languages reportedly fell from 20.6% to 6.8%. The summary does not identify those languages or show how reliably the model handles accents, background noise, or specialist vocabulary in real working conditions.
Pros and Cons of cheaper multimodal AI
Pros:
- Lower processing costs – Long audio and video recordings can be analyzed at prices that may support new recurring uses.
- Fewer tool changes – One model can consider sound, images, and text together instead of processing each component separately.
- More automation – Agent features connect analysis with tasks such as editing, translation, and structured note-taking.
- Greater choice – General-purpose models now compete with specialized offerings for transcription, decisions, and individual industries.
Cons:
- Difficult price comparisons – Token, hourly, and subscription prices cover different services and do not always include response costs.
- Provider benchmarks – Many performance figures come from the companies themselves and have not been independently verified.
- Privacy risks – Audio, video, and documents may contain particularly sensitive information about people, businesses, or clients.
- Rapid model turnover – Multiple releases within a few weeks make long-term choices and stable workflows harder.
Cheaper does not necessarily mean more versatile. TypeSafe AI is taking a narrower approach with Jev: the model answers predefined types of questions with probabilities instead of generating open-ended prose. According to the introduction to Jev, input costs $0.042 per million tokens and output tokens are free. That pricing structure is more relevant to clearly defined decisions than to video summaries or creative writing.
Smaller models that can be run locally are also adding price pressure. PrismML says Ternary Bonsai 2 27B occupies 5.93 gigabytes, compared with 53.80 gigabytes for the FP16 version of its parent model. Ternary weights store model values in a heavily reduced form, lowering memory requirements. The provider reports that it retains 98.2% of the parent model’s average performance across 20 benchmarks. It accepts text and images and has a 262,000-token context window; these results are also provider claims.
What the prices mean in practice
For beginners: Start with a narrowly defined task in Qwen Studio or a comparable ready-made interface. You could summarize a short, nonconfidential video or have the system review the audio track from your own vlog. This lets you see whether it reliably identifies speakers, events, and relevant scenes without uploading an entire media archive.
Transcribing a meeting is another practical example. A specialized speech-to-text model may be cheaper and easier to assess than a broad multimodal system. Before using it, you should establish whether everyone involved has agreed to the processing and whether the recording includes names, customer information, or internal decisions.
For advanced users: Compare models using a small, consistent collection of your own test tasks. This could include a video with several speakers, a recording with background noise, and a long document. Evaluate errors, required corrections, response length, and the handling of confidential data alongside price. A cheap model that makes you correct every result is merely applying a creative definition of savings.
Specialized professions are also likely to care about the combination of a general model and a curated search source. OpenAI combines GPT-6 Astra with a legal search index in Astra for Law. The index searches US case law, statutes, and regulations across more than 230 million URLs. In a test conducted by OpenAI, the system passed 54% of 200 questions, compared with 38.7% for GPT-6 Astra using ordinary web search. The report on Astra for Law also describes a trusted-access program with privacy controls such as Zero Data Retention, meaning submitted data is not stored permanently.
What remains unresolved for Switzerland
For users in Switzerland, price and multilingual performance are particularly relevant, but the sources leave central questions unanswered. Grok Voice Transcribe 2.0 was evaluated across 19 languages, yet the specific language list is not provided. It is therefore unclear how reliably it handles German, Swiss Standard German, or conversations that switch between languages.
The sources also provide no details about Swiss data locations, retention periods, or the availability of individual services to accounts in Switzerland. Qwen identifies Studio, Cloud, and API access but does not confirm any special conditions for Switzerland in the supplied information. For schools, businesses, and public organizations, a low token price is therefore not enough; the treatment of uploaded voices, faces, and documents also matters.
Astra for Law further illustrates the limits of regional specialization. Its described index focuses on US law, so it cannot simply be transferred to questions about Swiss law. The service still demonstrates why industry-specific search collections can be useful: they give a general model a narrower and potentially more traceable information base. Its published test results, however, were produced by OpenAI itself.
Qwen3.8-Omni-Flash makes combined audio, video, and text processing considerably more affordable, while Gemini, Grok, Jev, and Bonsai offer different responses to the same cost pressure. For you, the best choice depends less on the lowest headline price than on task fit, reliable quality, and control over data flows. The main unresolved risks are independent verification of performance, regional privacy terms, and the durability of announced prices amid such a rapid release cycle.
Sources
- Qwen3.8-Omni-Flash: Multimodales KI-Modell mit Agenten-Fähigkeiten fordert Gemini Flash heraus – The Decoder, 2026-09-19
- Alibaba Qwen Releases Qwen3.8-Omni-Flash – MarkTechPost, 2026-09-18
- Halb so teuer wie Claude: Google veröffentlicht Gemini 3.8 Flash – t3n, 2026-09-19
- SpaceXAI Releases Grok Voice Transcribe 2.0 – MarkTechPost, 2026-09-19
- TypeSafe AI Releases Jev – MarkTechPost, 2026-09-19
- PrismML Releases Ternary Bonsai 2 27B – MarkTechPost, 2026-09-18
- Neuer juristischer Suchindex: OpenAI zielt mit Astra for Law auf den Rechtsmarkt – The Decoder, 2026-09-18


Image: Amar Preciado via Pexels
