New multimodal Artificial Intelligence (AI) models can process several media types, including text, images, audio, and video. Google, Qwen, and other providers are releasing cheaper, specialized, or openly licensed alternatives. For you, the key issue is therefore not which model is supposedly best overall, but which one can complete your particular task at a reasonable cost.
Why cheaper multimodal models matter
Google presents Gemini 3.8 Flash as a cheaper alternative to Claude. According to Google, the model focuses on complex reasoning, coding, and AI agents; an AI agent can plan tasks and call tools on its own. Gemini 3.8 Flash Cyber arrived alongside it and specializes in finding and fixing software vulnerabilities. For translation or document work, however, the Cyber edition is hardly the obvious place to spend extra money.
Gemini 3.8 Flash has an introductory price of $0.75 per million input tokens and $3.75 per million output tokens. Tokens are small units of text used to meter access through software interfaces. According to the Qwen comparison, those prices are due to double on January 1, 2027. The “half the price of Claude” headline offers a useful signal, but not a complete cost comparison for every task.
Qwen3.8-Omni-Flash sets considerably lower rates: $0.15 per million input tokens and $0.47 per million output tokens. Qwen estimates that one hour of audio costs less than $0.01, while processing 720p video with audio at one sampled frame per second costs $0.20. Neither estimate includes the cost of the response, and both are vendor claims rather than independently verified total costs.
A low price per token does not automatically make a model cheaper to use. Long responses, repeated attempts, and video processing can consume more resources than a short text summary. A comparison of different AI models therefore offers a sensible rule: use fast models for simple work and more specialized models for complex tasks.
What separates the available options
Qwen3.8-Omni-Flash can understand audio and video together, reason about them, and use tools. Qwen lists editing a vlog, translating short videos, and summarizing movies as specific applications. Its context window covers one million tokens, meaning the amount of material the model can consider in one operation. According to the provider, it approaches Gemini 3.8 Flash on combined audio and video tasks.
Access is available through Qwen Studio, Qwen Cloud, or an application programming interface (API), a connection that lets another service call the model. Open Qwen-MM plug-ins add video editing, speaker recognition, PDF video notes, and reusable workflows to existing agents. Qwen-Live Harness also connects a camera and microphone for real-time interaction. In this case, “open” describes the plug-ins and does not automatically mean that the entire online service runs locally or avoids sending data to a provider.
For conversations across language barriers, Qwen3.8-LiveTranslate is the more focused option. According to Qwen, the model understands 60 languages, speaks 29, and reduces average delay from 2.8 to 2.3 seconds. It separates speakers in real time, displays two languages in sync, and is designed to resolve names and specialized terms across longer passages. It is available through APIs on Alibaba Cloud Model Studio and QwenCloud.
That feature set fits a multilingual meeting in which comments need to be translated continuously and assigned to the correct participants. It is potentially useful in Switzerland because of the country’s multilingual environment. However, the sources do not list the individual supported languages, confirm Swiss German recognition, identify data locations, or describe special availability terms for Switzerland.
For transcription alone, Grok Voice Transcribe 2.0 offers a narrower alternative. Batch audio costs $0.10 per hour, while streaming costs $0.20 per hour. For short phrases across 19 languages, the provider reports that the word error rate fell from 20.6% to 6.8%, along with an overall claim of twice the accuracy of version 1.0. These claims have not been independently verified.
A speech-to-text model of this kind may be sufficient for transcribing a recorded interview, without the overhead of a full audio and video system. At a live event, streaming support becomes essential, while bilingual display and speaker separation favor LiveTranslate. The task should determine the model, rather than the length of its feature list.
What price and openness actually mean
Local use can offer more control over data and ongoing costs, but it requires a suitable model format and appropriate hardware. PrismML’s Ternary Bonsai 2 27B occupies 5.93 GB, compared with 53.80 GB for the Qwen3.8 27B parent model in a 16-bit format. It accepts text and images, supports a 262,000-token context, and is released under the open Apache 2.0 license, according to the source.
PrismML says the compact model retains 98.2% of the parent model’s average performance across 20 benchmarks. That figure comes from the provider and does not establish how reliably the model will handle your invoices, scanned documents, or images. It also covers text and images, not the audio and video tasks supported by Qwen3.8-Omni-Flash.
The smaller file size comes from ternary weights, which store model values in a heavily simplified form. This is a type of quantization, meaning that numerical precision is reduced to make a model smaller. A guide to GGUF, GPTQ, AWQ, and EXL2 effectively cautions against treating file containers and quantization methods as the same thing. The right format depends partly on whether you use a Mac, a consumer graphics card, or a production server.
An open license, a compact file, and local execution are therefore three separate properties. A license may permit use and modification even when the available hardware runs the model too slowly. Conversely, an inexpensive online service may be more practical while giving you less control over processing and future prices.
The cheapest model is not necessarily suitable for every kind of output either. TypeSafe AI’s Jev is designed for structured decisions, returning probabilities instead of freely written text. Input costs $0.042 per million tokens, and output tokens are free. That may suit clearly defined selection tasks, but it does not replace translation, transcription, or a readable document summary.
Pros and Cons of cheaper and open models
Pros:
- Lower usage costs – Qwen3.8-Omni-Flash substantially undercuts the stated introductory rates for Gemini 3.8 Flash.
- Useful specialization – LiveTranslate focuses on interpretation, Grok Voice on transcription, and Jev on structured decisions.
- Local options – A 5.93 GB text-and-image model can make local processing more realistic than a 53.80 GB file.
- Broader media handling – Multimodal models can connect documents, images, audio, and video within one workflow.
Cons:
- Vendor-supplied results – The performance, accuracy, and latency figures in the reports have not been independently confirmed.
- Inconsistent pricing units – Tokens, audio hours, video length, and input and output are billed in different ways.
- Unclear data paths – The sources do not fully describe storage location, retention, or compliance with Swiss data protection requirements for cloud services.
- Limited interchangeability – A cheap transcription model cannot analyze images, while a local text-and-image model cannot provide live interpretation.
What this means for you
For beginners: Start with a narrowly defined task and a small, nonsensitive test file. Compare not only answer quality, but also processing time, billing method, and the effort needed to correct errors. A cheap transcript full of mistakes may cost more in practice than a higher hourly rate that produces usable output.
For advanced users: Separate workflows by media type and difficulty. You could use a specialized model for transcription, for example, and send the resulting text to another model for summarization. For recurring image and document tasks, it can also be useful to compare a compact local model with a cloud API, provided your hardware and data protection requirements allow it.
Step 1: Define the task and output
- Decide whether you need to process text, images, audio, video, or several media types together.
- Specify whether the result should be a translation, transcript, summary, or structured decision.
- Determine whether processing must happen live or whether later batch processing is sufficient.
Step 2: Make the costs comparable
- For text, calculate input and output tokens; for audio, compare the hourly rate.
- For video, include resolution, sampled frames, and the cost of the generated response.
- Account for announced changes such as the planned price increase for Gemini 3.8 Flash.
Step 3: Test your own examples
- Use several representative files rather than one unusually easy sample.
- Check names, figures, technical terms, and the assignment of comments to different speakers.
- For local models, verify which format suits your device and whether performance is fast enough.
The new options reduce the cost of individual tasks and make local text-and-image processing more accessible. One model still cannot handle translation, transcription, documents, and structured decisions equally well. Independent confirmation of vendor claims remains limited, while language coverage, data locations, and long-term pricing are still open concerns.
Sources
- Halb so teuer wie Claude: Google veröffentlicht Gemini 3.8 Flash – t3n, 2026-09-19
- Qwen3.8-Omni-Flash: Multimodales KI-Modell mit Agenten-Fähigkeiten fordert Gemini Flash heraus – THE DECODER, 2026-09-19
- Alibaba Qwen Team Releases Qwen3.8-LiveTranslate – MarkTechPost, 2026-09-20
- SpaceXAI Releases Grok Voice Transcribe 2.0 – MarkTechPost, 2026-09-19
- PrismML Releases Ternary Bonsai 2 27B – MarkTechPost, 2026-09-18
- GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained – MarkTechPost, 2026-09-19
- KI-Modelle im Vergleich: Wie du das passende Tool findest – ohne zu viel zu zahlen – t3n, 2026-09-19
- TypeSafe AI Releases Jev – MarkTechPost, 2026-09-19


Image: Eduardo Rosas via Pexels
