New models from Google, Cognition, Cohere, and Sakana AI are intensifying the price-performance contest in Artificial Intelligence (AI). That does not make your choice easier: The cheapest model is not automatically the best, and a leading coding model may be the wrong tool for translation or sales forecasting. What matters is the task you need to complete and how dependable the result must be.
What is changing in the model market?
Google is positioning Gemini 3.8 Flash as a fast model for coding, AI agents, and complex reasoning. An AI agent is a system that carries out multiple steps with some autonomy, while reasoning refers to working through multi-stage problems. According to a report on Gemini 3.8 Flash, the model is intended to narrow the gap with Anthropic’s Mythos 5.1 and Fable 5.1 at roughly half the price.
The report does not provide the complete absolute pricing structure, however. You cannot turn that comparison into a reliable monthly bill; it only describes Google’s relative position against Anthropic. This is also Google’s third Flash release within a few weeks, with Gemini 3.7 Flash arriving only three weeks earlier according to the report. Product lines are moving quickly enough to make a static comparison age rather badly.
Google has also released Gemini 3.8 Flash Cyber, which was specifically trained to find and fix software vulnerabilities. At the same time, the company introduced the Fairwind Program, a security initiative for selected partner companies. That specialization does not make the Cyber version automatically better for ordinary chat, writing, or translation tasks.
Cognition is also applying price pressure in the coding category. The company behind the Devin coding agent says SWE-2 scored 50.0 percent on FrontierCode 1.1 Main, less than one point behind Fable 5.1, while costing 64 percent less. These SWE-2 performance and cost claims come from the provider and were not independently verified in the supplied summary.
Which model fits which task?
For coding and agent-based workflows, Gemini 3.8 Flash, SWE-2, and Sakana AI’s Fugu models are the most obvious candidates in this group. One practical case would be a coding assignment in which a model does more than suggest code and must coordinate several steps within the task. The broader rule still applies: A fast chat model can perform worse on complex work than a dedicated reasoning model, as the guide to selecting an AI model also explains.
Sakana AI takes a different approach with Fugu Max and Fugu Ultra v2. Both use learned multi-agent orchestration, meaning a control layer distributes work among several models. Fugu Max routes individual tasks to lean open or specialist models, including NVIDIA Nemotron; its published prices are $2 and $6 per one million tokens, with tokens being small units of text.
Fugu Ultra v2 instead targets maximum capability. Sakana AI reports scores of 48.3 on Chartography and 74.3 on DeepSWE. The published details for Fugu Max and Fugu Ultra v2 do not provide a complete cost calculation for real workflows, and the benchmark results should likewise be treated as provider claims.
For machine translation, Cohere’s North Small Translate is the clearest specialist in this selection. The provider says it covers 50 languages and scores 83.6 on Cohere’s own WMT26 evaluation. It is a Mixture-of-Experts model, in which only selected specialist parts are active for each input: The model has 218 billion parameters, meaning learned model settings, but uses 25 billion of them per token.
The North Small Translate weights are free for noncommercial use. Commercial access is offered through Cohere Model Vault or RWS Language Weaver, although the supplied summary gives no specific commercial prices. For a business translating product descriptions or internal documents, that specialization may be more suitable than a general chatbot. The released information about North Small Translate does not identify the 50 languages or show how well each language pair performs.
For forecasting from measurements collected over time, Google offers TimesFM-3. These time series might include daily sales figures, weather data, or store traffic. The model can also use known future events such as holidays, weather forecasts, and planned discounts, and it returns nine values for each point in time to represent forecasting uncertainty.
Google illustrates this with a retailer forecasting ice cream sales from past sales, cones, syrup, customer traffic, weather, and promotions. TimesFM-3 handles several related measurements simultaneously and, according to Google, requires no additional training for a new task; this is known as a zero-shot approach. The overview of TimesFM-3 describes a model with 330 million parameters, trained on more than one trillion real and synthetically generated data points.
How do you compare price and performance?
Model comparisons become misleading when price, benchmark results, and tasks are mixed together. A low price per million tokens says little about how many attempts you will need before getting a usable result. Likewise, a strong score in one test does not prove that a model will reliably handle your language, data, or workflow.
Step 1: Define the task and quality threshold
- Describe one specific task, such as translating a product description, summarizing a research paper, completing a coding assignment, or forecasting sales.
- Decide which mistakes you can tolerate and when human review is mandatory.
Step 2: Record comparable costs
- Look beyond the listed price and note the required text volume, repeated attempts, and any access conditions.
- Separate free noncommercial use from commercial offers, and do not treat relative claims such as “64 percent lower cost” as a complete calculation.
Step 3: Test your own examples
- Give every candidate the same inputs and assess output quality, speed, and required rework.
- For translation, test individual language pairs; for forecasting, inspect the stated uncertainty; and for agents, check whether intermediate steps remain controllable.
This small practical test tells you more than a leaderboard alone. The sources rely on different measures, including FrontierCode, WMT26, Chartography, and DeepSWE, which cannot be directly compared. Several results also come from provider statements or evaluations designed by the companies themselves.
Pros and Cons of specialist AI models
Pros:
- Better-matched capabilities – Translation, forecasting, and coding models focus on more narrowly defined tasks.
- More pricing options – Google, Cognition, and Sakana AI explicitly emphasize lower-cost performance in their positioning.
- More targeted model selection – Orchestration can route simple work to lean models and demanding tasks to stronger systems.
- Clearer uncertainty – TimesFM-3 produces a range for forecasts rather than one deceptively precise value.
Cons:
- Incompatible tests – Different benchmarks, meaning standardized performance tests, measure different capabilities.
- Incomplete pricing – Relative discounts and token prices do not fully capture access fees, retries, and human rework.
- Provider-dependent claims – Several performance figures were reported by the companies involved and have not been independently confirmed.
- More selection work – Translation, coding, forecasting, and agent workflows may each require a different tool.
What does this mean for your work and Switzerland?
If you are getting started, begin with a frequent task whose result is easy to check. For example, you can process the same short text with your current general chatbot and a translation model, or use a historical dataset with known results for a forecasting test. This lets you identify a real quality improvement before changing an entire workflow.
If you are an advanced user, build a small evaluation set from recurring real tasks. Record not only output quality and token price but also correction time, failed attempts, and the effort needed to supervise agent steps. For forecasting, examine whether the model’s range of possible results is plausible instead of relying only on its central estimate.
A model covering 50 languages is potentially relevant in multilingual Switzerland, particularly for product information, internal documents, and customer-facing text. The supplied source for North Small Translate does not list those languages, so it cannot confirm support or quality for all Swiss national languages. None of the six sources provides details about data location, privacy terms, or general availability in Switzerland.
It is similarly unclear how easily these models can be used in education or smaller businesses. Free model weights for noncommercial translation may be attractive for trials, but they do not automatically provide a straightforward path to commercial deployment. Before confidential documents, source code, or sales records are processed, the relevant access and data conditions still need to be established.
The current competition offers more specialist capability and more visible pressure on prices, but not a simple ranking. Gemini 3.8 Flash and SWE-2 target coding, Fugu distributes complex work, North Small Translate handles translation, and TimesFM-3 focuses on forecasting. The largest unresolved risk is the limited comparability of provider benchmarks, total operating costs, and results under your own working conditions.
Sources
- Google gegen Anthropic: Gemini 3.8 Flash schliesst Lücke – für die Hälfte des Preises – t3n, 2026-09-12
- Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost – MarkTechPost, 2026-09-12
- Googles neues KI-Modell sagt die Zukunft aus Verkaufszahlen, Wetter und Rabattplänen vorher – THE DECODER, 2026-09-12
- Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages – MarkTechPost, 2026-09-11
- Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration – MarkTechPost, 2026-09-11
- Welches KI-Modell passt am besten zu deinen Anforderungen? So findest du es heraus – t3n, 2026-09-11


Image: Google DeepMind via Pexels
