Warning: simplexml_load_file(): /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/themes/newsup/wpml-config.xml:12: parser error : Opening and ending tag mismatch: key line 9 and admin-texts in /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/plugins/polylang/src/modules/wpml/wpml-config.php on line 123 Warning: simplexml_load_file(): in /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/plugins/polylang/src/modules/wpml/wpml-config.php on line 123 Warning: simplexml_load_file(): ^ in /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/plugins/polylang/src/modules/wpml/wpml-config.php on line 123 Warning: simplexml_load_file(): /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/themes/newsup/wpml-config.xml:13: parser error : Opening and ending tag mismatch: key line 8 and wpml-config in /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/plugins/polylang/src/modules/wpml/wpml-config.php on line 123 Warning: simplexml_load_file(): /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/themes/newsup/wpml-config.xml:13: parser error : Premature end of data in tag key line 7 in /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/plugins/polylang/src/modules/wpml/wpml-config.php on line 123 Warning: simplexml_load_file(): /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/themes/newsup/wpml-config.xml:13: parser error : Premature end of data in tag key line 3 in /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/plugins/polylang/src/modules/wpml/wpml-config.php on line 123 Warning: simplexml_load_file(): /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/themes/newsup/wpml-config.xml:13: parser error : Premature end of data in tag admin-texts line 2 in /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/plugins/polylang/src/modules/wpml/wpml-config.php on line 123 Warning: simplexml_load_file(): /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/themes/newsup/wpml-config.xml:13: parser error : Premature end of data in tag wpml-config line 1 in /home/httpd/vhosts/ai-funghi.com/httpdocs/wp-content/plugins/polylang/src/modules/wpml/wpml-config.php on line 123 Decision Models: What Clef, Jev, and Strands Decide - ai-funghi.com
  • Deutsch
  • English
  • Clef, Jev, and Strands Decider are Artificial Intelligence (AI) models that choose among predefined options instead of composing answers. They are mainly intended to help AI agents select their next step quickly and attach a measurable probability to that choice. This affects you whenever an automated system sorts requests, selects tools, or decides whether a person needs to intervene.

    What are decision models?

    A decision model receives a question with a limited set of possible answers. Instead of producing a paragraph, it might return “support,” “sales,” or “suspected fraud,” along with a probability for each option. It can also handle yes-or-no decisions and numeric scores.

    These models occupy a niche between large language models (LLMs) and traditional classifiers. An LLM can write freely, explain relationships, and call different tools, but it may take more time and does not always return a consistent structure. A traditional classifier, meaning a system specialized for fixed categories, is fast but often needs to be adapted or retrained for each new task.

    Decision models combine the broad language processing of an LLM with a closed answer space. Their output remains machine-readable: another system can accept the selected option, require a minimum probability, or send the case to a person. “Calibrated” means that the stated probabilities should correspond reasonably well to actual success rates; it does not guarantee that any individual decision is correct.

    TypeSafe AI introduced Jev as an early model in this category. Competitors now include Cloudflare’s Clef and Clef-flash, Amazon’s Strands Decider 2B, Fastino’s GLiDE, and GLiNER2.5-Decide. According to TechCrunch’s report on the expanding model category, researchers produced dozens of related models after Jev appeared.

    How do Clef and Strands make decisions?

    Cloudflare released Clef with 27 billion parameters and Clef-flash with 9 billion. Parameters are learned internal weights; their number offers a rough indication of model size but does not establish quality by itself. According to the provider, both models are based on Qwen, accept images as well as text, and return typed probabilities rather than free-form sentences.

    The interface is compatible with Jev, which is intended to make switching easier for organizations already using that model. Cloudflare offers Clef through Workers AI and reports a median response time, also called latency, of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash. These model size and speed figures come from the provider and have not been independently verified in the supplied sources.

    Strands Decider 2B follows the same basic idea with a smaller model. The AWS Strands Agents team released it under the Apache 2.0 license, using Qwen3.5-2B-Base as its foundation. It returns choices, yes-or-no probabilities, and scores with confidence values, but never generates text.

    Strands Decider reportedly has a median runtime of about 115 milliseconds on an RTX 3090. It scored 0.723 on the public JevBench set, according to MarkTechPost’s summary of Strands Decider. Without details about each test task, however, that number does not tell you how reliably the model would sort your inbox, customer service queue, or document collection.

    What can these models do in everyday use?

    Customer service provides one concrete example. A model can assess an incoming message by urgency and responsible team. Based on the probabilities, the surrounding system can route the ticket, trigger an escalation, or hand an uncertain case to an employee.

    Cloudflare’s threat intelligence team is also testing Clef for website classification. In one example, the model rated a domain as a fashion site with 95 percent probability, an online store with 85 percent probability, and a phishing site with less than 1 percent probability. The values can overlap because several characteristics may apply at once.

    Fetching, displaying, and classifying the website took 2.2 seconds in total, according to Cloudflare. The company’s fastest general-purpose language model needed 4.7 seconds for the same workflow and returned only two categories. This comparison between Clef and a general-purpose language model is illustrative, but it remains an internal provider test.

    Other proposed uses include selecting a tool for an AI agent, choosing the next workflow step, and enforcing guardrails. Guardrails are control rules that block risky or unwanted actions or send them for review. An agent is an AI system that gathers information and uses it to perform multiple steps or call external tools.

    Speed matters because an agent may make many small choices during a single task. The sources do not present one consistent performance picture, though: Cloudflare reports that Jev took just over 524 milliseconds, while an overview of Jev and competing decision models gives a range of 70 to 500 milliseconds. Different hardware, tasks, or measurement methods could explain the discrepancy, but the supplied summaries do not describe them in enough detail.

    Pros and Cons of decision models

    Pros:

    • Speed – Short, structured decisions can arrive much faster than detailed responses from a general-purpose language model.
    • Clear options – A predefined answer space makes it easier for another system to route a ticket or select a tool.
    • Measurable uncertainty – Probabilities allow you to set thresholds below which a person reviews the case.
    • Potential cost savings – Jev costs $0.042 per million input tokens, according to MarkTechPost; a token is a small unit of text. The sources provide no directly comparable usage prices for Clef or Strands.

    Cons:

    • Limited answers – The model can only choose within the supplied options and does not explain its reasoning in free-form text.
    • Consequential errors – A fast misclassification can automatically trigger further actions if no review threshold is in place.
    • Hard-to-compare metrics – Providers use different hardware, tests, and workflows, limiting the value of simple rankings.
    • Unsettled language quality – AWS itself points to the trade-off among speed, accuracy, calibration, and understanding different languages.

    A closed answer space makes a workflow more predictable, but not automatically reliable. A confidence score of 90 percent does not prove that the particular decision is right. TypeSafe founder Diogo Almeida therefore disputes the idea that the wave of new releases already amounts to strong competition: in his view, observers underestimate how difficult it is to make these models genuinely smart.

    What does this mean for your work?

    If you are a beginner, start with a narrow task whose outcome is easy to check. Sorting support messages by team and urgency is a reasonable example, as long as uncertain or particularly sensitive cases still go to people. Do not compare speed alone; test already resolved examples to see how often the categories and probabilities are useful.

    If you are an advanced user, you can apply different confidence thresholds to decisions with different consequences. A harmless internal sorting task may tolerate less certainty than blocking an account for suspected fraud. You can also place a decision model in front of a more expensive general-purpose language model so that only complex cases are passed on.

    For Switzerland, the sources provide no information about data center locations, compliance with Swiss data protection requirements, or quality in German, French, Italian, and Romansh. AWS says Strands Decider is small enough to run locally, which may offer greater control over processed data; that does not automatically make a deployment compliant. Organizations considering Cloudflare’s hosted service would still need to determine where data is processed and which contractual terms apply.

    Costs are only partly comparable as well. The sources quote a per-million-input-token price for Jev but provide no equivalent figure for Clef or Strands. Local operation may avoid recurring model usage fees, but it requires suitable hardware and maintenance, whose costs are not quantified in the reports.

    Decision models are a useful specialization for frequent, narrowly defined choices, not a replacement for every language model. Clef-flash, Clef, and Strands Decider show that speed, structured output, and local options can coexist. The unresolved risk is whether their calibration holds up outside provider benchmarks, across multiple languages, and in decisions with serious consequences.

    Sources

    AI-FunghiAI-Funghi

    Image: Rafael Minguet Delgado via Pexels

    © 2024 - 2026 ai-funghi.com | All Rights Reserved | Impressum | Datenschutz