Your Agent Pays Frontier Rates Just to Decide ‘Should I Call This Tool.’ A New Model Class Answers in Milliseconds and Never Writes a Word.


blue circuit board

On September 15, a startup called TypeSafe shipped a model that cannot write a sentence. Sixteen days later it had four open-weight competitors. That is the fastest a new model category has gone from one proprietary API to commodity that I can remember, and almost nobody is talking about what the thing actually does or where it belongs in your stack.

The category is decision models, and the pitch is narrow on purpose. A decision model does not generate prose. You hand it a block of state (a support ticket, a tool call it is about to make, a chunk of retrieved context) and a bounded question, and it returns a typed answer with a probability for every allowed option. Yes or no, with the probability of yes. Pick one of these five routes, with a score for each. Rate this against a rubric from one to ten. That is the whole job. No tokens streamed, no JSON to parse, no “as an AI language model” preamble to strip.

If you are building agents, you are already making these decisions hundreds of times per run. The question is what you are using to make them. For most teams the answer is a full frontier model, billed at frontier rates, generating a paragraph of reasoning to arrive at a single bit.

What the model returns instead of text

The mechanical difference is the reason this works. A normal large language model answers autoregressively: it predicts one token, feeds it back in, predicts the next, and so on. To get a yes/no out of it you generate text and then parse the text, which is slow and occasionally wrong in ways that break the agent loop.

A decision model skips generation. AWS, whose Strands Decider 2B shipped October 1, takes a pretrained Qwen3.5-2B torso, removes the language-model head, and bolts on a small pointer head (just over a million parameters) that scores each offered option in a single forward pass. Cloudflare’s Clef and Clef-flash, released the same day, use a prefill-only, non-autoregressive step: the frozen backbone runs one pass, then a joint schema head routes evidence to each question and scores every option at once. Cloudflare’s models answer three fixed question shapes, a boolean (probability of yes), a choice (pick one named option, with per-option probabilities and a confidence value), and a score (a probability-weighted rating against a rubric).

Because there is no token-by-token generation, latency collapses. Clef-flash returns in about 39 ms on Workers AI; Strands Decider 2B lands around 115 ms on a single RTX 3090, per the numbers in this side-by-side comparison. Compare that to a frontier model generating even a short reasoning trace before you can read its decision.

Sixteen days from one closed API to four open clones

TypeSafe, founded by Diogo Almeida, introduced Jev on September 15 as the first of what it calls System One models, fast structured decisions for software rather than open-ended text. Jev is closed: no published weights, no parameter count, a 32,000-token text-only context, and access through TypeSafe’s API and OpenRouter at $0.042 per million input tokens with free output. TypeSafe also published JevBench, a benchmark for decision-only models that hands the model state plus a rubric and checks the typed answer.

Then October 1 happened. Within a single day, Cloudflare open-sourced Clef (27.4B) and Clef-flash (9.4B) under Apache 2.0 on Workers AI and Hugging Face; AWS open-sourced Strands Decider 2B along with its training data and scripts; and Perplexity launched a Decisions API on pplx-decider-v1-27b, also Apache 2.0, also on Hugging Face. Several of the new models are deliberately Jev-API compatible, which tells you how fast the industry decided this was a standard interface worth cloning rather than a product worth paying a premium for.

Model Params License Latency Price (input)
Jev (TypeSafe) undisclosed, closed Proprietary API ~524 ms median $0.042/M
Clef (Cloudflare) 27.4B Apache 2.0 ~209 ms $0.24/M
Clef-flash (Cloudflare) 9.4B Apache 2.0 ~39 ms $0.09/M
Strands Decider 2B (AWS) 1.9B Apache 2.0 ~115 ms Free (local)
pplx-decider-v1-27b (Perplexity) 27B Apache 2.0 n/a $0.04/M

The capability claims are real but they are vendor-run, so read them the way you would read any benchmark a company publishes about its own model. Perplexity reports its decider scoring 85.71% against Jev’s 84.51% across its own eleven-benchmark panel of 7,210 samples; AWS reports Strands Decider ranking third of 33 in its size class on JevBench v1.4.2. Those are leaderboard numbers on someone else’s rubric. What matters is how a decider performs on the specific questions you intend to ask it.

Where it belongs in the stack, and where it does not

AWS is refreshingly direct about the fit. Strands Decider is built to sit inside the agent loop and answer bounded questions in milliseconds: model routing, tool selection, evaluations, guardrails, memory and context management, policy classification. The pattern it names is a hybrid agent, where the large model makes the hardest calls and the decider handles the rote ones, cutting latency and token spend. Put plainly, the decider answers “should we?” and the frontier model answers “how.”

It is equally direct about the limits. Because the model generates no text, AWS says it is unsuited for coding, chatbots, and document summarization, and its single-pass design makes it worse at genuinely hard reasoning than a model that can think out loud. A decision model does not replace your reasoning model. It gates it. That distinction is the entire design decision, and getting it wrong is how teams will misuse these.

The thing a decider actually replaces is the habit of pointing a frontier model at a yes/no. That habit is what made agent pipelines expensive and slow in the first place, the same way an opaque prompt router quietly decides which model your request hits. Every tool-gate, every “is this retrieved chunk relevant,” every “does this output violate policy” was a round trip to a model priced for essay writing.

The cost and latency math

Here is where it stops being abstract. Route a tool-gate decision through a flagship and you pay somewhere between $2 and $10 per million input tokens, plus generated output tokens, plus the time to generate and parse them. Route it through Perplexity’s decider and you pay $0.04 per million input, zero for output, and get the answer in a non-autoregressive pass. Run Strands Decider 2B locally and the marginal cost of a decision is roughly electricity.

Multiply by the number of decisions in a real agent run. An orchestrator managing a dozen sub-agents makes thousands of bounded choices per task, and the failures in those systems live in the handoffs, not the headline model. Moving the rote decisions onto a sub-penny model is both a cost lever, like watching your cache hit rate, and a reliability lever, because a model that can only return an allowed option cannot hallucinate a nineteenth route.

I spent two decades in IT operations, and the shape of this is familiar. You never asked a senior engineer to read a paragraph to decide whether to page someone at 3am. You set a threshold on a signal and let the threshold fire, and you reserved human judgment for the cases the threshold flagged as ambiguous. Using a 2-trillion-parameter model to decide whether to call a weather API is the enterprise equivalent of paging the director to ask whether a disk is full. The right tool for a bounded decision is a cheap, fast, calibrated classifier, and that is exactly what this category is.

A probability you can threshold beats “the model decided”

The governance upside is the part I did not expect to care about and now think is the real story. When a frontier model makes an agent decision, your audit trail is a paragraph of plausible prose. When a decision model makes it, you get a probability per option and, in Clef’s case, an explicit confidence value. You can set a threshold: act automatically above 0.9, escalate to a human between 0.6 and 0.9, refuse below. That is a defensible control, and it is the same confidence-gated escalation any decent ops runbook already uses.

The catch is calibration, and it is a real one. A probability is only useful if it means what it says, and the published numbers are the vendors’ own, measured on their own panels. Before you let a decider auto-approve anything, you have to check its calibration on your data: feed it a few hundred of your actual decisions with known answers and see whether “0.9” is right nine times out of ten or six. A sloppy question schema produces confident garbage just as readily as a confident correct answer, because the model will always return one of the options you offered, even when none of them fit. The discipline that makes this safe is the same one that makes any effort or routing dial safe: measure on your own workload, do not trust the leaderboard.

When to add one

If you are running agents in production, pick one bounded decision that currently flows through a frontier model, most likely tool gating or relevance filtering, and A/B it against a decider on your own logged cases. Score accuracy and calibration, not vibes. If the decider holds up, you have cut the latency and the cost of that step by an order of magnitude and gained a threshold you can audit.

The open releases make this low-risk to try: Strands Decider runs on a single consumer GPU with weights and training data you can inspect, so there is no vendor to commit to before you know it works. What you are buying is not a smarter agent. It is a cheaper, faster, more legible answer to the thousands of small questions your agent was already answering the expensive way. Sixteen days of competition just made that answer nearly free.

Ty Sutherland

Ty Sutherland is the Chief Editor of AI Rising Trends. Living in what he believes to be the most transformative era in history, Ty is deeply captivated by the boundless potential of emerging technologies like the metaverse and artificial intelligence. He envisions a future where these innovations seamlessly enhance every facet of human existence. With a fervent desire to champion the adoption of AI for humanity's collective betterment, Ty emphasizes the urgency of integrating AI into our professional and personal spheres, cautioning against the risk of obsolescence for those who lag behind. "Airising Trends" stands as a testament to his mission, dedicated to spotlighting the latest in AI advancements and offering guidance on harnessing these tools to elevate one's life.

Recent Posts