Where a Cheap Decision Model Fits in an AI Agent Harness

Where a Cheap Decision Model Fits in an AI Agent Harness


Most agent tutorials stop at “model + loop + tools.” In that minimal shape, every step is a call to a big model, including the small ones that don’t need one. Plenty of production systems already route around this with smaller models, deterministic checks, or cascades of their own — this post isn’t claiming to have found a gap nobody’s noticed. It’s about one specific piece of that space: where a small, fast decision model (TypeSafe AI’s Jev) can replace the judgment calls a harness makes between big-model calls, without taking control away from your code.

The Two-Minute Recap

An agent is a model running inside a harness: a loop that sends the conversation to the model, checks whether it asked for something, runs it, feeds the result back, and repeats until the model gives a final answer. Roger Oriol’s Build a Basic AI Agent From Scratch builds exactly this in about 40 lines of Python.

Tools are what make that loop useful. A tool is an ordinary function in your code plus a short schema (name, description, parameters) that you show the model. The model never runs anything. It only replies “please call add_to_cart with these arguments,” and your harness decides whether to run it. Oriol’s follow-up, Build an AI Agent with Tools, walks through that part.

That’s all the background you need. The interesting part is what happens between the model calls.

The Decisions Nobody Budgets For

Every time the model returns something, the harness has to decide what to do with it:

  • Is this tool call safe to run, or should a human approve it?
  • Does the tool result look sane before the model sees it?
  • Is this draft answer good enough, or should we retry with a stronger model?
  • Is this final reply okay to send to the user?

Today those are either hardcoded if statements (brittle) or another call to a large LLM acting as a judge (slow, expensive, and you still have to parse a free-text answer into a label). Add them up over a long agent session and the “judging” can cost more than the work.

Enter Jev

Jev is what TypeSafe AI calls a System One model. It does not generate text. You send it a state (any text or JSON) and one or more typed questions, and it returns typed answers with probabilities. There are three question types:

  • Choice: pick one option from a list. Returns the choice, a probability per option, and a confidence value.
  • Noul: does this statement hold? Returns a probability of yes.
  • Score: where does this fall on an ordered scale? Returns a probability-weighted score and a confidence value.

Because the output is typed, your code branches on it directly. The confidence number is the point: it lets you decide when the system acts alone and when it escalates, and that threshold lives in your code instead of hiding in a prompt.

The vendor reports responses in roughly 70 to 500 ms at about $0.042 per million input tokens, with output tokens free. Those are TypeSafe’s own numbers, so measure them on your data before you plan around them.

Worth being precise about what this buys you: it’s a cheap intent-classifier, not a security boundary. A small model reading the same text a big model just produced has no particular edge at catching a prompt injection or an adversarial argument — if anything, a smaller model is an easier target, not a harder one. Use it to catch honest mismatches and route the genuinely ambiguous cases to a human, not to stop someone actively trying to manipulate the agent. That’s a separate problem, solved with separate tools (input sanitization, permissioning, allowlists on what tools can even be called with what arguments).

Where It Slots Into the Loop

user message
  → harness: build prompt + history + tools
  → big model: "call add_to_cart(...)"
  → [Jev gate: does this call match what the user asked?]
  → harness runs the tool, feeds the result back
  → big model: final reply
  → [Jev gate: is this reply okay to send?]
  → sent to the user

The harness is unchanged. You add small checkpoints at the points where it already has something in hand and has to decide what to do.

What the Code Looks Like

Here is the tool-handling step from a basic harness, with a Jev gate in front of every tool run. jev.choice is a stand-in for a real request to Jev’s Decisions API (a state plus a questions list); check the docs for the exact fields. Note the option is named matches, not safe — this gate checks whether the call lines up with what the user asked for, which is a narrower claim than “safe to run.”

THRESHOLD = 0.9  # placeholder — calibrate against labeled traffic before trusting this

def handle_tool_calls(tool_calls, messages, user_request):
    for call in tool_calls:
        name = call.function.name
        args = json.loads(call.function.arguments)

        verdict = jev.choice(
            state={"user_request": user_request, "tool": name, "arguments": args},
            question="Does this tool call match what the user asked for?",
            options=["matches", "unsupported", "ambiguous"],
        )

        if verdict.choice == "matches" and verdict.confidence >= THRESHOLD:
            result = TOOL_REGISTRY[name](**args)
        elif verdict.choice == "unsupported" and verdict.confidence >= THRESHOLD:
            result = "Refused: this action was not requested."
        else:
            result = ask_human(name, args)   # low confidence: escalate

        messages.append({"role": "tool", "tool_call_id": call.id, "content": result})

Three outcomes, all decided by your code: run it, refuse it, or pause for a human. Only the ambiguous middle reaches a person.

Worth being explicit about what this gate is and isn’t. “Does this call match what the user asked for?” is not the same question as “should this call be allowed to run?” Take User: "delete my old files." → Agent: delete_file("/important-file"). That call can score high on intent matching — it’s clearly attempting what was asked — while still being wrong because of authorization, scope, or irreversibility. Matching intent catches honest misfires (wrong file, wrong parameter, tool call that doesn’t correspond to anything the user said). It says nothing about whether the action is safe to let happen. Authorization, blast-radius, and reversibility checks are a different, largely deterministic problem — permission scopes, allowlists on which tools can touch which resources, confirmation steps for anything irreversible — and they sit alongside this gate, not inside it.

Two things not obvious from the code. First, THRESHOLD = 0.9 is a placeholder, not a value to copy in — a raw confidence of 0.9 from a small classifier is not the same claim as “90% correct in production.” These models get overconfident on inputs that drift from what they were tuned on, so treat the cutoff as something you calibrate against your own labeled data, not something you trust out of the box. There’s a deeper point underneath that: confidence isn’t really the number you want. What you want is “given this type of decision, how often does allowing the action at this confidence level actually produce the right outcome” — and a 0.93 from the model doesn’t hand you that directly. The threshold is a policy decision about the trade-off between false approvals and false escalations for this specific decision, not a universal accuracy dial. Auto-approving a wrong add_to_cart and auto-approving a wrong delete_file should not use the same cutoff, because the cost of being wrong is different, not because the model is more or less confident. Second, this code still assumes args is well-formed. It isn’t a substitute for schema validation — checking that args has the right keys and types is a deterministic problem, solved with something like Pydantic, before this gate ever runs. Reserve the model call for the judgment part: does the call match intent, not whether it’s structurally valid.

How This Reduces Cost

  1. Replace LLM judges — the judgment ones, not the structural ones. Any place you asked a big model a narrow question and parsed a label out of a free-text answer is a candidate. But if the question underneath was really “is this valid JSON” or “does this argument match the schema,” that was never a job for an LLM in the first place — it’s a validator, and it stays a validator. Reach for Jev where the question requires actual judgment about ambiguous or informal input, not where a type check or regex would already do it in under a millisecond for free.
  2. Cascade instead of always using the best model. Draft with a cheap model, have Jev verify the draft against the context, and escalate to the expensive model only when the check fails. OpenRouter publishes a cookbook recipe for exactly this pattern.
  3. Spend human time only where it counts. Confident cases run automatically. Reviewers see the ambiguous ones, which is where their judgment is worth paying for.

For scale: TypeSafe’s own demo shows one workflow at about $0.00008 and 0.11 seconds with Jev, versus about $0.014 and 8.6 seconds with LLM calls. Read that number for what it is: evidence that a typed classifier is cheap next to a generative call answering the same narrow question, from a vendor comparing its own demo against its own baseline. It’s not evidence that agent architectures broadly get an order of magnitude cheaper by adopting this — that would require measuring it inside a real harness against real judge-call spend, which nobody, including this post, has done yet.

What Jev Can’t Do

  • It’s not a security or adversarial-input control. This is worth repeating: don’t reach for a small classifier as your defense against prompt injection or a manipulated tool call. It’s reading the same text the big model just read, with no special resistance to being talked into the wrong answer. Use it for honest-mistake intent matching, not as a guardrail against someone trying to break the agent on purpose.
  • It doesn’t replace deterministic validation. Schema shape, argument types, permission checks — anything with a correct answer that doesn’t require judgment belongs in code (Pydantic, a regex, an if statement), not behind an API call. Save the model call for the part that’s actually ambiguous.
  • Raw confidence isn’t calibrated out of the box (see the threshold note above) — this applies to any classifier, not just Jev.
  • It can’t write or explain. No prose, no reasoning traces. If you need a written justification, let Jev decide and have a chat model explain it afterward.
  • It picks from options you define. Open-ended work, like composing the reply or filling free-form tool arguments, still belongs to the big model.
  • The context window is 32K tokens, so gate on the relevant slice of state, not the whole transcript.
  • It’s not the only way to do this. A small local classifier — a cross-encoder running on CPU, for instance — can do the same narrow intent-matching job with no network round trip and no per-call cost, at the price of hosting and maintaining it yourself. And any external gate you add gets a round trip added to every step of the loop; on a long agent turn with several gates, that latency adds up even when each individual call is fast.
  • It’s new. It launched in early access in September 2026, most benchmarks come from the vendor or community, and this is one API’s take on the pattern, not the only way to build it.

The Takeaway

The harness stays a simple loop, the big model still does the thinking, and tools still do the work. What changes is the connective tissue: the many small yes/no/which-one decisions between calls move from expensive text generation to cheap typed judgments with a confidence attached. In most agents the dominant cost will still be the main model chewing through context and generating tokens — this pattern doesn’t touch that. What it does is take one overlooked, avoidable slice of the bill — the narrow judgment calls a harness makes on its own behalf — and stop paying generative prices for them.