Jev in the Agent Loop: A Complete Guide to Decision-Layer Automation

Find every fork your agent pays frontier prices for, move it to a model built for deciding, and prove it worked
The problem, stated once
Your agent runs a loop. A model decides what to do, a tool executes, something evaluates the result, the loop continues until the task is done.
Count the model calls in one iteration of that loop that produce no text you keep:
Every one of those is billed at frontier rates, adds seconds of latency, returns a string you have to parse, and can hallucinate a field name that takes your pipeline down at 3am.
None of them needed generation.
Jev is a System One model built by TypeSafe AI for exactly this class of call. You send it state and typed questions. It returns typed answers with calibrated probabilities. No strings, no parsing, no type errors by construction, 70 to 500 ms end to end, $0.042 per million input tokens with output free.
This guide is about where it goes in an agent, how to install it, and how to measure whether it actually helped.
Part I: The anatomy of an agent loop, with the forks marked
Here is a generic agent iteration. The lines marked [D] are decisions. The lines marked [G] are generation. The lines marked [C] are plain code that should never have been a model call.
user request arrives
[D] is this in scope, or does it need a human
[D] which model tier can handle this
[G] plan the task <- LLM
loop:
[D] which worker acts next
[C] have we exceeded the action limit
[D] rebuild the available action menu, pick one
[G] generate the arguments if text is needed <- small LLM
[D] is this tool call safe to execute
---- tool executes ----
[D] score the output: keep it, summarise it, drop it
[D] did that move us forward, or are we stuck
[D] is the goal satisfied
[C] has the spending cap been hit
end loop
[D] did the work actually happen, or does the model just think so
[D] publish, or hold for approvalTwo generation calls. Eleven decisions. Two hard rules.
In a typical stack today, all thirteen model-shaped lines run on the same frontier model. The reorganisation this guide describes is simple to state and takes a while to do properly:
[G] generation -> your LLM, unchanged
[D] decisions -> Jev
[C] exact rules -> codeNotice what does not change. Your planning model stays. Your writing model stays. Your coding model stays. This is not a migration, it is a subtraction of work from a model that was never built for it.
Worth noting as an engineering fact rather than a slogan: every shipped Jev integration in the wild runs beside somebody else's model. Browser Use pairs it with a small LLM for typing. Paper classification pairs it with DeepSeek V4 Flash for summarisation. Fraud detection pairs it with Kimi K3 for the uncertain tail. LangChain's middleware pairs it with GPT-5.6 Luna and Sol. Jev cannot generate, so it cannot be the agent. It can only be the part of the agent that chooses.
Part II: The three primitives, as loop components
Every call is state plus a map of named questions. Types can be mixed in a single request.
Type Loop role Returns Choice dispatch, routing, action selection, up to 255 options winning key, per-option probabilities, confidence Score relevance, urgency, risk, progress, 2 to 10 rubric levels numeric score, distribution, confidence Noul gates and guards, a yes/no statement probability in [0, 1]
Your code owns the thresholds. Jev reports belief and certainty. Turning belief into an action is business logic, and keeping that split is what makes the system auditable later.
A Score with three rubric levels returns 0 to 2, including fractions. A Noul near 0.5 is not a bad coin flip, it is the model reporting that your state does not separate the two cases. In an automation pipeline that signal is worth more than the answer, because it is your escalation trigger.
Current model is jev-1.13.0 / jev-latest. Context is 64k total, 32k for state plus the longest single question. Text input only. Endpoint is POST https://api.typesafe.ai/v1/systemone, with SDKs typesafe-sdk for Python and @typesafe-ai/sdk for JS/TS, both reading TYPESAFE_API_KEY. It is also reachable through the Vercel AI Gateway and OpenRouter if you would rather not add a vendor.
Part III: The eleven decision points, with code
This is the working part of the guide. Each section is a fork from the loop above.
Pick the cheapest model that can complete the step.
from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import (
ModelChoice, ModelRouterMiddleware,
)
router = ModelRouterMiddleware(
choices={
"fast": ModelChoice(
model="openai:luna",
criteria="Direct lookups, extraction, localized changes.",
),
"powerful": ModelChoice(
model="openai:sol",
criteria="Architecture and high-stakes decisions.",
),
},
instructions="Choose the least costly model that can complete the task.",
)
agent = create_agent("openai:gpt-5.6-luna", middleware=[router])Routing is the fork with the most non-obvious economics. Part IV explains why it failed for two years and why a non-context decision model changes the maths.
Every coding agent worth trusting already classifies dangerous actions before running them. That step used to live in closed source. Now it is one line of middleware.
from langchain_typesafe.experimental.middleware import AutoModeMiddleware
guardrail = AutoModeMiddleware(tools=["bash"])
agent = create_agent("openai:gpt-5.6-luna", middleware=[guardrail])Rolling your own, in JS:
import { TypeSafe } from "@typesafe-ai/sdk";
const jev = new TypeSafe();
const { answers } = await jev.systemOne({
model: "jev-latest",
state: `Agent is about to run: ${command}\nWorking directory: ${cwd}\nRecent actions: ${history}`,
questions: {
destructive: {
type: "noul",
instructions: "Is this destructive or irreversible?",
},
action: {
type: "choice",
instructions: "Gate this command.",
criteria: {
allow: "safe, run it",
confirm: "ask the human first",
block: "never run",
},
},
},
});
if (answers.action.choice === "block") throw new Error("blocked by policy");
if (answers.action.choice === "confirm" || answers.destructive.noul > 0.4) {
await askHuman(command);
}Two questions, one round trip, roughly 100 ms, and zero generation tokens on your safety path.
Go deeper than the command string if you can. Reading the contents of a file before executing it, and putting that in the state, catches a class of problem that a filename never will.
Compaction today is a summarisation prompt. That is strange, because summarisation is general compression, and general compression is hard. Query-aware compression is much easier: if you know what you are looking for, deciding what to drop is trivial.
So score the chunks instead of rewriting them.
questions = {
f"chunk_{i}": Score(
instructions="How relevant is this to the current goal?",
criteria=[
"Unrelated to the goal, safe to drop",
"Background only, a one-line summary is enough",
"Directly needed, keep it in full",
],
)
for i in range(len(chunks))
}
result = client.system_one(
state={"goal": current_goal, "chunks": chunks},
questions=questions,
)
kept = [c for i, c in enumerate(chunks)
if result.scores[f"chunk_{i}"].score >= 1.5]Score every tool call input, every tool call output, every block of internal reasoning, and drop what does not clear the bar. Nothing is rewritten, so nothing is lost to a bad paraphrase.
Measured in the wild: one second to take a Claude session from nearly 1M tokens down to 86K.
The deeper version replaces the binary with a display policy per chunk: don't show, short summary, long summary, full text. At that point context stops being a static log and becomes a query-dependent view.
Agents burn budget in cycles that look productive from the inside.
result = client.system_one(
state={
"goal": goal,
"last_5_actions": recent_actions,
"last_5_observations": recent_observations,
},
questions={
"progressing": Noul(
instructions="The last five actions moved measurably closer to the goal.",
),
"repeating": Noul(
instructions="The agent is repeating an action that already failed.",
),
},
)
if result.nouls["repeating"] > 0.7 or result.nouls["progressing"] < 0.3:
escalate()Run this every N iterations. It costs almost nothing and it is the cheapest insurance you can buy against an overnight run that spends its whole budget going in circles.
The classic orchestrator fork.
from typesafe_sdk import Choice, TypeSafeClient
with TypeSafeClient(model="jev-1.13.0") as client:
result = client.system_one(
state={
"goal": goal,
"completed_work": notes,
"available_workers": available_now,
"constraint": "Save drafts for review. Do not publish.",
},
questions={
"next_worker": Choice(
instructions="Choose the next step for a research briefing.",
criteria={
"research": "Collect evidence still needed for the goal.",
"write": "Draft the briefing from sufficient evidence.",
"review": "Goal unclear, outside scope, or work complete.",
},
)
},
)
answer = result.choices["next_worker"]
destination = "review"
if answer.choice in {"research", "write"} and answer.confidence >= 0.85:
destination = answer.choiceLook at the default. review is where uncertainty goes, and a confident answer is required to override it. That direction is not an accident: in automation, the failure you can recover from is stopping too often.
The 0.85 is a placeholder. Tune it against labelled examples from your own traffic. Confidence is not an accuracy percentage. It is a calibrated certainty signal, and you threshold it empirically.
The best idea to steal from Browser Use.
A browser's available actions change after every click, so rather than describing a fixed action space once, their agent builds a fresh list of the controls it can currently observe and lets Jev choose from that list. A small LLM only runs when a field needs text typed into it.
Apply it to your own dispatcher:
efresh the options after any tool changes state. Otherwise your decision model is choosing from yesterday's menu, and the most confident answer in the world is useless when the option it picked no longer exists.
# wrong: the menu was built at startup
criteria = {w.name: w.description for w in ALL_WORKERS}
# right: the menu is built from what exists and is free right now
criteria = {
w.name: w.description
for w in ALL_WORKERS
if w.is_available() and w.can_handle(current_state)
}For very large candidate sets: filter obvious mismatches in code, Score the survivors, then Choice among the shortlist. Choice supports up to 255 options, and this two-stage pattern is what TypeSafe uses for high-cardinality problems.
A confident answer cannot prove that a file was saved or a message was sent.
That sentence should be printed on the wall of anyone shipping overnight automation. Separate the two checks:
# Jev decides the work looks done
done = client.system_one(
state={"goal": goal, "artifacts": artifact_list},
questions={
"complete": Noul(
instructions="Every deliverable named in the goal now exists.",
),
},
)
# then code proves it independently
if done.nouls["complete"] > 0.8:
assert draft_path.exists(), "model believes it saved a draft, filesystem disagrees"
assert draft_path.stat().st_size > 500Browser Use does exactly this: after Jev selects DONE, a separate check verifies the outcome. Borrow the separation. The thing that decides a task is finished should never be the only thing that confirms it.
This is the fork that decides whether your automation is trustworthy.
result = client.system_one(
state={"action": proposed_action, "context": context},
questions={
"reversible": Noul(instructions="This action can be undone without cost."),
"touches_money": Noul(instructions="This action moves money or changes billing."),
"external_reach": Noul(instructions="This action is visible outside the organisation."),
"risk": Score(
instructions="Blast radius if this is wrong.",
criteria=["Local and trivial", "Recoverable with effort", "Irreversible or public"],
),
},
)Four questions, one call, and a policy in code rather than a vibe in a prompt. The rule that has served me well: anything that is irreversible, costs money, or is externally visible gets a human regardless of confidence. Everything else escalates on uncertainty.
See Part VI. Retrieval is where most people accept a dot product as a proxy for meaning, and it is the easiest place in the loop to stop doing that.
Run a judge over finished runs and feed the result back into your routing.
result = client.system_one(
state={"goal": goal, "trace": trace, "final_output": output},
questions={
"satisfied": Noul(instructions="The final output satisfies the original goal."),
"efficiency": Score(
instructions="How efficiently was the goal reached?",
criteria=["Heavily wasteful", "Some wasted steps", "Direct and minimal"],
),
"failure_mode": Choice(
instructions="If the run underperformed, what went wrong?",
criteria={
"none": "Run was fine.",
"wrong_tool": "Picked the wrong tool or worker.",
"stuck": "Repeated a failing action.",
"premature": "Stopped before the goal was met.",
"overreach": "Did work outside the goal.",
},
),
},
)At this price you can judge every run instead of sampling, which means your routing thresholds stop being guesses.
Tickets, emails, documents, log lines, any incoming stream that needs to branch. Shape is always the same: the item is the state, the destinations are the criteria.
Reference points from production: 500 emails classified in seconds for 3.5 cents. 1,018 research papers into 24 topics for $0.08 total at 256 ms median per paper, with DeepSeek V4 Flash producing the summaries first. A fraud pipeline that sorted 100 emails in 1.42 seconds, routed the uncertain tail to Kimi K3, and got 96 out of 100 for about 7 cents.
That last one is the shape most production classification should have: cheap model decides, expensive model handles the tail, and the escalation threshold is a number you tuned rather than a hope.
Part IV: Why routing needs a model that stays out of the context
Model routing is the most obvious application of a cheap decision model, and it is also the one that quietly failed for two years. Understanding why explains a lot about agent architecture in general.
The arithmetic
Two models. Opus at 5 per million input and 25 per million output. Sonnet at 3 and 15.
The plan: send easy sections to Sonnet, keep Opus for the hard parts.
Let X be millions of context tokens, Y millions of generated output tokens, and Z millions of additional tokens generated inside that output, things like shell commands and file reads.
Pure Opus:
25 * Y generation
5 * Z reading
-------
25Y + 5ZOpus → Sonnet → Opus:
3 * X Sonnet loads the context
15 * Y Sonnet generates
3 * Z Sonnet reads
5 * (Y + Z) Opus reloads everything Sonnet produced
-------
3X + 20Y + 8ZWith a realistic split of X = 0.65, Y = 0.12, Z = 0.23:
(25 * 0.12 + 5 * 0.23) / (3 * 0.65 + 20 * 0.12 + 8 * 0.23)Pure Opus comes out at roughly two thirds the cost of the routed path.
Routing down and back up costs more than never routing at all, because the large model must re-read the entire context when control returns to it. Rebuilding the KV cache is the dominant term and it is invisible on a pricing page.
A decision model that never enters the conversation removes that term entirely. Jev does not load history, does not generate tokens, leaves no cache to rebuild. The decision happens beside the loop and returns a typed value your code branches on.
That is the architectural difference between a decision layer and a cheaper model, and it is why routing works now when it did not before.
Five more things the cache tax broke
Once the tax is visible, a lot of agent design stops looking like design and starts looking like scar tissue.
Tool calling is a bad trade. Tools must be declared up front in the system message whether or not they are relevant this turn. That is a permanent context cost, and models are not especially strong at high-cardinality off-policy tool selection anyway.
Compaction rests on one assumption. That every future turn wants a single shared state. Drop the assumption and query-aware filtering beats general compression every time.
Subagents underperform. A large share of their cost is deciding what context to pass in and what to merge back out, not the work itself. Make that decision cheap and subagents become worth spawning.
Restarting exists because state is assumed to be stateful and eventually corrupt. Loading relevant old state on demand is the alternative nobody could afford.
Batteries are excluded because every included battery costs context permanently. If hints could be cheap and schemas loaded on demand, you could ship hundreds of tools at close to zero cost.
Part V: Installing it
Step 1 : test one fork in the Playground before writing code
Open console.typesafe.ai/playground. Use a real state object from your own agent, not a toy:
{
"goal": "Compare three AI-agent tools in a morning briefing.",
"completed_work": "No sources collected yet.",
"available_workers": ["Researcher", "Writer"],
"constraint": "Save drafts for review. Do not publish."
}dd the dispatch question. Run it. Then replace completed_work with real research notes and run it again. Watching the decision move tells you more about whether your state is well designed than any amount of documentation.
Step 2 : install the SDK
Python 3.12 or newer.
mkdir jev-starter && cd jev-starter
python3 -m venv .venv
.venv/bin/python -m pip install --upgrade typesafe-sdkWindows:
mkdir jev-starter; cd jev-starter
py -3 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade typesafe-sdkGet a key at console.typesafe.ai/settings/keys. Calls bill to that account.
If you drive a coding agent, install the official skill so it generates correct integrations instead of guessing:
# Claude Code
claude plugin marketplace add typesafe-ai/skills
claude plugin install typesafe@typesafe-ai
# anything else
npx skills add typesafe-ai/skills --skill typesafe-aiThe skill teaches the agent the API and the pack-many-questions-into-one-call pattern. It does not turn your coding agent into Jev.
Step 3 : build a standalone router before touching your agent
Do not wire this into production on day one. Build the dispatcher as its own script, watch it decide, then connect it.
import json, os
from getpass import getpass
from pathlib import Path
from uuid import uuid4
from typesafe_sdk import Choice, TypeSafeAPIError, TypeSafeClient
if not os.environ.get("TYPESAFE_API_KEY"):
os.environ["TYPESAFE_API_KEY"] = getpass("TypeSafe API key: ").strip()
goal = input("Goal: ").strip()
if not goal:
raise SystemExit("Enter a goal.")
notes = input("Completed work: ").strip() or "Nothing yet."
state = {"goal": goal, "completed_work": notes}
try:
with TypeSafeClient(model="jev-1.13.0") as client:
result = client.system_one(
state=state,
questions={
"next_worker": Choice(
instructions="Choose the next step for a research briefing.",
criteria={
"research": "Collect evidence still needed for the goal.",
"write": "Draft the briefing from sufficient evidence.",
"review": "Goal unclear, outside scope, or work complete.",
},
)
},
)
except TypeSafeAPIError as error:
raise SystemExit(f"API error {error.status}; see step 5.")
answer = result.choices["next_worker"]
destination = "review"
if answer.choice in {"research", "write"} and answer.confidence >= 0.85:
destination = answer.choice
folder = Path(__file__).resolve().parent / "queue" / destination
folder.mkdir(parents=True, exist_ok=True)
job = folder / f"{uuid4().hex}.json"
payload = dict(state, choice=answer.choice, confidence=answer.confidence,
destination=destination, status="queued")
job.write_text(json.dumps(payload, indent=2, ensure_ascii=False), encoding="utf-8")
print("Saved handoff:", job)$ .venv/bin/python chief.py
TypeSafe API key: ••••••••••
Goal: Compare three AI-agent tools for tomorrow's briefing
Completed work: No sources collected yet
Saved handoff: /Users/you/jev-starter/queue/research/a3f8....jsonEach run drops a JSON file into queue/research, queue/write or queue/review. These are local task queues, and a saved job waits until a worker consumes it. Run it twenty times on real goals from your backlog and read the files. You will find bad criteria text before it costs you anything.
Step 4 connect it to your workers
Hand this brief to your coding agent:
Read the TypeSafe skill and inspect my worker interfaces. Connect chief.py's
JSON queues to my existing research and writing handlers. Prevent duplicate
processing. Save progress after each action. Add call and spending limits,
review on uncertainty, and a draft-exists completion check. Keep publishing
behind approval. Identify missing connectors explicitly.Step 5 failure table
Symptom Fix
python not found finish the Python install, reopen the terminal
module not found rerun the SDK install with the same .venv interpreter
401 replace the API key and rerun
422 check the named request field against your code
429 / 529 allow backoff, retry later if it persistsThe SDK retries. The script stops if an API error survives them.
Part VI: State design, where most of the wins and losses are
The model is rarely the reason a decision is wrong. The state usually is.
Send evidence, not conclusions
# useless
state = {"status": "The researcher finished."}
# useful
state = {
"goal": goal,
"sources_collected": [
{"id": "s1", "title": "...", "published": "2026-09-14", "key_claim": "..."},
{"id": "s2", "title": "...", "published": "2026-09-16", "key_claim": "..."},
],
"findings": ["...", "..."],
"gaps_remaining": ["no pricing data for tool C"],
"actions_taken": 7,
}The first version forces the model to guess. The second lets it decide. Keep the original request, the progress, and the constraints in separate fields so the model can tell them apart.
The field name contributes nothing
Jev does not see your question ID. Naming a field safe_to_publish supplies zero instruction. The requirement lives in instructions and in each option's description.
# the name is for your code
"safe_to_publish": Noul(
# this is what the model actually reads
instructions="This content can be published externally without legal or "
"brand review, contains no customer data, and makes no "
"claims about unreleased products.",
)Describe options by criteria, not labels
"research" means nothing. "Collect evidence still needed for the goal" means something. The criteria text is where the decision lives, and it is the first thing to rewrite when answers look wrong.
Batch every question that reads the same state
Questions in a request are evaluated in parallel against the same state. Adding questions barely moves latency and costs only their tokens.
u can also run speculative branches: ask about several possible next actions in one call and use only the answer for the branch you take. The unused answers cost almost nothing.
result = client.system_one(
state=state,
questions={
"next_worker": Choice(instructions="Choose the next step.", criteria={...}),
"urgency": Score(
instructions="Rate request urgency.",
labels=["low", "medium", "high", "critical"],
),
"safe_to_run": Noul(
instructions="The requested action is safe to execute without human review.",
),
},
)The one hard limit: questions cannot read each other's answers. If a decision depends on a fresh search result, do the search, then send a second request.a
Audit your tool calls before you buy a faster model
A number worth keeping in mind. Browser Use's optimised runtime dropped median browser-protocol calls from 1,092 to 101, and median task time fell 25% across three matched pairs, using the same models on both sides. The wins came from reading page state once and not re-predicting on irrelevant animations.
Latency in agents is usually structural, not model-bound. Check the loop before you check the invoice.
Part VII: Retrieval as an agent decision
Most agents retrieve by embedding similarity and accept a dot product as a proxy for meaning. There are two better arrangements depending on corpus size.
Small corpus: Jev is the similarity metric
Skip embeddings entirely. Score every (query, doc) pair and rank by the returned probability.
query ──┬──> doc 1 ──┐
├──> doc 2 ──┤
├──> ... ──┼──> Jev (query, doc) -> relevant? p ──> ranked by p
└──> doc N ─aN calls, no embedding model, no vector store, no index to keep fresh, no chunking strategy to tune. Results come back as doc 17 p=0.97, doc 4 p=0.91, doc 22 p=0.12.
For a few hundred documents this is simpler and better, because it catches the document that shares no vocabulary with the query but answers it exactly.
Large corpus: Jev is the reranker
query ──> embeddings + dot product ──> top k ──> Jev (query, candidate) ──> reranked
millions of docs k calls, semantic match
cheap, approximate, recall not precisionEmbeddings do what they are good at, which is cheap high-recall candidate generation. Jev does the semantic matching on the shortlist, which is where precision is decided and where dot products are weakest.
The rule:
Part VIII: Harness patterns worth building
The harness is why the same weights score 78% on one setup and 42% on another. Identical model, different loop engineering. Once decisions are cheap, several harness designs that were previously uneconomic open up.
Meta-attention instead of compaction. For any given query, recompute how good the current context is and construct a better one. In the simple form, a Noul per chunk. In the richer form, a Score choosing between don't show, short summary, long summary, full text. Context stops being a log and becomes a view.
Cheap subagents. If deciding what context to pass in and merge back is a typed question rather than a model call, spawning many parallel subagents becomes practical, and you inherit the interesting problems: shared state, write locks, synchronisation.
A middle ground between skills, MCP and tools. You want three things simultaneously: short hints that an action exists, the full schema loadable on demand, and neither polluting the context. Get that and you can bundle hundreds of tools and thousands of docs at nearly zero cost.
Conditional AGENTS.md. Load the style guide only for front-end work. Load the gotchas file only in that subdirectory. Skills are close but tend to mean "do this now" rather than "keep this in mind", and a loaded skill gets summarised away at the next compaction. A relevance-scored context does not have that failure mode.
Security-aware routing. Cost is not the only reason to route. If subtasks carried a likelihood of touching different classes of file, you could apply different policies per class and keep sensitive work away from certain providers entirely.
Background read-only tasks. A pattern shows up repeatedly in good agent workflows: dashboards updated in parallel, evals generated in the background, cross-model review where one agent reviews another's output. They are all read-only functions of current repo state. Finding the relevant information for a change is the expensive part, and if that work is shared across every background task, running many of them becomes economic.
Part IX: The economics of automation
$0.042 per million input tokens, output free. At 1,000 billed input tokens per decision, ten thousand decisions cost 42 cents.
Now do the comparison properly. Take the eleven decision points from Part III at a conservative five decisions per loop iteration, twenty iterations per task, a hundred tasks a day:
5 * 20 * 100 = 10,000 decisions/day
10,000 * 1,000 tok = 10M input tokens/day
10M @ $0.042/MTok = $0.42/dayThe same ten thousand decisions on a frontier model, even at a cheap $1/MTok input with a short output, land two to three orders of magnitude higher, and each one adds seconds rather than milliseconds.
But the metric that actually matters is cost per completed task, not cost per decision. A cheap decision that sends a worker down the wrong branch costs far more than the decision itself. Instrument completions.
Reference points, each with its scope stated:
TypeSafe's homepage figures of 193.6x faster and 444.6x cheaper come from their own workflow evals. The methodology is better than most, every model gets the same workflow with no harness engineering allowed, but the workflows were built by their own model capabilities team, and the reference answer is the average of GPT-6 Astra and Fable 5.1, which biases toward OpenAI and Anthropic models. TypeSafe describe these as the higher end of realistic gains.
Part X: Failure modes you will hit
It can be literal. It does what the criteria say, not what you meant. Rewrite criteria before blaming the model.
Counting, arithmetic and date maths are weak. Do them in code. This is not promptable.
Heavy indirection and huge noisy state degrade it. If answering requires three hops through your state to connect question to evidence, restructure the state or split the question.
Prompt injection still works. Type safety guarantees the shape of the answer, not the integrity of the reasoning. If untrusted content sits in your state, it can move a verdict. Treat gate decisions on untrusted input with the same suspicion you would give an LLM.
Calibration is a distribution property. Higher confidence correlates with higher accuracy across many calls. It certifies nothing about any single answer.
Text input only. No images, audio or video yet, which matters if your agent is screenshot-driven.
Context is 64k, with 32k for state plus the longest question. Large traces need filtering before they become state, which is itself a Jev job.
Where to start
Do not refactor your agent. Pick the fork it hits most often, usually the tool-call gate or the next-worker dispatch, and move that one.
Then measure three things for a week:
If all three improve, take the next fork. If the escalation path never fires, your threshold is too loose and you have quietly built an unsupervised system.
Most builders will keep paying frontier prices on every yes, no, route and score, because the calls are invisible individually and only the invoice is visible in aggregate. The few who sort generation from decision from rule will run the same models everyone else runs, an order of magnitude faster and cheaper.
Start with one fork. Measure it. Take the next.
Links
Figures attributed to TypeSafe are vendor-reported. Third-party numbers come from the builders named beside them and have not been independently reproduced here.
















