JEV actually doesn't even run deterministically Exact same prompts...

@neural_avb
AVB@neural_avb
118 views Sep 21, 2026 ~1 min read
Advertisement
1
JEV actually doesn't even run deterministically

Exact same prompts give you different probabilities when you run it multiple times

The ORDER of the choices DRASTICALLY changes the output probs

I am more and more confused by the their "no-hallucination" claim
Media image
@neural_avb
AVB@neural_avb
Guys it’s a misconception that Jev can’t hallucinate

It’s a non-generative LM ie it models probabilities of things given language as context

The guarantee is the type-safety, and that answers are deterministic - ie no stochastic sampling. The deterministic answers can still be confidently wrong!

It can still hallucinate if the training data on the topic was under cooked
2
Easy to reproduce:

curl -sS api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Your annual plan looks interesting, but I cannot find whether it supports five users. The pricing page will not load for me either.",
"model": "jev-1.13.0",
"questions": {
"classification": {
"type": "choice",
"instructions": "What is the main intent of this message?",
"criteria": {
"bug_report": "Reporting a feature or page that does not work",
"information": "Asking about product capabilities or pricing",
"purchase": "Considering buying a plan"
}
}
}
}
EOF

Switch the order of the choices, make bug_report appear below information... and model's pick changes
3
To be clear, it makes total sense in a network architecture sense. JEV is a language model. Token representations depend on the input state, and TypeSafe team is likely packing the full list of choices as a list of options into the prompt... So choice order changes probabilities.
4
My issue really is them claiming "no-hallucination" coz of the schema safety.

Even Autoregressive LLMs can guarantee schema using low-level guidance libraries like outlines.

Since JEV is a classification model, or a parallel constraiend decoder - the schema guarantee is largely a inferencing algorithm. Its NOT the capability of the underlying model, but the harness around it.
5
While we are on the topic, the "confidence" stuff you have seen people post on JEV is also a farce. The "confidence" value the API returns is a derived metric from the predicted class probs... it does not "model" uncertainty like a bayesian net would.


@neural_avb
AVB@neural_avb
I was super intrigued when I noticed that JEV returns not just probs for candidate choices, but also a CONFIDENCE SCORE of it's prediction...

I was wondering if TypeSafe AI brought back Bayesian Networks or something to learn uncertainty.

Turns out - NO. Acc to their docs, confidence is just a derived metric from the classification scores. It carries ZERO extra information. Their demo page shows:

confidence = (3 × largest probability − 1) / 2

Confidence is tightly coupled with the classification scores it is measuring the uncertainty for.

There is no magic here. The model can produce high probability on a wrong choice and just be confidently wrong.
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial

Advertisement