JEV actually doesn't even run deterministically Exact same prompts...

Exact same prompts give you different probabilities when you run it multiple times
The ORDER of the choices DRASTICALLY changes the output probs
I am more and more confused by the their "no-hallucination" claim

It’s a non-generative LM ie it models probabilities of things given language as context
The guarantee is the type-safety, and that answers are deterministic - ie no stochastic sampling. The deterministic answers can still be confidently wrong!
It can still hallucinate if the training data on the topic was under cooked
curl -sS api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Your annual plan looks interesting, but I cannot find whether it supports five users. The pricing page will not load for me either.",
"model": "jev-1.13.0",
"questions": {
"classification": {
"type": "choice",
"instructions": "What is the main intent of this message?",
"criteria": {
"bug_report": "Reporting a feature or page that does not work",
"information": "Asking about product capabilities or pricing",
"purchase": "Considering buying a plan"
}
}
}
}
EOF
Switch the order of the choices, make bug_report appear below information... and model's pick changes
Even Autoregressive LLMs can guarantee schema using low-level guidance libraries like outlines.
Since JEV is a classification model, or a parallel constraiend decoder - the schema guarantee is largely a inferencing algorithm. Its NOT the capability of the underlying model, but the harness around it.

I was wondering if TypeSafe AI brought back Bayesian Networks or something to learn uncertainty.
Turns out - NO. Acc to their docs, confidence is just a derived metric from the classification scores. It carries ZERO extra information. Their demo page shows:
confidence = (3 × largest probability − 1) / 2
Confidence is tightly coupled with the classification scores it is measuring the uncertainty for.
There is no magic here. The model can produce high probability on a wrong choice and just be confidently wrong.
