Jev-Mem: Memory Decisions Without an LLM

Jev-Mem shows that memory decisions don't need an LLM that writes sentences. Its thresholds were tested with one model on one benchmark.
Every agent memory system makes a ton of small decisions, all the time.
Most setups hand all of that to a general-purpose LLM. The model writes out an answer token by token, and then some code parses it.
For a yes/no question, that's a lot of work!!
A new paper from UT Dallas, Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents, basically says: stop doing that for most of these decisions.
Here's the rundown
What is Jev?
Jev is a model from TypeSafe AI that answers structured questions without writing any text. You hand it some state, like a query and a memory, a batch of questions, and it hands back numbers. No prose.
The paper uses two question types:
Since every question has a small, known output space, there's nothing to parse. And questions that share the same state go out in one batched call.
The paper borrows the System One and System Two split as an analogy. System One is a lightweight controller that makes the small memory decisions, and Jev is the implementation the authors use for it. System Two is the LLM that writes the final answer. The cost is all the little memory decisions that happen before the answer ever gets written.
How Jev-Mem uses it for memory
The key observation in the paper is pretty simple. A lot of memory operations are semantic but not generative, which means they need judgment, but the output is a label or a score, not a paragraph.
Write path (when a memory comes in):
Read path (when a query comes in):
Where does Jev-Mem sit?
Every memory system does three things: write memories, retrieve them, and hand them to an LLM to answer. Jev-Mem leaves the answer to the LLM and moves the decisions inside the other two steps to Jev.
Jev-Mem is a full architecture with its own multi-relational memory store, not a drop-in controller.
Two strong ideas
Report
On LoCoMo, with GPT-4o-mini as the answering model:
Nice numbers. But a few things before you run with them:
The real catch: thresholds
The whole design runs on cutoffs:
And the appendix is pretty candid about this: the returned values are not assumed to be calibrated probabilities.
That's the whole thing, honestly.
A 0.95 cutoff only means something if 0.95 means the same thing every time. These thresholds were evaluated using only one model on a single benchmark. Swap the controller model, or move to data that looks nothing like LoCoMo, and those same cutoffs can stop retrieval too early or let noise in.
There are caps on the loop (depth 8, 60 visited nodes, 16 controller calls, 15 seconds), so it won't run forever. But it can still make the wrong call, and since the LLM only sees what the controller hands it, that stopping threshold is exactly where speed and correctness trade off.
So our takeaway isn't "swap your LLM for Jev." It's: a fast classifier can absolutely make memory decisions, but calibrate its scores on your own data before you trust its thresholds.
Where this could fit if you already use a memory layer
You don't have to rebuild your memory setup to try this. If you use Mem0, search already gives you ranked memories with a relevance score and timestamps. On Platform, that score is a combined 0 to 1 value from semantic, BM25 keyword, and entity signals.
So you could drop a typed classifier between that search and your model, and have it check whether each memory actually matters for the current query before the model ever sees it:
from mem0 import MemoryClient
client = MemoryClient(api_key="your-api-key")
THRESHOLD = 0.5 # placeholder: tune on your own labeled queries
def relevance(query: str, memory: str) -> float:
"""Swap in Jev or any classifier that returns a score from 0 to 1."""
raise NotImplementedError
query = "What should I cook tonight?"
results = client.search(query, filters={"user_id": "user123"})
memories = [
r["memory"] for r in results.get("results", [])
if relevance(query, r["memory"]) >= THRESHOLD
]
# Pass `memories` to your LLM as context.To be clear, this is a pattern you'd build and test yourself, not an existing integration.
Mem0's automatic path is additive too, so an old fact and its update can both come back from search. That's why the timestamps matter: they let your model see which version is newer. If your app knows a fact changed, you can also correct it directly with update or delete.
Summing up
Jev-Mem makes a really good case that frequent memory decisions don't need a model that writes sentences. Typed, batched, bounded decisions are a clean way to take cost off the memory path.
What's still open is how much the controller alone contributes, how it holds up under real production load, and whether those thresholds hold outside the setup they were tested in.
Fast memory decisions are the easy part. Calibrating them is the work.
If you're poking at memory control in your own stack, we'd love to hear what you're seeing.
References
Paper and code
Mem0



