Jev-Mem: Memory Decisions Without an LLM

@mem0ai
mem0@mem0ai
25 views Oct 08, 2026 ~8 min read
Advertisement

Jev-Mem shows that memory decisions don't need an LLM that writes sentences. Its thresholds were tested with one model on one benchmark.

Media image

Every agent memory system makes a ton of small decisions, all the time.

  • Is this new message a preference or an event?
  • Is it related to something already stored?
  • Which part of memory should this query even search?
  • Is there enough evidence to stop looking? … and many more.
  • Most setups hand all of that to a general-purpose LLM. The model writes out an answer token by token, and then some code parses it.

    For a yes/no question, that's a lot of work!!

    A new paper from UT Dallas, Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents, basically says: stop doing that for most of these decisions.

    Here's the rundown


    What is Jev?

    Jev is a model from TypeSafe AI that answers structured questions without writing any text. You hand it some state, like a query and a memory, a batch of questions, and it hands back numbers. No prose.

    The paper uses two question types:

  • Noul: The paper's term for a yes/no question. You give it an instruction and clear criteria for true and false, and Jev returns a value between 0 and 1.
  • Choice: Pick one from a fixed set of options. You get back the chosen label and a distribution over all the options.
  • Since every question has a small, known output space, there's nothing to parse. And questions that share the same state go out in one batched call.

    The paper borrows the System One and System Two split as an analogy. System One is a lightweight controller that makes the small memory decisions, and Jev is the implementation the authors use for it. System Two is the LLM that writes the final answer. The cost is all the little memory decisions that happen before the answer ever gets written.


    How Jev-Mem uses it for memory

    The key observation in the paper is pretty simple. A lot of memory operations are semantic but not generative, which means they need judgment, but the output is a label or a score, not a paragraph.

    Media image

    Write path (when a memory comes in):

  • Every observation becomes one memory node.
  • Jev gives it four overlapping type scores: episodic, semantic, procedural, preference. So one memory can be more than one type.
  • Plain vector search, keyword overlap, shared entities, and timestamps pick at most 10 existing memories as candidates.
  • Jev only judges those pairs. Are they related? Did one cause the other? Same episode?
  • An edge gets created when the score hits 0.60.
  • Anything that doesn't need judgment stays in plain code. Timestamps set the order, and exact shared IDs create entity links.
  • Read path (when a query comes in):

  • Jev scores which kinds of relations the query needs (semantic, temporal, causal, entity), and how much it leans on multi-hop reasoning and recency.
  • Search effort gets split across those relation types based on the scores.
  • Then retrieval runs as a loop: find entry points, check if the evidence is enough, expand to neighbors, score the new candidates, and check again.
  • Only when that loop stops do the selected memories go to the LLM.

  • Where does Jev-Mem sit?

    Every memory system does three things: write memories, retrieve them, and hand them to an LLM to answer. Jev-Mem leaves the answer to the LLM and moves the decisions inside the other two steps to Jev.

  • On write, Jev types the new memory, plain search picks up to 10 related candidates, and Jev decides which links to draw (semantic, temporal, causal, entity).
  • On retrieval, Jev routes the query to the right link types, sets a search budget, scores what it finds, and decides when to stop. Only then does the LLM see anything.
  • Jev-Mem is a full architecture with its own multi-relational memory store, not a drop-in controller.

    Two strong ideas

  • Keep everything, pick later: Their config turns admission filtering off, so every valid, nonempty observation gets stored. The picky part happens when relations are built and when memory is read. The reasoning makes sense because something that looks useless today might be exactly what a future query needs, and a write-time filter would lose it for good.The catch here is corrections. If a user updates a fact, both versions stay in memory, and the paper doesn't test whether the updated one actually wins at answer time.
  • Consolidation that doesn't delete. Every 20 writes, Jev scores pairs of related memories for redundancy, contradiction, obsolescence, and whether a link would help. It then picks one of four options: keep separate, merge, promote, or uncertain. By default, it only records links, and the originals always stay. A merged or promoted memory is written only if you plug in an LLM summarizer, and only when Jev picks merge or promote with a score of at least 0.85 and the contradiction score stays below 0.85. Even then, the new memory sits next to the originals rather than replacing them.

  • Report

    On LoCoMo, with GPT-4o-mini as the answering model:

    Media image
  • Answer quality: An LLM-as-a-Judge score of 0.777 vs 0.700 for the strongest baseline, an 11.0% relative gain.
  • Build time: 158 seconds to construct memory vs 1,044 seconds for the fastest competing system, so about 6.6× faster.
  • Query latency: 0.93 seconds on average vs 1.47 seconds for the fastest memory-based baseline, 36.7% lower. This includes retrieval and answer generation.
  • Nice numbers. But a few things before you run with them:

  • One benchmark, one setup: Only LoCoMo, one answering model, and an LLM judge. Compare these only with numbers from the same setup.
  • The 6.6× is construction time, not query speed: The paper credits typed, batched decisions, but the implementation also caches, so it's hard to say how much each one contributes.
  • Latency is an average: There are no tail-latency numbers and no tests under concurrent writes, and that's what production actually cares about.
  • The controller's contribution is only partly isolated: The strongest baseline, MAGMA, is from the same group, and the repo says running without Jev gives you the MAGMA baseline. But the paper doesn't report write-only or read-only ablations. Also, the prose and Table 1 disagree on a few category scores, so trust the table.

  • The real catch: thresholds

    The whole design runs on cutoffs:

  • An edge is created at 0.60.
  • A relation type gets activated for a query at 0.10.
  • Retrieval stops when evidence sufficiency is at least 0.95and missing-evidence and contradiction scores are both under 0.15, or when the "keep going" score drops under 0.15.
  • And the appendix is pretty candid about this: the returned values are not assumed to be calibrated probabilities.

    Media image

    That's the whole thing, honestly.

    A 0.95 cutoff only means something if 0.95 means the same thing every time. These thresholds were evaluated using only one model on a single benchmark. Swap the controller model, or move to data that looks nothing like LoCoMo, and those same cutoffs can stop retrieval too early or let noise in.

    There are caps on the loop (depth 8, 60 visited nodes, 16 controller calls, 15 seconds), so it won't run forever. But it can still make the wrong call, and since the LLM only sees what the controller hands it, that stopping threshold is exactly where speed and correctness trade off.

    So our takeaway isn't "swap your LLM for Jev." It's: a fast classifier can absolutely make memory decisions, but calibrate its scores on your own data before you trust its thresholds.


    Where this could fit if you already use a memory layer

    You don't have to rebuild your memory setup to try this. If you use Mem0, search already gives you ranked memories with a relevance score and timestamps. On Platform, that score is a combined 0 to 1 value from semantic, BM25 keyword, and entity signals.

    So you could drop a typed classifier between that search and your model, and have it check whether each memory actually matters for the current query before the model ever sees it:

    from mem0 import MemoryClient
    
    client = MemoryClient(api_key="your-api-key")
    THRESHOLD = 0.5  # placeholder: tune on your own labeled queries
    
    def relevance(query: str, memory: str) -> float:
        """Swap in Jev or any classifier that returns a score from 0 to 1."""
        raise NotImplementedError
    
    query = "What should I cook tonight?"
    results = client.search(query, filters={"user_id": "user123"})
    memories = [
        r["memory"] for r in results.get("results", [])
        if relevance(query, r["memory"]) >= THRESHOLD
    ]
    # Pass `memories` to your LLM as context.

    To be clear, this is a pattern you'd build and test yourself, not an existing integration.

    Mem0's automatic path is additive too, so an old fact and its update can both come back from search. That's why the timestamps matter: they let your model see which version is newer. If your app knows a fact changed, you can also correct it directly with update or delete.


    Summing up

    Jev-Mem makes a really good case that frequent memory decisions don't need a model that writes sentences. Typed, batched, bounded decisions are a clean way to take cost off the memory path.

    What's still open is how much the controller alone contributes, how it holds up under real production load, and whether those thresholds hold outside the setup they were tested in.

    Fast memory decisions are the easy part. Calibrating them is the work.

    If you're poking at memory control in your own stack, we'd love to hear what you're seeing.


    References

    Paper and code

  • Jiang, Li, and Li, Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents, arXiv:2609.23986 (v1, 21 Sep 2026)
  • Jev-Mem repo: github.com/libingzheren/Jev-Mem
  • Mem0

  • Mem0, Quickstart
  • Mem0, Search memories
  • Mem0, How Mem0 works
  • Actions
    What You Can Do
    • Export as PDF or Markdown
    • Batch Export to Notion
    • Bookmark & Highlight
    • Screenshot Tweet
    Create Free Account

    Includes 7-day Premium trial

    Advertisement