September 29, 20265 min read

Building Retold: Agent Memory That Decides Before It Retrieves


A memory system should stay quiet on the turns that need nothing, and surface the right records on the turns where they matter.

TL;DR

An earlier essay of mine, Related Is Not Useful, measured five agent memory systems at their documented defaults. Two results held across all of them. Every system returned its full quota of records on nearly every turn, including turns the store could not answer at all. And the further a saved rule sat from the wording of the request, the fewer systems brought it back, until the most distant rule tested came back from none of them.

That essay raised a question none of the five systems answered.

I am building Retold in the open as an answer to it. It is early and public under MIT, and I would like you to use it, break it, and help build it.

Instead of ranking records by one similarity score, Retold makes two decisions. A planner asks whether the turn needs anything stored. Search runs only if it says yes. A judge then compares what came back against the answer already drafted, and keeps only the records that would change it. Separately, preferences that apply to every answer skip that path entirely and sit in a small profile from the start of the session.

I am still designing the experiments to measure it, but early results are promising on one half of the problem: Retold stays quiet on turns that need nothing from the store. The other half is not there yet. When a turn does need a stored record, Retold does not fetch it as reliably as it should, and that is where I could use help.

GitHubadimyth/retoldA local, provider-neutral memory layer for AI agents. It decides whether a turn needs memory before it searches, and uses a record only when that record would change the answer.

Architecture

How does a memory system decide whether a turn needs anything stored?

Put a model in front of the search. Ask it whether the turn needs anything the store might hold, and search only when it says yes. Then, before any record reaches the answer, ask a second model whether that record would change what the assistant was about to say.

  • Serving model. The host writes a draft with no retrieved records in it.
  • Planner. It decides whether this turn needs anything stored. It never sees the text of the records it might fetch, only the categories of thing this user has stored, such as "time zone and working hours". It can tell that a question about staging versions might be answerable without learning that Node is pinned to 20.
  • Retrieval. Only if the planner says yes does Retold search, and only this user's records.
  • Judge. It reads the draft and every candidate record together, and keeps the ones that would change the answer.
  • Second generation. If anything survives the judge, the answer is written once more with those records in hand.
  • Fallback. A no from either the planner or the judge serves the original draft untouched.

The path builds on two papers that also reject retrieving by similarity and injecting whatever comes back. RUMS picks the stored memories that most reduce the model's uncertainty about its next response. TRACE-Memory works out what user-specific information a question is missing, retrieves for those gaps, and admits a record only if it makes the answer better than one written from public information alone. Retold keeps TRACE-Memory's two stages, a planner and then a judge. It adds two things neither paper has: the planner also sees which categories of information this user has stored, without reading any of them, and preferences about how to answer get their own path.

That path is built for facts: a time zone, an owner, a pinned version. Something the turn is reaching for, even when it does not name it.

It is the wrong path for a preference about how to answer. Nobody ever asks "how should you answer me?", so a check that runs once per turn rarely says yes for the kind of memory that applies on the most turns. Retold therefore treats records as two kinds:

  • A conditional record is a fact, a decision, or an event. It travels the path in the diagram and earns its place on the turns that need it.
  • An ambient record is a preference about how to answer. It skips the path. The host assembles a small, capped profile of these once at the start of the session, and it is in the draft before the planner has finished thinking.

Every record starts as conditional. It moves into the ambient profile only when all three of these are true:

  • The user said it. Retold finds the preference in the user's own words in the conversation. A paraphrase or a guess by the model does not count.
  • It applies everywhere. "Use Python for code examples unless I ask for another language" qualifies. "When reviewing pull requests, keep answers short" covers one kind of task, and "keep answers short for now" expires, so both stay conditional.
  • It is about how to answer. It sets the answer style, the response language, an accessibility need, or the language for code examples.

If any of the three fails, the record stays conditional. A fixed rule checks the wording first, looking for phrases like "unless I ask", "when reviewing" or "for now". When it does not recognise the wording, a model judges it, and an unsure call sends the record to a review queue, where it stays conditional until a person decides. A detailed post on that machinery follows soon.

The two kinds cover each other's blind spots. The planner only fires when a turn is reaching for something, and no turn reaches for a preference about how to answer, so a preference left conditional would rarely arrive. The profile goes into every answer, so a fact placed there would come back on every turn, which is the failure Retold exists to avoid. Preferences sit in the profile because they apply everywhere, and facts wait behind the planner and the judge because they apply only sometimes. Together, every answer starts with the user's preferences and picks up facts only on the turns that need them.

Where it stands

I am designing the experiments for this now. They will measure both halves of the problem: whether Retold stays quiet when a turn needs nothing, and whether it fetches the right record when a turn needs one. I will share the numbers once they are done.

Try it, break it, help build it

Two commands, and no API key. The first run downloads about 640 MB of local models.

bash
python -m pip install "retold[local-models] @ https://github.com/adimyth/retold/releases/download/v1.2.1/retold-1.2.1-py3-none-any.whl"
python
from retold import Retold

with Retold.open("memory.sqlite", profile="lite") as retold:
    with retold.session(user_id="aditya") as memory:
        memory.remember(
            "I prefer concise answers.",
            evidence="I prefer concise answers.",
            attribute="answer_style",
        )
        print(memory.search("What kind of answers do I prefer?").text)

The result carries the reason it came back, not just the text:

text
Recalled 1 memory for "What kind of answers do I prefer?".

[01a08...] semantic · confirmed · user_statement · event 2026-09-29 · scope agent:assistant/aditya
I prefer concise answers.
matched: dense 0.53 (rank 1), lexical 2/3 (rank 1); fused rank 1; passed dense 0.53 ≥ 0.32 (semantic)

That snippet is plain search. The decision path is off by default, and examples/utility_aware_quickstart.py turns it on with deterministic policies and real retrieval, so you can watch it choose to use memory on one turn and decline on the next without spending anything on models.

I am building this in public, and there is more open work than I can do alone. Three places where help would matter most:

  • Fetching memory when a turn needs it. This is the weak half right now. A turn where the planner should have asked for a record and stayed quiet is the most useful thing you can send me.
  • Experiments and benchmarks. I am designing the experiments that measure both halves of the problem, and I want to evaluate Retold against public memory benchmarks such as LongMemEval and LoCoMo. Help with either is welcome.
  • Storage. Retold is one SQLite file and one process. A Postgres option is the next storage work.

Cases that break it are just as welcome: a record the judge admitted that made your answer worse, or a preference the promotion rule got wrong. Open an issue or a pull request on the repository.