Dev

Agent memory as Markdown documents: a readable memory is still a memory

Moving agent memory out of a vector store and into Markdown files makes it easy to inspect, and that is worth having. Staleness, conflicting writers and the page nobody opened were never problems of storage format, so they move into the folder intact, where a clean file can pass for a checked one.

Kevin Liao published an essay on 3 October with a title that does most of the arguing: agents don't need memory, they need documentation. His target is the memory plugin as it usually ships, a pipeline that cuts session transcripts into snippets, embeds them, and attaches the closest few to every prompt. His list of what goes wrong with that is short and mostly right. The line I would keep from it is this one:

Similarity search ranks how close two snippets are in embedding space. That’s it. You don’t know which is correct, current, or what’s missing.

His replacement is a workspace of plain Markdown that the agent reads before it works and revises after, a loop that goes from build-and-forget to consult, build, update. He shipped it as Operator Memory, and it is more careful than the essay has room to show: a fixed orientation loaded into every session, catalogs that say when each document is worth opening, private and shared partitions. Claude Code's own auto memory has the same basic shape, typed notes in Markdown files under a MEMORY.md index, all of it editable by hand.

I agree with the direction. I maintain a memory server for agents, agent-recall, and there is no vector database in it either: it is a SQLite file with full-text search over entities and the facts attached to them. But the essay credits the format and the loop with more than they deliver. Coherent documents do answer two of its five failures, the lost context and the ranking by similarity, and readable files answer a third, the audit. What they do not give you is any guarantee that a page is current, that the agent opened the right one, or that two writers agreed. Those were never about how the bytes are stored, so they move into the folder along with everything else, and they are harder to see there, because a file you can open looks like a file somebody has checked. What a Markdown folder gives you is a very good view. Whether it also keeps a sound record of how things changed is a separate question.

What does a Markdown agent memory actually fix?

The fifth failure on Liao's list is the one Markdown fixes.

The store is unauditable.

A vector store is not literally sealed; most keep the original text next to each embedding and let you list and delete records. But nobody reads ten thousand payloads through a database client, and a folder of Markdown you can read on a Sunday afternoon, correct by hand and commit. That difference matters more than it sounds. An entry nobody reads is an entry nobody corrects, and a memory nobody corrects only accumulates.

Look at what kind of fix it is, though. Auditability is something the store offers a reader. It does nothing until someone reads, and the reader it assumes is a person with time set aside. Operator's own workflow guide is candid about the gap: agents prefer the status quo, it says, and hesitate to consolidate documents, split one that has grown too large, or trim one that has gone stale, so those moves need your judgment. The format makes the audit possible. Doing it is still on you.

Why is the memory map still a retrieval layer?

The fourth failure is the one the essay most needs documentation to solve.

Agents can’t search for what they don’t know.

Both systems answer it the same way, with a map. Claude Code loads the first 200 lines or 25KB of MEMORY.md at the start of a session, whichever comes first, and opens the topic files only when Claude decides it needs them. Operator puts its catalogs and main index into every session and gives each subindex a read_if condition, a sentence saying when it is worth opening. This beats top-k similarity, and not by a little. The agent sees what exists before it forms a query, and the routing rules are sentences you can read and fix.

It is still retrieval. The agent decides which page to open by judging how well a one-line description fits the task in front of it, and that judgment is made by the model instead of by cosine distance. The unknown-unknowns problem survives the move in a narrower form: if the description does not look relevant to the task as the agent reads it, the page stays closed, and nothing records that it stayed closed. The map also has a size. Past Claude Code's 200 lines, an entry is no longer announced at startup, and from then on a fact is found only by an agent that already thought to go looking for it.

Is stale agent memory a format problem or a time problem?

The third failure is the one the format can date but not cure.

The past is treated as truth.

A Markdown page that says the auth service uses one token format is exactly as stale as an embedded snippet saying the same thing, the morning after the format changes. Readability does not make a sentence current. Liao's answer is the update step, the agent revising what is outdated while the whole picture is still in its context, and that works for whatever a session touched. For a fact that changed with no session in the room, a teammate's commit or a decision taken in a meeting, it guarantees nothing: a later session can repair the page only if something sends it looking. Operator's architecture says as much, plainly, and its workflow asks you to have the index refreshed after a large pull. The Brain can still go stale after a git pull, and the design's answer is that the stale page is at least an ordinary, inspectable file. I think that is the right answer and an incomplete one. A page you could inspect and a page somebody did inspect are different states, and the folder looks the same in both.

Dates help, and they are narrower than they look. Since v2.1.214, Claude Code writes a modified timestamp into a memory file's frontmatter each time Claude writes a file that has frontmatter. That says when the file last changed, not when any sentence in it was last true. agent-recall goes one step further on the record side. When a fact's value changes, the old value is not overwritten: it is closed with the time it was replaced and kept, with an optional label for its source, so you can ask what the store believed about that fact on any given day. That holds for those key-value facts, not for everything in the store. I built it that way because I needed to know what was true when, not only what is true now. The honest limit is the same as the timestamp's. Both clocks record when the store learned something, not when the world changed. A clock on the write is still worth a great deal, because it is what turns a stale fact from invisible into datable.

The quieter cost of a document is the overwrite. Editing a sentence is the natural thing to do with a readable page, and on a page without version control it removes the evidence an audit would need. Operator's shared partition lives in Git, so its history survives. Its private partition, by design, is added to the global Git ignore. Whether a document memory keeps its past depends on where each page happens to live, which is a storage policy, and the format does not set it for you.

Who is allowed to change an agent's memory page?

The problem I would add is not on Liao's list at all, because it only appears once more than one agent writes. agent-recall checks scope on writes because two agents once wrote conflicting data to the same entity. That check is narrower than it sounds: it stops an agent writing outside its scope, while two agents that are both allowed to set the same fact still race, and the later write wins, with the earlier value kept in the history. Free-text notes are not even that strict: two contradictory ones simply sit side by side. A shared folder of Markdown has the same problem in a quieter form. The conflict is two paragraphs that disagree in one file, or an edit that silently replaces another. Operator settles which instruction partition outranks which. Which agent may rewrite which page, nothing settles.

There is also the question of who the writer is. The update step is the agent writing an account of its own work, which comes back into a later window with the standing of something you wrote, and I have argued that at length already. A readable memory makes the account easy to check. It does not change who wrote it.

Why is the read in the harness but the memory write still a sentence?

Liao is pointed about what Operator leaves out: no summarizers, no curators, no overnight rewriters, no background process burning tokens. agent-recall has one of those, a language model that writes a briefing from the store, and it exists because agents could not make sense of a raw dump of hundreds of entries. A folder of documents mostly avoids that by loading selectively, but the individual pages and the maps still grow, and a page too long to read whole has the same problem in miniature. Someone keeps them short. You can do that by hand, the agent can do it as it updates, or a summarizer can do it at read time. Operator puts the work on the agent and on your direction, which is a fair call, and the tokens get spent all the same, just inside the session instead of overnight.

On getting memory into the window, the two designs agree, and I think both are right. Operator injects its orientation into every model call in its OpenCode integration. agent-recall delivers its briefing through a SessionStart hook in Claude Code. In both, the loading happens without asking the model, which matters, because a prompt is not an invariant. The writing does not work that way in either. agent-recall's server carries instructions telling the agent to save what matters as it goes, and those instructions exist because agents ignored the memory tools until they were told to use them. Being told is still advisory. Operator's update step is an instruction as well. The read side has moved into the harness and the write side is still a request, and the write side is the one the whole design depends on.

What should agent memory look like: a record with a clock and a page on top?

So the line I would draw is not between memory and documentation. It is between what you read and what keeps the history of it. The record needs what the essay's list was really pointing at: values that keep their past instead of being overwritten, a rule about who may change what, a way into the window that the harness enforces, and some signal from outside the sessions when the world moves. Versioned Markdown can be that record, if every page lives under version control and someone owns the trigger that checks it against the world. A database with a Markdown view on top is another way to get there. agent-recall's exporter writes current fact values to plain files and leaves their revision history in the database, which is the split I mean. What does not work is the readable page standing in for all of it, unversioned and maintained only when a session happens to pass by. That setup inherits every overwrite, every unopened page and every confident stale sentence the essay blamed on vectors.

A readable memory is a real improvement. It is still a memory, with every problem a memory has.

Discussion

No comment section here — all discussions happen on X.

Max Nardit

Max Nardit

@mnardit

More articles