Agents Don't Need Memory, They Need Documentation
A widely shared essay by Kevin Liao argues that RAG-based agent memory is architecturally flawed and that agents need structured Markdown documentation instead; he shipped the argument as an open-source plugin called Operator Memory.
Every agent memory plugin makes the same architectural bet: extract snippets from transcripts, embed them into a vector database, inject the top similar snippets into each new prompt. Kevin Liao’s essay “Agents Don’t Need Memory. They Need Documentation,” published October 3 and widely shared on Hacker News, argues that the entire category is built on the wrong premise, and shipped the alternative: Operator Memory, a free open-source plugin that replaces embeddings with plain Markdown files the agent reads, updates, and commits.
The Failure Modes Are Specific
The essay’s critique is stronger than a vibe; it enumerates why similarity search is structurally wrong for memory. Retrieved snippets strip away the context and motivations that made a lesson worth storing. Stored memories go stale as codebases change daily, while the vector index keeps confidently resurfacing them. The agent “doesn’t know what it doesn’t know,” so it cannot formulate the query that would retrieve the memory it needs. And the whole store is unauditable: when an agent repeats a mistake, nobody can say which stored memory was wrong or why it was never retrieved. Against the industry premise that “agents forget, that’s the problem,” Liao’s counter is anthropologically sound: humans do not rewatch old meeting recordings either, they write things down.
Documentation Has Different Properties Than Recall
The alternative is not “more context,” it is a different data structure. A structured Markdown brain, instructions, specs, research, and indexes, has properties an embedding store cannot offer: it is auditable (you can read every memory), correctable (fix the file, not a vector), version-controlled (committed and diffed like code), and shareable with a team. It also changes the agentic loop itself, from “prompt, build, forget” to “prompt, consult, build, update,” which makes memory maintenance an explicit agent behavior rather than a background daemon’s guess. Liao reports using the approach for over a year, with no vector databases, no embeddings, and, in his framing, no “token-burning background daemons.”
The Industry Context Makes This Timely
The argument lands in a market that has been sprinting the other way: Claude shipped built-in memory across its agent surfaces, unified memory across chat and cowork, and memory startups raised on the embedding thesis. The essay is a useful counterweight precisely because it does not deny the problem agents have with continuity; it denies that statistical retrieval is the right storage format for procedural knowledge. The distinction is between episodic facts (what happened, retrievable) and operational knowledge (how this codebase works, which needs curation). RAG may be fine for the first; the essay’s case is that only documentation works for the second, and that most of what people are stuffing into memory plugins is the second kind.
What to Watch
Three things. First, adoption of the AGENTS.md-style documentation pattern beyond one plugin: if agent frameworks start shipping documentation-first memory as a default alongside retrieval, the hybrid answer will emerge in practice. Second, whether the major memory vendors respond with auditability features (visible stores, editable entries), which would concede the essay’s core point. Third, team workflows: documentation that lives in the repo, commits itself, and diffs like code makes agent memory a collaboration artifact, and the first platform to make agent-authored documentation a first-class citizen of the code review process will have found the actual product.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
Context Engineering: The Most Important AI Skill in 2026
Context engineering is replacing prompt engineering as the critical AI skill. Learn what it is, why it matters more than prompting, and how to manage state, memory, and information flow in AI systems.
DeepSeek's 242k-Star Agent Harness Reaches the Desktop
DeepSeek Harness v0.2 preview shipped September 29 with official macOS and Windows desktop apps, in-app plugin management without Node or pnpm, and scheduled automations, maturing the MIT-licensed framework that has drawn 242k GitHub stars.
System Prompts: How to Write Effective LLM Instructions
System prompts define how your LLM behaves. Here's how to structure them, what mistakes to avoid, and how provider-specific behavior affects your prompt strategy.
Chain of Thought Prompting: A Developer Guide
Chain of thought prompting makes LLMs reason through problems step by step. Here's when it works, when it doesn't, and how to implement it with practical patterns.
Few-Shot Prompting: How to Guide LLMs with Examples
Few-shot prompting teaches LLMs by example instead of instruction. Here's how to choose examples, format them, and know when few-shot is the right approach vs. fine-tuning.