Ashish · May 2026
A story about getting stuck, and what I found when I stopped trying to get unstuck the obvious way.
Scroll to begin ↓
About six months ago, I was standing at a whiteboard.
There was a graph on it. Entities, relations, arrows — the usual furniture of a system diagram. I had been adding to it for weeks. That afternoon, I couldn't draw the next box.
Not because I didn't know what should go there. Because I didn't believe it would help.
I was trying to give an AI agent a memory.
Not loosely — I mean something specific. A structured representation of what it had learned. A way for it to retain things across conversations, reason over what it knew, not start from scratch every time.
The kind of memory that would make it feel, finally, like something was accumulating.
The architecture was technically fine.
The agent could read back everything in the graph. Ask it a question, it would answer. The structure made sense on paper. Every review went well.
But every time I looked at it, something was missing. Not wrong. Not broken. Lifeless. That was the word I kept not writing down.
The list we kept returning to
We had a shorthand for what current AI systems genuinely can't do. These three kept appearing on every whiteboard, in every serious conversation about what was missing.
I picked memory because it felt tractable.
The other two felt diffuse — learning and search are almost philosophical when you push on them. Memory had the shape of an engineering problem. I could build something, measure it, show it to someone.
In retrospect, that should have been a warning sign.
My first attempt at fixing the lifelessness was to add interrogation.
Instead of the agent looking facts up in the graph, it would ask the graph questions — questions shaped by whatever situation it was in, answers derived by composition. A verb at the front of the data, finally.
I built a prototype. I ran it on a benchmark. The numbers improved. I wrote it up and felt good about it for about a week.
Then I noticed what was still missing.
The questions were only as good as whoever was framing them. When I framed them well, the agent looked intelligent. When I didn't, it was useless. The agent had no taste in questions.
It didn't know which question to ask next, because nothing was at stake for it. Every question was equally worth asking — which meant none of them really were.
The thing I kept polishing
I had built a library. A well-indexed one, with an articulate librarian who could navigate it. Nobody in the building had an errand.
I sat with that for longer than I want to admit.
The fix kept arriving as more machinery. Better question generators. Better retrieval. Sharper summarization. Each version more articulate, same emptiness underneath.
At some point I realized I was decorating the same problem, not solving it. I just didn't know yet what the problem actually was.
The unit is not the fact stored.
It is the prediction made —
and the error noticed
when the prediction fails.
The word I had been refusing to look at was goal.
I had avoided it because it felt too vague — too far from the concrete engineering problem I'd scoped. Goal felt like something for philosophers or cognitive scientists, not for someone trying to ship something.
But it was the right word. I had been building storage for an agent that had no errand. Of course it felt lifeless. It was lifeless.
Memory doesn't have an intrinsic shape.
What you choose to store, what counts as a useful question to ask of it, what counts as a good summary — all of these are downstream of what the system is trying to do.
I had been designing the shelving system before anyone had decided what the library was for.
Once I put the errand back in, the picture rearranged.
You don't need memory for its own sake. You need it because you're trying to anticipate something — and you'd like to be a little less wrong about it tomorrow than you were today.
Everything else — the graph, the embeddings, the retrieval — is bookkeeping in service of that. Not the point. The infrastructure.
Call it what it is
This is older than any of the current excitement about AI. The control theorists have been drawing this loop for fifty years. What's new is how completely the rest of the agenda falls out of taking it seriously.
The loop has to close.
A system that predicts but never acts is rehearsing. One that acts but never observes is just guessing on repeat. The closure — the moment the world talks back — is what makes any of it count as learning.
Open-loop systems can be impressive the way a parrot is impressive. They've absorbed a lot. But nothing is updating. Nothing is at stake.
The old list doesn't disappear. It reorganises.
Memory is what the predictor kept because it reduced future error. Not everything — just what was worth keeping given what it was trying to do.
Learning is what happens when a prediction fails and the model updates. Not accumulation — correction.
Search is the predictor running forward through possibility space, looking for the action most likely to confirm or refute its model of the world.
I don't yet know the right way to encode the world model.
I have three candidates and I'm dissatisfied with all of them. Symbolic representations are interpretable but brittle. Learned embeddings are flexible but opaque. Hybrid approaches carry the downsides of both.
I also don't know how much of the world a useful agent actually needs a model of — versus how much it can offload to the environment by acting and looking. That turns out to be a surprisingly deep question.
So I've stopped calling what I work on a memory system.
More honestly, it's an understanding system. And understanding, for me now, means one thing: can it predict, act on the prediction, notice when it was wrong, and update.
If yes — even crudely, even slowly — something real is happening. If no, we're back in the library with the articulate librarian and no one with an errand.
Write back if any of this strikes you wrong. I would genuinely rather find out from you than from a year of code.
— Ashish