I accidentally turned LLM memory into program analysis
Over the past few months I have been playing around quite a bit with LLM agents, particularly for vulnerability research.
They are becoming surprisingly good at navigating large codebases, explaining unfamiliar subsystems and helping explore potential attack surfaces. However, once an investigation starts taking a few hours, I kept running into the same problem: the model would slowly lose track of what we had actually established.
It might suggest an approach that we had already ruled out, forget that an assumption turned out to be false, or confidently continue reasoning from an observation that was no longer valid. Obviously, telling an LLM that something is wrong does not necessarily mean that it will stop believing all of the things that depended on it :)
I initially started looking into memory systems because I wanted to make LLMs more useful for complex vulnerability research and reduce this type of hallucination.
There are of course already plenty of solutions for giving LLMs memory. Usually this involves storing old conversations or observations somewhere, embedding them, and then retrieving the most relevant pieces whenever the model needs them again.
This works reasonably well, but there was something about it that bothered me.
During a vulnerability research sesh, I don’t just want the model to remember what we said.
I want it to maintain what we currently know.
Imagine that during an investigation we establish the following:
From this, we may conclude that the attacker can control a kernel object.
A normal memory system could store all of these observations and retrieve them again whenever we ask about the exploitability of the bug. The LLM then figures out the same conclusion.
However, suppose that two hours later we discover in LLDB that object_a does not actually point to object_b, and that our previous observation was based on a wrong assumption.
At that point our memory may contain something like:
Now we retrieve some subset of these memories and hope that the LLM correctly figures out which conclusions are still valid.
This started to feel a little familiar to me.
A lot of the work I normally do involves program analysis.
When analysing a program, we usually have a bunch of facts about the program and some rules that derive additional facts from them.
For example, imagine we know:
We could define a rule stating that if one function calls another function, which itself can reach a third function, then the first function can reach the third function as well.
Eventually we calculate a fixed point containing everything we can derive from the program. More importantly, if one of our input facts changes, there are plenty of techniques for updating only the affected results instead of rerunning everything from scratch.
This is also exactly what I wanted from an LLM during vulnerability research.
If an observation changes, I don’t want the model to reconstruct the entire investigation from a transcript and hopefully notice all of the consequences. I want the affected conclusions to become invalid automatically.
When looking at the problem from this perspective, I started wondering why we were making the LLM reconstruct its entire state over and over again.
What if we just maintained it?
And this is how I somehow ended up writing a Datalog engine for LLMs :)
Before we continue, it is probably useful to briefly explain what Datalog actually is.
Datalog is a declarative logic programming language. Instead of writing instructions describing how something should be calculated, we describe facts and rules from which new facts can be derived.
For example, we could store the following facts:
And then define the following rule:
From our existing facts, the engine can therefore derive:
Nothing particularly exciting yet.
However, suppose we later discover that:
If controls_kernel_object(attacker) was derived from that fact, we know exactly which conclusion depends on the observation that just changed, and we can automatically invalidate it.
This is considerably nicer than putting all of the old information into a prompt and asking an LLM to hopefully notice the same thing.
This eventually turned into Lemmalog.
The basic idea is that an LLM should not necessarily be responsible for maintaining its own knowledge. Instead, I split the problem into two parts.
The LLM handles the fuzzy part:
And Lemmalog handles the deterministic part:
This means that the LLM is still responsible for understanding natural language, source code, debugger output and all the other messy information that appears during an investigation.
LLMs happen to be quite good at this.
But once that information has been converted into structured facts, we no longer need the model to repeatedly determine all of its consequences. The database can do that instead.
One of the first interesting problems I ran into was removing facts.
Adding facts to a Datalog database is relatively straightforward: add the new fact and evaluate any rules which may now produce additional results.
Removing something is a little more annoying.
Take the following example:
Here c has two separate reasons for being true.
If we remove a, we cannot simply remove c, because b still provides another derivation for it. However, if we remove both a and b, c should disappear as well.
This turns out to be quite important during vulnerability research, because a conclusion may be supported by multiple observations.
may remain true even if one particular exploit primitive turns out not to work, because there is another independent path to the same result.
So Lemmalog has to keep track of how facts were derived and update their support when something changes.
Conveniently, this also gives us another useful property:
we can ask why something is true.
Imagine we have been running an agent for a few hours while investigating something and it eventually concludes:
That is nice, but I would also quite like to know why.
Because Lemmalog already tracks the dependencies of derived facts, we can ask it for the provenance of a conclusion. For example, we may get something that conceptually looks like this:
If observation_41 later turns out to be incorrect, we know that this conclusion may no longer be valid, and because the database knows this as well, it can remove the affected conclusions automatically.
This was originally mostly necessary to make incremental evaluation work correctly, but it turns out that being able to ask an AI agent why it believes something is quite useful as well :)
It also addresses one of the more annoying failure modes I encountered with LLM-assisted research. Sometimes a model will confidently say something like:
when that is not actually true.
If a conclusion exists in Lemmalog, I can ask where it came from. If there is no provenance supporting it, then it is not part of the maintained state.
This obviously does not prevent an LLM from hallucinating during extraction, but it does make it much harder for unsupported conclusions to silently become part of the investigation.
Another issue is that replacing old facts is not always the same as deleting them.
Suppose we originally believe:
For most current queries, we probably only care about the second statement. However, if we want to understand why we previously explored a particular exploit strategy, the old state is still useful.
For this reason Lemmalog can associate facts with validity intervals.
Conceptually, we can represent the state as something like:
This allows us to answer both:
without keeping two apparently contradictory facts around and asking the LLM to decide which one we meant.
Again, this is not really a language model problem.
It is mostly a database problem.
Vector databases are very useful.
semantic search is probably exactly what I want.
But cosine vibe similarity and truth are not quite the same thing.
A vector database can retrieve:
because it is relevant to my question. It does not inherently know that the statement was disproven two hours later, or that five other conclusions depended on it and should therefore no longer be considered valid.
This made me realise that there are really two different problems hiding under the term “memory”.
Retrieval is very good at the first problem.
Lemmalog is mostly an experiment in solving the second one.
The two can also be combined, which is what I currently do.
The more I worked on this, the more similarities with program analysis started appearing.
During a vulnerability investigation we have observations:
This maps surprisingly well to the things we already do in program analysis.
and when an input changes, we perform incremental evaluation:
Because we track dependencies, we can also explain where results came from:
At some point it became fairly obvious that I had approached the problem like a static analysis engine without intentionally meaning to.
This also changed how I thought about the role of the LLM itself.
You can almost think of the whole system as a slightly strange compiler.
The LLM acts as the front-end:
Lemmalog is the intermediate representation and analysis engine:
Another LLM invocation can eventually turn that state back into natural language, suggest the next experiment, or use it to perform some action.
The amusing part is that our parser is probabilistic, while everything after it does not necessarily have to be.
This is of course the important question.
The engine itself now supports incremental evaluation, retractions, provenance, temporal facts, aggregations, entity reconciliation, hybrid retrieval, demand-driven queries and a bunch of other things that I probably added because implementing Datalog features is more fun than I expected.
There is also an MCP server which allows agents to use Lemmalog directly.
But none of that matters very much if giving an LLM this memory does not actually improve anything.
So I plugged it into MemEval and tested it on both LongMemEval and LoCoMo using their standardized reader models and evaluation setup. Extraction during ingestion is Claude Sonnet 4.6 (chunked and file-cached, so it is paid once per conversation); everything after extraction uses the benchmark’s own standardized readers and judges.
The results were a little better than I expected.
LongMemEval tests whether an LLM can answer questions about information spread across long conversation histories. The split I used contains 102 questions, divided equally between user facts, assistant facts, preferences, multi-session questions, temporal reasoning and knowledge updates.
Because 17 questions per category is not exactly a massive sample size, I ran Lemmalog three times rather than getting excited about whichever run happened to score highest.
For comparison, the published memory-system results are:
My own full-context GPT-4.1 run scored 0.197 F1.
So Lemmalog is not beating PropMem yet, and it is still slightly behind SimpleMem, but it gets more than twice the F1 of giving GPT-4.1 the entire conversation.
More amusingly, the context passed to the answering model is roughly 38 times smaller.
Apparently maintaining state instead of repeatedly rereading the entire history is useful :)
The category results from one representative run looked like this:
The result I found most interesting was Knowledge Update.
Lemmalog scored 0.579, compared with 0.528 for PropMem and 0.202 for full context.
Knowledge Update is basically the situation I originally cared about:
So seeing Lemmalog top the published field on the category that most closely resembles maintained program state was rather satisfying.
Single-session factual memory also worked surprisingly well. Lemmalog reached 0.790 on user facts and 0.672 on assistant facts, while temporal reasoning reached 0.416, almost identical to PropMem’s 0.424 in that run.
The obvious remaining problem is multi-session reasoning:
Diagnosing those failures was interesting: the information usually was not mis-connected, it was simply never extracted. If the extractor never emits a fact for the Airbnb booking, no amount of derivation is going to answer a question about it.
Which brings us to one of the more amusing parts of running benchmarks.
At one point LongMemEval suddenly dropped to 0.371 F1.
After going through the failures, I discovered that 32 of the 102 questions were being refused.
All 32 were answerable.
The problem was an instruction I had added to reduce hallucinations. I told the reader to make sure that the answer was actually supported by the retrieved facts before answering.
Unfortunately, the model interpreted this as:
If no single fact literally contains the final answer, refuse.
There is obviously no fact saying:
if the memory instead contains:
The answer exists. It just requires counting.
The fix was to separate two cases:
If the premise is absent or misattributed, refuse.
If the evidence exists but requires co