Why the last mile of AI memory repair is not cleanup. It is judgment.
At the end of the first repair, the audit landed on seven.
Twenty-six memory stores. One hundred eighty-five active files. No orphaned active memories. No broken index entries.
Seven unresolved wikilinks.
The tempting move was obvious.
Create seven missing files. Change seven links. Get the green check. Go home.
Clean.
Also potentially wrong seven different ways.
Because a broken memory link tells you exactly one thing:
The link is broken.
It does not tell you what the memory should become.
A broken link is a question, not an instruction.
So instead of bulk cleanup, we built an adjudication harness. Every memory gets evidence, an independent challenge, and a rollback path before it moves an inch.
The link that looked easy
One unresolved link was [[mani-voice]].
That sounds like a missing file.
Just create mani-voice.md, right?
Except there was already a canonical voice corpus. There were project-specific writing rules. There were runtime instructions. There were stable user preferences that had just been canonicalized across Claude, Codex, and Hermes.
Creating another file would fix the link by rebuilding the duplication problem we had just spent days removing.
So maybe the link should point to the voice corpus.
Unless the project memory system cannot retrieve that cross-store target.
Maybe it should point to the canonical user rule.
Unless the source record is about writing procedure, not user preference.
Maybe it should become plain text because it was never supposed to be a link.
Every option makes the scanner happy.
Every option creates a different system.
The scanner cannot choose because the scanner does not know what mani-voice owns.
It can beep.
It cannot tell you whether it found a coin, a pipe, or a land mine.
That is where cleanup stops being mechanical.
Clean is not correct
Most cleanup logic is built from signals that are easy to count.
Old file. Archive it.
Unindexed file. Add it.
Duplicate text. Merge it.
Broken link. Replace it.
Stale claim. Delete it.
Beautiful report.
Potentially disastrous memory.
An old incident may explain a current safeguard. An unindexed file may be intentionally quarantined. Two similar records may belong to different profiles. A procedure may deserve a tested skill instead of a longer paragraph. A stale installation claim may need to become dated history because the sequence of what broke and what fixed it still matters.
The filename does not know that.
The modification date does not know that.
The index does not know that.
Clean is not correct.
A system can have zero broken links and still feed the agent obsolete claims, misplaced rules, and polished little lies.
That is worse than a visible broken link.
At least the broken link admits there is a question.
Every memory gets a case
The solution is slower than bulk cleanup.
It is faster than recovering from a bad one.
Each unresolved memory, or inseparable cluster of conflicting memories, gets a case ID.
The case includes the complete record. Not the filename. Not the opening paragraph. Not the one-line index summary.
The whole damn thing.
Frontmatter. Backlinks. Conflicts. Possible destinations. Current evidence. Retrieval path. Every factual claim that could disappear during a tidy little merge.
Then the work separates.
A coordinator reserves the source and destination stores so two workers cannot edit the same index or collide during a cross-store move.
A researcher separates durable facts from dated history, volatile state, procedure, instruction, sensitive material, and claims nobody can currently prove.
Then the researcher has to propose more than one outcome.
Keep it. Correct it. Split it. Merge it. Move it. Preserve the incident as history. Turn the volatile part into a live verification command. Point across stores. Recommend a skill. Quarantine the thing until the evidence catches up.
Or do something custom because the memory refuses to fit the boxes.
The categories serve the memory.
The memory does not serve the categories.
Then an independent evaluator tries to kill the recommendation.
That part matters.
The person who fell in love with a clever reorganization should not be the only person deciding whether it loses information.
Is the proposed owner correct? Is the evidence current? Did the researcher search the destination for collisions? Will the right sessions still retrieve the record? Is a project rule about to leak into every profile? Is a resolved incident being erased because closed looked the same as useless?
The evaluator returns one of four answers: pass, revise, blocked, or no change.
No pass, no mutation.
Only then does the implementer touch the file.
One logical case at a time. Original bytes preserved. Quarantine instead of deletion. Indexes updated. Backlinks checked. Retrieval tested. Rollback proved.
Seven links.
Seven evidence trails.
Seven reversible decisions.
Not seven reflexes.
A green check has to earn the right to be trusted
There is one more problem.
What if the harness approves everything because the harness is blind?
So before it touches live memory, it gets planted cases.
A clear keep. A real move. A mixed record that must split. A stale claim that needs correction. A genuine duplicate. A resolved incident that must remain history. A procedure that belongs in a skill. A protected instruction that requires approval. A sensitive record that must stop the workflow. A weird case where every standard answer is wrong.
The evaluator has to call each one correctly.
Every scanner that claims zero hits has to detect a planted hit first.
A green check counts only after a red check works.
That rule would have prevented the first false-clean audit, when a shell loop mishandled project paths and the scanner treated TOML inside a code fence like a missing memory.
The check was green.
The checker was wrong.
Now the checker has to prove it can fail.
What this changes when an agent answers you
The first tier made the agents remember me consistently.
This tier helps them remember reality consistently.
Fewer contradictions between related project memories. Fewer dead references. Less obsolete runtime state presented as current. Less project information leaking into the wrong profile. Better preservation of why a decision was made. Fewer procedures duplicated as prose.
And when something is history, the system can say it is history without pretending it never happened.
That matters because memory has to answer more than What is true?
It also has to answer:
Why did we decide this?
What did it replace?
Who needs to know it?
What breaks if we undo it?
Delete the wrong old memory and the next agent gets to relearn the lesson with fresh damage.
Preserve everything indiscriminately and the agent has to fight through every version before it can act.
The job is neither hoarding nor cleaning.
The job is judgment.
The machine found seven questions.
Now each one gets an answer specific enough to preserve the truth and reversible enough to survive being wrong.
Clean is not correct.
And a broken link is a question, not an instruction.
Leave a Reply