A companion to Part 4 of the Building the AI Memory Stack series.
Part 4.5 of the series.
Part 4 argued that agentic systems need a Reasoning Ledger: a layer that preserves why a decision happened, not just what was decided.
The comment thread that followed turned into something more specific and more useful, a working design conversation about what a single ledger record should actually contain.
This piece consolidates that.
Several of the strongest ideas below arrived from other people, and I have tried to credit them where they land.
The easy version of this article is a schema.
Here are the fields, copy them, done.
I want to resist that, because the field list is the least durable thing I could hand you.
Implementations differ, field names drift, and a record shape copied without its reasoning becomes cargo-cult structure that nobody maintains.
The useful thing is the set of design tensions that decide what belongs in the record and what does not.
Get those right and you can derive the fields yourself.
Get them wrong and no schema will save you.
So this is principles first, record second.
At the end there is a worked record and a field reference, tagged for what is core and what is genuinely optional.
A Starting Point Here is the baseline record from Part
4.
It is a reasonable start and, as the thread quickly established, incomplete in instructive ways.
Every principle below is, in effect, a thing this record does not yet say.
Principle 1: The Ledger Witnesses, It Does Not Enforce The first tension is architectural, and it is the one I would defend hardest.
A reasoning ledger must not be able to block, veto, or gate the action it records.
Its job is to preserve what happened and what evidence surrounded it.
The moment the ledger can prevent an action, it stops being an independent witness and becomes part of the mechanism it is supposed to describe, and its own records stop being examinable as neutral fact.
This came up when pm25coder noted, correctly, that a ledger that only narrates can quietly become fiction, and that trust comes from being able to gate rather than merely describe.
I agree with the diagnosis and draw the boundary one step earlier: enforcement is real and necessary, but it belongs at the policy and tool boundary, not inside the witness.
The ledger preserves that the boundary was evaluated and what it returned.
The boundary decides whether the action proceeds.
The practical consequence for the record: a ledger entry can contain a result showing that a check ran and what it concluded, but it never contains the enforcement decision as its own authority.
It reports; it does not rule.
Core.
This is not a field, it is a constraint on the whole design.
Principle 2: Supersession Is a New Event, Never a Rewrite A superseded decision should become a new record that points back at the old one.
It should never overwrite the original. "We decided A, and later decided B instead" is two events with a relationship between them, not one field that changed value.
This matters because "wrong now" does not mean "was never decided then." If you rewrite the March record when you change course in August, you have destroyed the ability to answer whether the March decision was reasonable given what was known in March.
The noisier history is the correct trade.
Compaction can always produce a clean current-state projection later, but once you have rewritten the historical evidence, you cannot reconstruct it.
This is the same append-only discipline that makes Forensic Receipts useful: preserve what was decided under which evidence and authority, then record the superseding decision as its own event with its own receipt.
Core.
Principle 3: Record How the Authority Was Obtained, Not Just Which One The baseline record says .
That tells a future reader what supposedly governed.
It does not tell them how the system established that version 7 was authoritative at decision time, and those are very different trust claims.
Self-Correcting Systems and pm25coder arrived at this from opposite directions and met in the middle: a policy version fetched fresh from its authority at 09:22, a version read from a five-minute cache, and a version inherited from session state can produce identical fields while supporting completely different claims about what the system could reasonably have known.
The fix is to treat the authority fetch itself as a recorded event.
The record should say which source was consulted, when, what came back, and whether cached state was involved.
This also exposes the sharpest failure mode in the thread, the one an otherwise perfect ledger cannot catch on its own.
If the external authority moved to version 8 an hour before your decision and nothing in your system observed that change, the record faithfully captures version 7 and stays perfectly self-consistent.
It is a flawless account of a decision that was already wrong when it was made.
The record cannot flag this, because there is no edge to preserve; nothing inside the system ever saw the change.
Recording how the version was obtained at least lets a later examiner distinguish "we checked and got stale data" from "we never checked." Core for the fact of how evidence was obtained.
The revalidation mechanism that catches silent version drift lives outside the record, and Principle 7 covers it.
Principle 4: Relationships Need Two Clocks If you ever want to reconstruct what the system could have known at a past moment, every relationship in the ledger needs two timestamps, not one.
This is standard bitemporal modeling, and Giulio D'Erme named exactly why it is not optional here.
Valid time is when a fact was true in the world.
Transaction time is when your system asserted or learned the relationship.
If a supersession edge carries only a single date, replaying last March will show March's decision annotated with August's supersessions, and the decision-maker will look like they ignored a policy that did not yet exist.
You will have judged a past decision using knowledge that arrived in the future, which is the precise thing a reasoning ledger exists to prevent.
So a supersession or correction relationship carries both (when the new state became true) and (when the system recorded the edge).
Reconstruction filters on to see only what was knowable then.
Core for any ledger whose purpose includes reconstructing historical decision context.
If you genuinely only ever query current state, you can defer this, but that is a smaller ambition than most of these systems have.
Principle 5: Preserve What Lost, Not Just What Won A ledger that records only the evidence supporting the final decision is a post-hoc justification engine wearing an audit trail.
You can reconstruct why the decision looked reasonable, and you have quietly lost what competed with it, what failed a threshold, and what stayed unresolved.
GnomeMan4201 made this case from the investigation side, and it reframed the record for me.
An immutable ledger can preserve history perfectly and still preserve a biased history if the losing evidence never gets written.
The distinction between "we chose A because of X" and "we chose A because of X, rejected B because of Y, and could not resolve Z" is enormous when someone later asks whether the decision was defensible given what was actually known.
The fields this implies: with a for each, for evidence that actively cut against the chosen path, and or for what the system could not resolve at decision time.
A scoping note, in answer to Kartik N V J K, who asked whether to capture rejected branches: capture the alternatives that were explicit parts of the decision process, not an exhaustive reconstruction of every path the model internally considered.
If the agent evaluated three tools and rejected two on policy grounds, those rejections are observable decision evidence and belong in the record.
The model's private deliberation does not.
Observable reasoning is architecture; private reasoning belongs t