From cd9d9d03b31f98a68f2e20a02a65d8ca9e842aae Mon Sep 17 00:00:00 2001 From: Michael Jordan Date: Wed, 29 Jul 2026 13:49:44 -0400 Subject: [PATCH] docs(agent): explain micro-compaction Covers what it does, the head/tail protection, the cursor and rolling summary, defrag, how the session DB is kept in step, and the failure paths. States the tradeoff up front: compression cost is amortized across turns, at the price of older detail becoming summarized earlier in a session than batch-only compaction would. Co-Authored-By: Claude Fable 5 --- docs/micro-compaction.md | 156 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 156 insertions(+) create mode 100644 docs/micro-compaction.md diff --git a/docs/micro-compaction.md b/docs/micro-compaction.md new file mode 100644 index 00000000000..5c712a8ad36 --- /dev/null +++ b/docs/micro-compaction.md @@ -0,0 +1,156 @@ +# Micro-compaction + +**A way to amortize the cost of compression.** + +Long conversations eventually outgrow the model's context window, and something +has to be thrown away or summarized. Hermes has always done this in one batch: +when the transcript crosses a threshold, the session stops, a large chunk of the +middle is summarized in a single call, and the conversation resumes. That works, +but the whole bill comes due at once — one visible pause, one big summarization +request, at whatever moment you happened to cross the line. + +Micro-compaction pays the same bill in instalments. After each completed turn, +Hermes folds the single oldest un-absorbed exchange into a running summary. The +work is the same work; it just happens continuously, in small pieces, during the +idle moment after a response instead of all at once in the middle of your +session. + +**The tradeoff is that knowledge gets a little earlier than you may be used to.** +Because compaction is always running, older parts of the conversation become +summaries sooner than they would under batch compaction — which leaves +everything verbatim until the window actually fills. Detail from earlier in the +session turns second-hand faster. You trade some of that fidelity for never +eating one long stall, and for a context window that stays consistently smaller +rather than sawtoothing up to the threshold and back. + +--- + +## What it does + +After every turn that finishes normally, `finalize_turn` asks the context +compressor to absorb **one** exchange: + +1. Find the oldest exchange that hasn't been summarized yet. +2. Send just that exchange, plus the current running summary, to the auxiliary + summarization model. +3. Replace those messages in the transcript with a single summary marker + carrying the updated running summary. + +One exchange per turn. The per-turn cost stays bounded no matter how long the +conversation gets. + +An **exchange** is an assistant message together with any tool results that +followed it. In tool-heavy work that's where the bulk of the tokens live — a +file read or a command's output dwarfs the surrounding prose — which is why +absorbing one exchange at a time is worth doing at all. + +## What it never touches + +Two regions are protected and stay verbatim: + +- **The head** — the system prompt and the opening messages, so the session's + founding instructions are never paraphrased. +- **The tail** — a token-budgeted window of the most recent messages, so + everything that's immediately relevant is still there in full. + +Micro-compaction only ever works in the middle, between those two. + +## How it works + +### The cursor + +The compressor keeps a cursor: the index of the first message not yet absorbed. +Each successful pass advances it past the exchange it just summarized. + +If that in-memory cursor is missing or out of range — a fresh process, a resumed +session — it's recovered by scanning the transcript for the last summary marker +and resuming just after it. The transcript itself is the source of truth, so +resuming a session doesn't re-summarize work already done. + +### The rolling summary + +Rather than keeping a pile of per-exchange summaries, there is exactly one +running summary that each new exchange is merged into. The summarizer is asked +to fold in the new material's decisions, requirements, file paths and open +questions, drop details that are no longer relevant, and preserve the existing +structure. It's also explicitly instructed to replace any credentials it +encounters with `[REDACTED]`. + +Because that summary is cumulative, only the newest marker is kept in the +transcript. Earlier markers are strictly redundant — the current summary already +contains everything they held — so they're dropped as they're superseded. This +matters more than it sounds: leaving them in place stacks near-duplicate copies +of the same text, each with its own heading and end-marker scaffolding, and the +transcript grows on every turn instead of shrinking. + +### Defrag + +Merge into a summary often enough and it gets baggy — repetitive, and larger +than the material justifies. When the running summary crosses a token threshold +(2000 by default), the next pass **defrags**: it re-summarizes the summary and +whatever middle remains in one shot, replacing it with a fresh compact version +and advancing the cursor to the tail. + +This is still much cheaper than full batch compaction. It only ever processes the +summary plus the un-absorbed middle, never the whole transcript. + +### Staying in step with the session database + +The in-memory splice alone isn't enough. Hermes's normal session flush is +append-only, so the original rows would stay marked active and a resume would +load *both* the summary and the messages it replaced — putting the session +straight over the context limit. + +So each pass also calls `archive_and_compact`, which atomically soft-archives the +active rows and inserts the compacted set. The messages are then stamped as +already-persisted so the append-only flush that follows skips them. If that +database step fails, it's logged and the session continues; the resume would +double-load until the next batch compression cleans up. + +### When the summarizer fails + +A summarization call can fail — the auxiliary model is unreachable, out of quota, +or the exchange itself is somehow unsummarizable. The transcript is left +untouched and the failure is counted. + +If the *same* exchange fails three times in a row, the cursor is advanced past it +anyway. Without that, one bad exchange would be retried on every single turn +forever. Those skipped messages stay in the transcript and get picked up by the +next defrag or batch compaction. + +## Interaction with batch compaction + +Micro-compaction doesn't replace batch compaction — it defers it. Threshold-based +compaction is still there and still fires if the window fills anyway, and its +summary markers are the same format, so the two interoperate. In practice +micro-compaction keeps the transcript far enough below the threshold that the +batch path fires much less often. + +## Configuration + +```yaml +compression: + micro_compact: true # default +``` + +Set it to `false` to disable micro-compaction and return to batch-only +compaction. Everything else about compression is unchanged. + +## What you'll see in the logs + +``` +Micro-compaction: 37 -> 36 messages +Micro-compaction defrag: rolling summary re-summarized (1843 chars) +Micro-compaction: skipping exchange at cursor 12 after 3 consecutive failures +``` + +Message counts move by small amounts — that's the point. The token count is where +the effect shows: absorbing one tool-heavy exchange can drop hundreds of tokens +while changing the message count by one or two. + +## Failure behaviour + +Micro-compaction is best-effort throughout. The call in `finalize_turn` is wrapped +so that any exception is logged and swallowed — a failure returns the conversation +unchanged and the turn completes normally. It can degrade, but it shouldn't be +able to break a session.