mirror of
https://github.com/NousResearch/hermes-agent.git
synced 2026-07-31 19:16:29 +00:00
docs(agent): explain micro-compaction
Covers what it does, the head/tail protection, the cursor and rolling summary, defrag, how the session DB is kept in step, and the failure paths. States the tradeoff up front: compression cost is amortized across turns, at the price of older detail becoming summarized earlier in a session than batch-only compaction would. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
parent
186cad02f9
commit
cd9d9d03b3
1 changed files with 156 additions and 0 deletions
156
docs/micro-compaction.md
Normal file
156
docs/micro-compaction.md
Normal file
|
|
@ -0,0 +1,156 @@
|
|||
# Micro-compaction
|
||||
|
||||
**A way to amortize the cost of compression.**
|
||||
|
||||
Long conversations eventually outgrow the model's context window, and something
|
||||
has to be thrown away or summarized. Hermes has always done this in one batch:
|
||||
when the transcript crosses a threshold, the session stops, a large chunk of the
|
||||
middle is summarized in a single call, and the conversation resumes. That works,
|
||||
but the whole bill comes due at once — one visible pause, one big summarization
|
||||
request, at whatever moment you happened to cross the line.
|
||||
|
||||
Micro-compaction pays the same bill in instalments. After each completed turn,
|
||||
Hermes folds the single oldest un-absorbed exchange into a running summary. The
|
||||
work is the same work; it just happens continuously, in small pieces, during the
|
||||
idle moment after a response instead of all at once in the middle of your
|
||||
session.
|
||||
|
||||
**The tradeoff is that knowledge gets a little earlier than you may be used to.**
|
||||
Because compaction is always running, older parts of the conversation become
|
||||
summaries sooner than they would under batch compaction — which leaves
|
||||
everything verbatim until the window actually fills. Detail from earlier in the
|
||||
session turns second-hand faster. You trade some of that fidelity for never
|
||||
eating one long stall, and for a context window that stays consistently smaller
|
||||
rather than sawtoothing up to the threshold and back.
|
||||
|
||||
---
|
||||
|
||||
## What it does
|
||||
|
||||
After every turn that finishes normally, `finalize_turn` asks the context
|
||||
compressor to absorb **one** exchange:
|
||||
|
||||
1. Find the oldest exchange that hasn't been summarized yet.
|
||||
2. Send just that exchange, plus the current running summary, to the auxiliary
|
||||
summarization model.
|
||||
3. Replace those messages in the transcript with a single summary marker
|
||||
carrying the updated running summary.
|
||||
|
||||
One exchange per turn. The per-turn cost stays bounded no matter how long the
|
||||
conversation gets.
|
||||
|
||||
An **exchange** is an assistant message together with any tool results that
|
||||
followed it. In tool-heavy work that's where the bulk of the tokens live — a
|
||||
file read or a command's output dwarfs the surrounding prose — which is why
|
||||
absorbing one exchange at a time is worth doing at all.
|
||||
|
||||
## What it never touches
|
||||
|
||||
Two regions are protected and stay verbatim:
|
||||
|
||||
- **The head** — the system prompt and the opening messages, so the session's
|
||||
founding instructions are never paraphrased.
|
||||
- **The tail** — a token-budgeted window of the most recent messages, so
|
||||
everything that's immediately relevant is still there in full.
|
||||
|
||||
Micro-compaction only ever works in the middle, between those two.
|
||||
|
||||
## How it works
|
||||
|
||||
### The cursor
|
||||
|
||||
The compressor keeps a cursor: the index of the first message not yet absorbed.
|
||||
Each successful pass advances it past the exchange it just summarized.
|
||||
|
||||
If that in-memory cursor is missing or out of range — a fresh process, a resumed
|
||||
session — it's recovered by scanning the transcript for the last summary marker
|
||||
and resuming just after it. The transcript itself is the source of truth, so
|
||||
resuming a session doesn't re-summarize work already done.
|
||||
|
||||
### The rolling summary
|
||||
|
||||
Rather than keeping a pile of per-exchange summaries, there is exactly one
|
||||
running summary that each new exchange is merged into. The summarizer is asked
|
||||
to fold in the new material's decisions, requirements, file paths and open
|
||||
questions, drop details that are no longer relevant, and preserve the existing
|
||||
structure. It's also explicitly instructed to replace any credentials it
|
||||
encounters with `[REDACTED]`.
|
||||
|
||||
Because that summary is cumulative, only the newest marker is kept in the
|
||||
transcript. Earlier markers are strictly redundant — the current summary already
|
||||
contains everything they held — so they're dropped as they're superseded. This
|
||||
matters more than it sounds: leaving them in place stacks near-duplicate copies
|
||||
of the same text, each with its own heading and end-marker scaffolding, and the
|
||||
transcript grows on every turn instead of shrinking.
|
||||
|
||||
### Defrag
|
||||
|
||||
Merge into a summary often enough and it gets baggy — repetitive, and larger
|
||||
than the material justifies. When the running summary crosses a token threshold
|
||||
(2000 by default), the next pass **defrags**: it re-summarizes the summary and
|
||||
whatever middle remains in one shot, replacing it with a fresh compact version
|
||||
and advancing the cursor to the tail.
|
||||
|
||||
This is still much cheaper than full batch compaction. It only ever processes the
|
||||
summary plus the un-absorbed middle, never the whole transcript.
|
||||
|
||||
### Staying in step with the session database
|
||||
|
||||
The in-memory splice alone isn't enough. Hermes's normal session flush is
|
||||
append-only, so the original rows would stay marked active and a resume would
|
||||
load *both* the summary and the messages it replaced — putting the session
|
||||
straight over the context limit.
|
||||
|
||||
So each pass also calls `archive_and_compact`, which atomically soft-archives the
|
||||
active rows and inserts the compacted set. The messages are then stamped as
|
||||
already-persisted so the append-only flush that follows skips them. If that
|
||||
database step fails, it's logged and the session continues; the resume would
|
||||
double-load until the next batch compression cleans up.
|
||||
|
||||
### When the summarizer fails
|
||||
|
||||
A summarization call can fail — the auxiliary model is unreachable, out of quota,
|
||||
or the exchange itself is somehow unsummarizable. The transcript is left
|
||||
untouched and the failure is counted.
|
||||
|
||||
If the *same* exchange fails three times in a row, the cursor is advanced past it
|
||||
anyway. Without that, one bad exchange would be retried on every single turn
|
||||
forever. Those skipped messages stay in the transcript and get picked up by the
|
||||
next defrag or batch compaction.
|
||||
|
||||
## Interaction with batch compaction
|
||||
|
||||
Micro-compaction doesn't replace batch compaction — it defers it. Threshold-based
|
||||
compaction is still there and still fires if the window fills anyway, and its
|
||||
summary markers are the same format, so the two interoperate. In practice
|
||||
micro-compaction keeps the transcript far enough below the threshold that the
|
||||
batch path fires much less often.
|
||||
|
||||
## Configuration
|
||||
|
||||
```yaml
|
||||
compression:
|
||||
micro_compact: true # default
|
||||
```
|
||||
|
||||
Set it to `false` to disable micro-compaction and return to batch-only
|
||||
compaction. Everything else about compression is unchanged.
|
||||
|
||||
## What you'll see in the logs
|
||||
|
||||
```
|
||||
Micro-compaction: 37 -> 36 messages
|
||||
Micro-compaction defrag: rolling summary re-summarized (1843 chars)
|
||||
Micro-compaction: skipping exchange at cursor 12 after 3 consecutive failures
|
||||
```
|
||||
|
||||
Message counts move by small amounts — that's the point. The token count is where
|
||||
the effect shows: absorbing one tool-heavy exchange can drop hundreds of tokens
|
||||
while changing the message count by one or two.
|
||||
|
||||
## Failure behaviour
|
||||
|
||||
Micro-compaction is best-effort throughout. The call in `finalize_turn` is wrapped
|
||||
so that any exception is logged and swallowed — a failure returns the conversation
|
||||
unchanged and the turn completes normally. It can degrade, but it shouldn't be
|
||||
able to break a session.
|
||||
Loading…
Add table
Add a link
Reference in a new issue