Self-review after the #51714 feedback found the reviewer's dead-table finding
was not isolated — the schema advertised far more than the code populates, and
our own tests hid it by hand-feeding fields production never sends. Make the
surface honest by subtraction.
Schema (10 tel_* tables -> 5):
- Delete tel_gateway_events, tel_cron_events, tel_skill_events,
tel_memory_events, tel_feedback_events — declared, never written, never read.
- Drop columns nothing populates: tel_runs.{profile_id,estimated_cost_usd,
cost_status}; tel_model_calls.{ttft_ms,estimated_cost_usd,cost_status,
cost_source,end_reason,retry_count}; tel_tool_calls.{backend,retry_count,
approval}; tel_spans.attrs_json. Cost duplicated the existing sessions
billing columns and was always NULL here.
- events.py / emitter _TABLE_COLUMNS / OTLP _span_attrs / rollup / preview
display all trimmed to match.
Correctness:
- end_reason no longer hardcodes "completed". Production finalize callers pass
`reason` (shutdown/session_expired/session_reset); _coarse_end_reason now
reads it and maps accordingly.
- Fix a latent bug the trim exposed: the model_call hook passed end_reason= to
ModelCallEvent, which the @_safe wrapper was silently swallowing — so
tel_model_calls dropped every row in real runs. Now writes correctly.
Tests:
- Stop hand-feeding estimated_cost_usd / turn_exit_reason that no production
call site sends. Finalize is now driven with the real `reason` kwarg, and
assertions cover only fields that are actually populated. This is what let
the model_call drop hide — the suite graded on a fictional contract.
Net: a smaller system that does what it says. Verified end-to-end over the real
dispatch path (runs + connected span tree + model/tool rows populate; dead
tables gone). 160 telemetry/state/insights tests green.
Addresses the review on #51714: the trace/span layer was declared but unwired —
tel_spans was never written, call rows had no timestamp, and nothing set parent
lineage, so the store was metrics-only and couldn't reconstruct a trace.
Wire the span layer (keeping the praised star-schema shape):
- New SpanEvent (span_id/trace_id/run_id/parent_span_id/name/kind/start_ns/end_ns)
mapped into tel_spans via the emitter's _TABLE_COLUMNS.
- The plugin mints a root span per run and, on each model/tool call, emits a
SpanEvent (timing + parent = the run's root) keyed by the SAME span_id as the
detail row, so tel_model_calls / tel_tool_calls JOIN to their span.
- Call hooks fire on completion, so end_ns = now and start_ns is reconstructed
from the measured latency/duration. The run's root span is emitted at finalize
with the true run start/end.
Result: tel_spans is a connected, single-trace_id, run -> calls tree a desktop
waterfall (or any reader) can render directly, ordered by start_ns. Existing
metrics rows (tel_runs/model_calls/tool_calls) are unchanged.
OTLP: spans now flow to the exporter with their trace/parent/timing attributes.
The exporter still emits one OTel span per event rather than reconstructing OTel
SpanContexts into a connected trace tree; that projection is left for a follow-up
and the module docstring now says so plainly instead of over-claiming.
Adds test_spans_trace.py (connected-tree + detail-row JOIN) over the real dispatch
path. Accurate (pre-hook) start times, real OTLP SpanContexts, and subagent
cross-run lineage remain follow-ups.
Aggregate metrics are derived from the local tel_* tables — they're a coarsened
view of local data, not an independent capture path. With telemetry.local=false
nothing is written, so an aggregate opt-in had nothing to aggregate, yet
may_upload_aggregate() returned True and `status` showed "Aggregate metrics: on".
The config could claim a state it couldn't fulfill.
Gate aggregate on local being enabled:
- may_upload_aggregate() now requires local_enabled AND allow_aggregate AND
consent_state == aggregate.
- `telemetry status` computes aggregate_enabled the same way and, when consent is
aggregate but local is off, prints "inert: local telemetry is off — nothing to
aggregate" instead of the opt-in hint.
Happy path is unchanged (local on + consent aggregate -> on). Adds policy and CLI
tests for the inert combo.
The existing hook tests call the plugin's _on_* callbacks directly, which passes
even if the bundled plugin stops auto-loading or a hook name drifts from what core
fires — real runs would go dark while the suite stays green.
Add test_plugin_e2e.py, which drives the real dispatch chain through public entry
points only (discover_plugins -> invoke_hook -> registered callback -> emitter ->
tel_* tables), exactly as core does:
- one completed turn produces tel_runs / tel_model_calls / tel_tool_calls rows
with real provider/model/tool values and correct counts;
- telemetry.local=false means the plugin does not load and nothing is written.
Verified robust against test ordering (singleton resets for the plugin manager and
the emitter in the fixture).
Rename the telemetry tiers away from the borrowed control-plane/data-plane
jargon to plain language, across code, CLI output, config, and docs:
- "local plane" -> "local telemetry"
- "aggregate plane" -> "aggregate metrics"
- "trajectories plane" -> "trajectories" / "telemetry.trajectories"
- "three planes with a hard wall" -> "three settings, isolated from each other"
User-facing `hermes telemetry status` now reads "Local telemetry: on" /
"Aggregate metrics: off" / "Content export: off (trajectories disabled)".
The OTLP resource attribute key telemetry.plane is renamed to telemetry.scope
(wire-level identifier; nothing consumes it yet).
No behavior change — wording only. Status renders identically apart from the
labels; tests updated to match the new strings.
policy.resolve() / TelemetryDecision was a read-only projection used only by
`hermes telemetry status` for display. The actual behavior gates already read
telemetry.* straight from config: the emitter (whether to write) and the plugin
loader (whether to auto-load) each call .get("local", True) on the loaded config,
never through policy.
Make config the single chokepoint the status command reads too: it now resolves
local/allow_aggregate/consent_state inline from the loaded config, the same way
the other gates do. policy.py keeps only what config can't express on its own —
the consent constants, ensure_install_id(), and may_upload_aggregate(config) as a
pure function (the gate a future uploader must consult). resolve() and the
TelemetryDecision dataclass are removed; policy.py drops 107 -> 70 lines.
No behavior change: status renders identically, and the default-on local plane is
still defaulted in DEFAULT_CONFIG plus a fail-safe .get(..., True) at each gate.
Add a built-in telemetry system that records what the agent does — workflows,
model calls, tool calls, errors — to the local machine, powers `/insights`, and
can export to an operator-chosen destination. Default-on locally; nothing leaves
the machine unless the user exports it or opts into the aggregate plane.
Three planes with a hard wall between them:
- local: full-fidelity observability (real model/provider/tool names), on by
default, never leaves the machine.
- aggregate: opt-in metadata, default off. No uploader ships — consent is
recorded via telemetry.consent_state, and `preview` shows what would be
produced, computed locally.
- trajectories: full message content, opt-in, exported only to the operator's
own destination.
Mechanism:
- Bundled `telemetry` plugin registers observational lifecycle hooks
(on_session_start / post_api_request / post_tool_call / on_session_finalize).
No core call sites are edited; hooks already carry the data.
- Fire-and-forget emitter: emit() returns in microseconds, never blocks or
raises into a model/tool call. A daemon thread writes events to an
append-only JSONL log and the tel_* tables in state.db (its own sqlite
connection, separate from SessionDB).
- tel_runs / tel_model_calls / tel_tool_calls live in the declarative
SCHEMA_SQL and are reconciled automatically; SCHEMA_VERSION 16 -> 17.
- metrics derives rollups for /usage and /insights; rollup builds per-run
summaries for `hermes telemetry preview`.
Consent is config, not a parallel command surface. The config file is the root
of trust: set telemetry.consent_state with `hermes config set`, or pin any
telemetry.* key (including allow_aggregate) via managed scope, which overrides
the user's value per key. `hermes telemetry` exposes only what config cannot:
status (report), preview (query), and export.
Export:
- exporter_bulk writes telemetry (and, when the trajectories plane is enabled,
session content) to ndjson/json.
- otlp_exporter streams spans to a configured OpenTelemetry Collector over
OTLP/HTTP. The SDK is an optional extra (hermes-agent[otlp]), lazily
installed via tools.lazy_deps on first use.
- Secrets are always redacted on every export path
(redact_sensitive_text(force=True)); content export is gated by the
trajectories plane, and PII scrubbing follows telemetry.content_redaction.
OTLP auth headers reference environment variable names, never inline values.
No outbound emission to Nous. The aggregate uploader is intentionally not built.