- Cut the leftover telemetry.* DEFAULT_CONFIG block (nothing reads it) and
the legacy telemetry-key fallback in policy.py.
- install_id: persist the minted UUID back to config.yaml on first use so
service.instance.id survives gateway restarts (fail-open when the write
is not possible); regression test covers the restart path.
- Remove the no-op gateway_health_export.redaction config keys. Redaction
is always-on by design and deliberately not configurable; status output
now says so.
- Collapse redaction to one unconditional secrets+PII scrub: drop the
none/pii content modes (they served the dropped trajectories plane) and
fold gateway_health.py's duplicate bearer/token/email/phone regex layer
into agent/monitoring/redaction.py.
Salvages the event-spine foundation from feat/telemetry-observability
(emitter, typed events, OTLP streaming, redaction — authorship preserved in
the preceding commits) and scopes it to the plane enterprise operators need
today: gateway Service Health Monitoring plus redacted Operational
Diagnostics, exported over OTLP.
Dropped from the salvaged branch, deliberately:
- run/model/tool trajectory capture (plugins/telemetry hooks, tel_spans)
- the local JSONL + state.db tel_* store (monitoring is egress, not storage)
- usage rollups/metrics, /insights integration, bulk export
- hermes telemetry CLI (replaced by hermes monitoring status)
Those planes — shared client usage metrics and enterprise trace telemetry —
are being designed on the NeMo Relay integration with distinct consent,
policy, and export boundaries; this keeps the monitoring plane content-free
and independently enableable.
Renames agent/telemetry -> agent/monitoring, config telemetry.* ->
monitoring.*, and pins the otlp extra at OpenTelemetry 1.39.1 (matching
uv.lock; 1.30.0 conflicts with mistralai>=2.4 on opentelemetry-api).