4.8 KiB
Gateway Monitoring
Service health monitoring plus structured operational diagnostics for the Hermes gateway daemon, exported over OTLP/HTTP to an operator-configured endpoint (OpenTelemetry Collector, DataDog, or any OTLP receiver).
This plane is content-free by construction. It exports gateway and cron
lifecycle state, platform connector health, and content-free warning/error
diagnostics. It never exports prompts, messages, tool arguments or results,
job names, destinations, schedules, raw errors, session history, usage
analytics, audit logs, or detailed execution traces. Run/model/tool trajectory
capture is a separate plane served by the NeMo Relay integration
(plugins/observability/nemo_relay/) and its Hermes-owned subscribers.
What gets exported
| Signal | OTLP route | Content |
|---|---|---|
| Gateway gauges | /v1/metrics |
hermes.gateway.up/state/busy/drainable/active_agents/restart_requested, hermes.platform.up/degraded with bounded error_code attributes |
| Health/lifecycle events | /v1/traces |
gateway.lifecycle state transitions (starting -> running -> draining -> stopped, startup_failed, exit), gateway.health_snapshot, platform state changes |
| Diagnostics | /v1/logs |
Warning/error gateway events with a constant body and bounded subsystem, severity, error class, and error code attributes; rendered log messages are never exported |
| Cron scheduler gauges | /v1/metrics |
Ticker heartbeat and last-success age (omitted when unavailable), enabled/running job counts, and overdue count derived from persisted next_run_at plus the scheduler's existing grace rule |
| Cron execution lifecycle | /v1/traces |
Durable claimed/running/completed/failed/unknown states, bounded source and error class, opaque hashed job key, elapsed duration when timestamps exist, and delivery outcome when the scheduler knows it; terminal states flush through a bounded fail-open barrier |
Signals carry service.name, version, supervision mode, and a stable one-way
hash of the install id so an operator can distinguish instances without
exporting account/profile identity or the raw install identifier.
Enabling
# config.yaml
monitoring:
gateway_health_export:
enabled: true
export:
otlp:
enabled: true
endpoint: http://collector-host:4318/v1/traces # metrics/logs derive
headers_env: {} # header name -> ENV VAR NAME (values never stored)
Check the posture any time:
hermes monitoring status
The OpenTelemetry SDK is an optional extra (pip install 'hermes-agent[otlp]'),
lazy-installed on first use. When the SDK is missing or the endpoint is down,
the gateway runs unaffected: every export path is fail-open and off the hot
path (events flow through a fire-and-forget in-process emitter; a slow or
failing exporter can never block gateway code).
Works identically under systemd/launchd/s6 supervision, containers, tmux, or
a plain hermes gateway run — the exporter lives in the gateway process, so
no sidecar, agent, or collector is required on the host.
Collecting into DataDog
Run a customer-owned OpenTelemetry Collector and forward:
# otel-collector config
receivers:
otlp:
protocols:
http:
exporters:
datadog:
api:
key: ${env:DD_API_KEY}
service:
pipelines:
metrics: {receivers: [otlp], exporters: [datadog]}
traces: {receivers: [otlp], exporters: [datadog]}
logs: {receivers: [otlp], exporters: [datadog]}
Point monitoring.export.otlp.endpoint at the collector. Alerts belong on
hermes.gateway.up, hermes.platform.up, and hermes.platform.degraded.
Local smoke test (no Docker)
# terminal 1: capture collector on :4318
python scripts/observability/otel_capture_collector.py \
--host 127.0.0.1 --port 4318 --log /tmp/hermes_otel_capture.jsonl
# terminal 2: drive the real exporter through lifecycle transitions,
# a fatal platform, and a structured warning event, then flush
python scripts/observability/gateway_health_export_probe.py \
--endpoint http://127.0.0.1:4318/v1/traces \
--log /tmp/hermes_otel_capture.jsonl --wait 8
# exit 0 prints: {"requests": 6, "paths": ["/v1/logs", "/v1/metrics", "/v1/traces"]}
Boundaries and roadmap
The hermes monitoring CLI intentionally exposes status only. This first
release covers only Hermes Agent-owned service-health and operational-diagnostic
signals. Team Gateway/shared-connector signals are explicitly out of scope, as
are product analytics, audit/quality reporting, and detailed execution traces.
Shared client usage metrics and enterprise trace telemetry are being designed on
the NeMo Relay integration with their own consent, policy, and export
boundaries; this monitoring plane stays narrow so an operator can enable it
without touching any content-bearing signal. The telemetry surface may be
reorganized as that lands.