Agent telemetry
The audit trail records what agents said to each other: every relayed message, every policy decision, every hop. It cannot see inside an agent. Agent telemetry closes that gap. An agent sends its own log lines and its steps (model calls, tool calls, delegations) to Swarmd over OpenTelemetry. Each one is stamped with the turn’s correlation id, and the trace view draws them on the same timeline as the audit events.Preview, and opt-in at both ends. The platform operator turns the
telemetry service on (see Agent telemetry for
self-hosting), and each agent turns on what it
sends with environment variables. Upgrading the SDK alone sends nothing.
Turn it on for a Google ADK agent
Agents built onswarmd-google-adk 0.4.0 or later need no code change.
serve() reads these variables at startup:
.env
A switch is on only for
1, true, yes or on. Anything else is off.
Where it goes
The SDK posts OTLP/HTTP JSON to<SWARMD_BASE_URL>/telemetry/v1/logs and
/telemetry/v1/traces, using the agent’s existing client-credentials token.
On a self-hosted install SWARMD_BASE_URL is your own gateway, so
telemetry never leaves your network. Agents get the TELEMETRY:WRITE
permission automatically, including agents registered before the upgrade.
What gets captured from Google ADK
ADK already opens spans for everything it does. Until something sets a real OpenTelemetry tracer provider, those spans go nowhere. SettingSWARMD_TELEMETRY_STEPS sets one, and each step is labelled with what kind of
step it is:
Spans from the A2A SDK’s own event-queue polling are dropped. They are noise,
and in testing they made up most of a turn’s spans.
Every log record and span carries
swarmd.correlation_id,
swarmd.context_id and swarmd.task_id. These are taken from the inbound
request and A2A ids, and are what joins them to the audit trail.
Alongside Datadog, Grafana or your own collector
Telemetry is an extra destination. It never replaces yours.- Logs. Swarmd uses its own
LoggerProviderand never takes over the global one, so your exporter receives each record exactly once. The SDK never changes logger levels, handlers orpropagate. - Spans. If your process already set a global
TracerProvider, Swarmd adds its processor to it and does not replace it. Your exporter sees spans unchanged, content included. Note that every span in that process is then also sent to Swarmd, except the A2A polling scope. - Reasoning content. ADK publishes prompts and completions as
gen_ai.*log records through the global logger provider. If you already set one, Swarmd does not take it over. It logs a warning and skips reasoning capture. When it does capture, it setsOTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=trueonly if you have not set that variable yourself. - Noise. The loggers
httpx,httpcore,urllib3,opentelemetryand the SDK’s own telemetry logger are never shipped. Each export also makeshttpxlog one line into your handlers, which you can filter.
It cannot take your agent down
- Export runs off the request path in bounded batches (queue 2048, every 5 s, 10 s timeout).
- On any error a batch is dropped and counted. Nothing is raised into your code.
- A
401gets one retry with a fresh token. - Pending records are flushed on shutdown.
- If telemetry cannot start at all (no credentials, an incompatible OpenTelemetry pin, an unreachable endpoint), the agent logs one line and starts without it.
opentelemetry-sdk >=1.30,<2, which swarmd-sdk installs.
Other frameworks, and wiring it by hand
The environment variables are read only byswarmd-google-adk’s serve().
For LangChain or your own server, call the SDK directly once at startup:
create_runtime(), and accept keyword overrides:
LangChain has no built-in spans. You get your log lines, and spans only if you
add a LangChain OpenTelemetry instrumentation yourself and pass
step_kinds
so they are labelled. swarmd-langchain does not yet stamp the A2A context
and task ids, so its records join a trace by correlation id only.
Reading it back
Only the Tenant Administrator group has
TELEMETRY:READ by default. Add
it to other groups under Manage › Groups if your developers need it.
When the trace shows no telemetry, it says why: you lack the permission,
telemetry is not enabled on this deployment, or the agent sent nothing for
that turn. It never shows an empty result as if the agent did nothing.
Next
Enable it on your cluster
ClickHouse, the two chart switches, sizing.
Environment reference
Every
SWARMD_* variable the SDK reads.