Skip to main content

Agent telemetry

The audit trail records what agents said to each other: every relayed message, every policy decision, every hop. It cannot see inside an agent. Agent telemetry closes that gap. An agent sends its own log lines and its steps (model calls, tool calls, delegations) to Swarmd over OpenTelemetry. Each one is stamped with the turn’s correlation id, and the trace view draws them on the same timeline as the audit events.
Preview, and opt-in at both ends. The platform operator turns the telemetry service on (see Agent telemetry for self-hosting), and each agent turns on what it sends with environment variables. Upgrading the SDK alone sends nothing.

Turn it on for a Google ADK agent

Agents built on swarmd-google-adk 0.4.0 or later need no code change. serve() reads these variables at startup:
.env
Restart the agent, send it a message, and open the trace from Observe › Trace explorer. Log lines appear between the audit events, and the steps appear as a tree under the timeline. The agent’s Logs tab also gets an Application logs view over a time range. A switch is on only for 1, true, yes or on. Anything else is off.
Leave SWARMD_TELEMETRY_REASONING off outside a test tenant. Redaction of reasoning content is not built yet. With it on, the full text of every prompt, completion and tool result is stored in ClickHouse for 30 days, readable by anyone with TELEMETRY:READ. The data stays in your cluster, but it is the most sensitive data your agents handle. Large call_llm steps can also go over the per-record size limit, and those steps are then dropped.

Where it goes

The SDK posts OTLP/HTTP JSON to <SWARMD_BASE_URL>/telemetry/v1/logs and /telemetry/v1/traces, using the agent’s existing client-credentials token. On a self-hosted install SWARMD_BASE_URL is your own gateway, so telemetry never leaves your network. Agents get the TELEMETRY:WRITE permission automatically, including agents registered before the upgrade.

What gets captured from Google ADK

ADK already opens spans for everything it does. Until something sets a real OpenTelemetry tracer provider, those spans go nowhere. Setting SWARMD_TELEMETRY_STEPS sets one, and each step is labelled with what kind of step it is: Spans from the A2A SDK’s own event-queue polling are dropped. They are noise, and in testing they made up most of a turn’s spans. Every log record and span carries swarmd.correlation_id, swarmd.context_id and swarmd.task_id. These are taken from the inbound request and A2A ids, and are what joins them to the audit trail.

Alongside Datadog, Grafana or your own collector

Telemetry is an extra destination. It never replaces yours.
  • Logs. Swarmd uses its own LoggerProvider and never takes over the global one, so your exporter receives each record exactly once. The SDK never changes logger levels, handlers or propagate.
  • Spans. If your process already set a global TracerProvider, Swarmd adds its processor to it and does not replace it. Your exporter sees spans unchanged, content included. Note that every span in that process is then also sent to Swarmd, except the A2A polling scope.
  • Reasoning content. ADK publishes prompts and completions as gen_ai.* log records through the global logger provider. If you already set one, Swarmd does not take it over. It logs a warning and skips reasoning capture. When it does capture, it sets OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true only if you have not set that variable yourself.
  • Noise. The loggers httpx, httpcore, urllib3, opentelemetry and the SDK’s own telemetry logger are never shipped. Each export also makes httpx log one line into your handlers, which you can filter.

It cannot take your agent down

  • Export runs off the request path in bounded batches (queue 2048, every 5 s, 10 s timeout).
  • On any error a batch is dropped and counted. Nothing is raised into your code.
  • A 401 gets one retry with a fresh token.
  • Pending records are flushed on shutdown.
  • If telemetry cannot start at all (no credentials, an incompatible OpenTelemetry pin, an unreachable endpoint), the agent logs one line and starts without it.
The SDK needs opentelemetry-sdk >=1.30,<2, which swarmd-sdk installs.

Other frameworks, and wiring it by hand

The environment variables are read only by swarmd-google-adk’s serve(). For LangChain or your own server, call the SDK directly once at startup:
Both read the endpoint and token from the runtime you already built with create_runtime(), and accept keyword overrides: LangChain has no built-in spans. You get your log lines, and spans only if you add a LangChain OpenTelemetry instrumentation yourself and pass step_kinds so they are labelled. swarmd-langchain does not yet stamp the A2A context and task ids, so its records join a trace by correlation id only.

Reading it back

Only the Tenant Administrator group has TELEMETRY:READ by default. Add it to other groups under Manage › Groups if your developers need it. When the trace shows no telemetry, it says why: you lack the permission, telemetry is not enabled on this deployment, or the agent sent nothing for that turn. It never shows an empty result as if the agent did nothing.

Next

Enable it on your cluster

ClickHouse, the two chart switches, sizing.

Environment reference

Every SWARMD_* variable the SDK reads.