Agent telemetry
telemetry-service takes agents’ own OpenTelemetry log lines and spans and
stores them in ClickHouse, keyed by the same correlation id as the audit trail.
The trace view then shows what each agent logged, and the steps it took, on the
same timeline as the policy decisions. The developer side is described in
Agent telemetry (Python SDK).
It is an opt-in preview in 0.4.x: off in the chart’s defaults, on in
values-prismforce-local.yaml, and in every case nothing is collected until
an agent opts in.
Nothing leaves your cluster. Agents post to your gateway, the gateway
forwards to
telemetry-service, and it writes to your ClickHouse. It calls no
other service and nothing at Swarmd.Before you start
External ClickHouse works the same way as for audit history: it must
define a cluster named
swarmd_cluster and the {shard}/{replica} macros,
because the tables are ReplicatedMergeTree ... ON CLUSTER swarmd_cluster. The
bundled ClickHouse already does this.
Turn it on
values-prismforce-local.yaml already has it on, together with
ClickHouse. With that file, skip to step 3.
For other values files:
1
Turn on ClickHouse and telemetry
2
Upgrade
TELEMETRY roles on first start. Every agent, including
ones registered before the upgrade, is given TELEMETRY:WRITE.3
Turn it on in your agents
Set
SWARMD_TELEMETRY_ENABLED=true and, for steps,
SWARMD_TELEMETRY_STEPS=true on each agent and restart it. See the
SDK page.The values
When telemetry is enabled the chart points the gateway at it. When it is not,
platform-ui is told so and the trace view says “Agent telemetry is not enabled
on this deployment” instead of reporting an error.
The endpoints
Agents reach it through the gateway, on the same host as the rest of the API:- JSON only. No protobuf and no gRPC. If you put your own OpenTelemetry
Collector in front, use the
otlphttpexporter withencoding: jsonand forward the agent’s bearer token. - Tenant comes from the token, never from the payload. A token without a tenant is refused.
- Ingest always answers
200. Records it refuses (too large, or a span without a trace or span id) are counted in the OTLPpartialSuccessreply. The rest of the batch is stored.
Who can read it
TELEMETRY:READ is granted to the Tenant Administrator group only. Editors
and Viewers see “You do not have permission to read telemetry” on traces until
you add it to their group under Manage › Groups. Agents can write but can
never read.
Limits of the preview
- No redaction of reasoning content. If an agent sets
SWARMD_TELEMETRY_REASONING=true, full prompts and completions are stored. Keep it to test tenants. - No per-tenant quotas, sampling or rate limit. A noisy agent can fill the ClickHouse volume. Watch its disk. A full ClickHouse volume also stops audit history and schema migrations.
- Retention is fixed at 30 days from ingest time. An agent’s clock cannot shorten or lengthen it.
- Logs and spans only. No metrics.
- Python only. Automatic capture is for Google ADK agents. LangChain agents can send log lines by calling the SDK, and the TypeScript SDK does not send telemetry.
Troubleshooting
Traces say telemetry is not enabled
Traces say telemetry is not enabled
services.telemetry.enabled is false in the values the release was
rendered with. Check with helm -n swarmd get values swarmd.Telemetry is enabled but nothing arrives
Telemetry is enabled but nothing arrives
Check the agent first. It logs
[telemetry] shipping logs for <name> to <url>
at startup when telemetry is on, and [telemetry] off or runtime not configured
when it is not. The URL must be your gateway (SWARMD_BASE_URL). If the
agent’s own logger.info() lines are missing but warnings arrive, raise
LOG_LEVEL to info (SDK 0.4.0+).Some steps are missing
Some steps are missing
A step whose attributes are larger than
TELEMETRY_MAX_ATTRIBUTE_CHARS is
rejected. This mostly happens to call_llm steps with reasoning content on.
Raise the limit in services.telemetry.settings, or leave reasoning off.The telemetry migrate Job never completes
The telemetry migrate Job never completes
It is waiting for ClickHouse. Either ClickHouse is being created in the
same upgrade on a split install (see Turn it on), or an
external ClickHouse lacks the
swarmd_cluster definition.Next
SDK side
The variables agents set, and what they capture.
Databases
ClickHouse alongside Postgres, and backups.
