Skip to main content

Agent telemetry

telemetry-service takes agents’ own OpenTelemetry log lines and spans and stores them in ClickHouse, keyed by the same correlation id as the audit trail. The trace view then shows what each agent logged, and the steps it took, on the same timeline as the policy decisions. The developer side is described in Agent telemetry (Python SDK). It is an opt-in preview in 0.4.x: off in the chart’s defaults, on in values-prismforce-local.yaml, and in every case nothing is collected until an agent opts in.
Nothing leaves your cluster. Agents post to your gateway, the gateway forwards to telemetry-service, and it writes to your ClickHouse. It calls no other service and nothing at Swarmd.

Before you start

External ClickHouse works the same way as for audit history: it must define a cluster named swarmd_cluster and the {shard}/{replica} macros, because the tables are ReplicatedMergeTree ... ON CLUSTER swarmd_cluster. The bundled ClickHouse already does this.

Turn it on

values-prismforce-local.yaml already has it on, together with ClickHouse. With that file, skip to step 3. For other values files:
1

Turn on ClickHouse and telemetry

2

Upgrade

audit, registry and telemetry wait for ClickHouse to answer before they start and then create their tables. The service creates its own Keycloak client and the TELEMETRY roles on first start. Every agent, including ones registered before the upgrade, is given TELEMETRY:WRITE.
3

Turn it on in your agents

Set SWARMD_TELEMETRY_ENABLED=true and, for steps, SWARMD_TELEMETRY_STEPS=true on each agent and restart it. See the SDK page.
Worker split and a brand-new ClickHouse. With the default worker split (services.*.worker.enabled: true), migrations run in pre-upgrade Jobs, and Helm runs those before it creates ClickHouse. If the same upgrade creates ClickHouse, those Jobs wait for a database that does not exist yet and the upgrade times out. On a split install, turn ClickHouse on in one upgrade and telemetry in the next. Installs without the worker split, such as values-prismforce-local.yaml, can do both at once.

The values

When telemetry is enabled the chart points the gateway at it. When it is not, platform-ui is told so and the trace view says “Agent telemetry is not enabled on this deployment” instead of reporting an error.

The endpoints

Agents reach it through the gateway, on the same host as the rest of the API:
  • JSON only. No protobuf and no gRPC. If you put your own OpenTelemetry Collector in front, use the otlphttp exporter with encoding: json and forward the agent’s bearer token.
  • Tenant comes from the token, never from the payload. A token without a tenant is refused.
  • Ingest always answers 200. Records it refuses (too large, or a span without a trace or span id) are counted in the OTLP partialSuccess reply. The rest of the batch is stored.

Who can read it

TELEMETRY:READ is granted to the Tenant Administrator group only. Editors and Viewers see “You do not have permission to read telemetry” on traces until you add it to their group under Manage › Groups. Agents can write but can never read.

Limits of the preview

  • No redaction of reasoning content. If an agent sets SWARMD_TELEMETRY_REASONING=true, full prompts and completions are stored. Keep it to test tenants.
  • No per-tenant quotas, sampling or rate limit. A noisy agent can fill the ClickHouse volume. Watch its disk. A full ClickHouse volume also stops audit history and schema migrations.
  • Retention is fixed at 30 days from ingest time. An agent’s clock cannot shorten or lengthen it.
  • Logs and spans only. No metrics.
  • Python only. Automatic capture is for Google ADK agents. LangChain agents can send log lines by calling the SDK, and the TypeScript SDK does not send telemetry.

Troubleshooting

services.telemetry.enabled is false in the values the release was rendered with. Check with helm -n swarmd get values swarmd.
Check the agent first. It logs [telemetry] shipping logs for <name> to <url> at startup when telemetry is on, and [telemetry] off or runtime not configured when it is not. The URL must be your gateway (SWARMD_BASE_URL). If the agent’s own logger.info() lines are missing but warnings arrive, raise LOG_LEVEL to info (SDK 0.4.0+).
A step whose attributes are larger than TELEMETRY_MAX_ATTRIBUTE_CHARS is rejected. This mostly happens to call_llm steps with reasoning content on. Raise the limit in services.telemetry.settings, or leave reasoning off.
It is waiting for ClickHouse. Either ClickHouse is being created in the same upgrade on a split install (see Turn it on), or an external ClickHouse lacks the swarmd_cluster definition.

Next

SDK side

The variables agents set, and what they capture.

Databases

ClickHouse alongside Postgres, and backups.