> ## Documentation Index
> Fetch the complete documentation index at: https://docs.swarmd.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent telemetry

> Turn on telemetry-service so agents can send their own OpenTelemetry logs and steps, stored in your ClickHouse beside the audit trail.

# Agent telemetry

`telemetry-service` takes agents' own OpenTelemetry log lines and spans and
stores them in ClickHouse, keyed by the same correlation id as the audit trail.
The trace view then shows what each agent logged, and the steps it took, on the
same timeline as the policy decisions. The developer side is described in
[Agent telemetry (Python SDK)](/sdks/python/telemetry).

It is an **opt-in preview** in 0.4.x: off in the chart's defaults, on in
`values-prismforce-local.yaml`, and in every case nothing is collected until
an agent opts in.

<Note>
  **Nothing leaves your cluster.** Agents post to your gateway, the gateway
  forwards to `telemetry-service`, and it writes to your ClickHouse. It calls no
  other service and nothing at Swarmd.
</Note>

***

## Before you start

| Need | Why |
| - | - |
| ClickHouse (`clickhouse.enabled: true`) | It is the only store. The chart refuses to render with telemetry capture on and ClickHouse off. |
| About 2 GiB more memory | ClickHouse (512 Mi–2 Gi) plus one telemetry pod (\~320 Mi requested). |
| Disk for 30 days of logs | Retention is fixed at 30 days from ingest. Size the ClickHouse volume for your agents' log volume, not just audit history. |
| Agents on `swarmd-google-adk` / `swarmd-sdk` 0.4.0+ | Older SDKs do not send telemetry. |

**External ClickHouse** works the same way as for audit history: it must
define a cluster named `swarmd_cluster` and the `{shard}`/`{replica}` macros,
because the tables are `ReplicatedMergeTree ... ON CLUSTER swarmd_cluster`. The
bundled ClickHouse already does this.

***

## Turn it on

**`values-prismforce-local.yaml` already has it on**, together with
ClickHouse. With that file, skip to [step 3](#turn-it-on-in-your-agents).

For other values files:

<Steps>
  <Step title="Turn on ClickHouse and telemetry">
    ```yaml theme={null}
    clickhouse:
      enabled: true
      storage:
        size: 20Gi
    services:
      telemetry:
        enabled: true          # deploy the service
        captureEnabled: true   # accept and serve telemetry
    ```
  </Step>

  <Step title="Upgrade">
    ```bash theme={null}
    helm upgrade swarmd $SWARMD_CHART -n swarmd -f my-values.yaml --wait --timeout 15m
    ```

    audit, registry and telemetry wait for ClickHouse to answer before they
    start and then create their tables. The service creates its own Keycloak
    client and the `TELEMETRY` roles on first start. Every agent, including
    ones registered before the upgrade, is given `TELEMETRY:WRITE`.
  </Step>

  <Step title="Turn it on in your agents">
    Set `SWARMD_TELEMETRY_ENABLED=true` and, for steps,
    `SWARMD_TELEMETRY_STEPS=true` on each agent and restart it. See the
    [SDK page](/sdks/python/telemetry).
  </Step>
</Steps>

<Warning>
  **Worker split and a brand-new ClickHouse.** With the default worker split
  (`services.*.worker.enabled: true`), migrations run in `pre-upgrade` Jobs, and
  Helm runs those before it creates ClickHouse. If the same upgrade creates
  ClickHouse, those Jobs wait for a database that does not exist yet and the
  upgrade times out. On a split install, turn ClickHouse on in one upgrade and
  telemetry in the next. Installs without the worker split, such as
  `values-prismforce-local.yaml`, can do both at once.
</Warning>

***

## The values

| Key | Default | Meaning |
| - | - | - |
| `services.telemetry.enabled` | `false` | Deploy the service at all. |
| `services.telemetry.captureEnabled` | `false` | Register the ingest and read endpoints. Set it back to `false` to stop accepting telemetry without uninstalling anything; the pods roll and keep running. |
| `services.telemetry.worker.enabled` | `true` | Core + worker + migrate Job, like the other services. Set `false` on single-node installs. |
| `services.telemetry.resources` | chart default | Requests/limits for the pod. |
| `services.telemetry.settings.TELEMETRY_MAX_BODY_CHARS` | `65536` | Largest log body accepted. Larger records are **rejected, not truncated**. |
| `services.telemetry.settings.TELEMETRY_MAX_ATTRIBUTE_CHARS` | `16384` | Largest serialised attribute map per record or span. |
| `services.telemetry.settings.TELEMETRY_MAX_RECORDS_PER_REQUEST` | `10000` | Records past this in one request are rejected. |

When telemetry is enabled the chart points the gateway at it. When it is not,
platform-ui is told so and the trace view says "Agent telemetry is not enabled
on this deployment" instead of reporting an error.

***

## The endpoints

Agents reach it through the gateway, on the same host as the rest of the API:

| Method and path | Auth | What |
| - | - | - |
| `POST /telemetry/v1/logs` | Agent token, `TELEMETRY:WRITE` | OTLP/HTTP **JSON** log export |
| `POST /telemetry/v1/traces` | Agent token, `TELEMETRY:WRITE` | OTLP/HTTP **JSON** trace export |
| `GET /telemetry/v1/logs` | User token, `TELEMETRY:READ` | Read back log lines |
| `GET /telemetry/v1/spans` | User token, `TELEMETRY:READ` | Read back steps |

* **JSON only.** No protobuf and no gRPC. If you put your own OpenTelemetry
  Collector in front, use the `otlphttp` exporter with `encoding: json` and
  forward the agent's bearer token.
* **Tenant comes from the token,** never from the payload. A token without a
  tenant is refused.
* Ingest always answers `200`. Records it refuses (too large, or a span without
  a trace or span id) are counted in the OTLP `partialSuccess` reply. The rest
  of the batch is stored.

***

## Who can read it

`TELEMETRY:READ` is granted to the **Tenant Administrator** group only. Editors
and Viewers see "You do not have permission to read telemetry" on traces until
you add it to their group under **Manage › Groups**. Agents can write but can
never read.

***

## Limits of the preview

* **No redaction of reasoning content.** If an agent sets
  `SWARMD_TELEMETRY_REASONING=true`, full prompts and completions are stored.
  Keep it to test tenants.
* **No per-tenant quotas, sampling or rate limit.** A noisy agent can fill the
  ClickHouse volume. Watch its disk. A full ClickHouse volume also stops audit
  history and schema migrations.
* **Retention is fixed at 30 days** from ingest time. An agent's clock cannot
  shorten or lengthen it.
* **Logs and spans only.** No metrics.
* **Python only.** Automatic capture is for Google ADK agents. LangChain agents
  can send log lines by calling the SDK, and the TypeScript SDK does not send
  telemetry.

***

## Troubleshooting

<AccordionGroup>
  <Accordion title="Traces say telemetry is not enabled">
    `services.telemetry.enabled` is false in the values the release was
    rendered with. Check with `helm -n swarmd get values swarmd`.
  </Accordion>

  <Accordion title="Telemetry is enabled but nothing arrives">
    Check the agent first. It logs `[telemetry] shipping logs for <name> to <url>`
    at startup when telemetry is on, and `[telemetry] off` or `runtime not configured`
    when it is not. The URL must be your gateway (`SWARMD_BASE_URL`). If the
    agent's own `logger.info()` lines are missing but warnings arrive, raise
    `LOG_LEVEL` to `info` (SDK 0.4.0+).
  </Accordion>

  <Accordion title="Some steps are missing">
    A step whose attributes are larger than `TELEMETRY_MAX_ATTRIBUTE_CHARS` is
    rejected. This mostly happens to `call_llm` steps with reasoning content on.
    Raise the limit in `services.telemetry.settings`, or leave reasoning off.
  </Accordion>

  <Accordion title="The telemetry migrate Job never completes">
    It is waiting for ClickHouse. Either ClickHouse is being created in the
    same upgrade on a split install (see [Turn it on](#turn-it-on)), or an
    external ClickHouse lacks the `swarmd_cluster` definition.
  </Accordion>
</AccordionGroup>

***

## Next

<CardGroup cols={2}>
  <Card title="SDK side" icon="python" href="/sdks/python/telemetry">
    The variables agents set, and what they capture.
  </Card>

  <Card title="Databases" icon="database" href="/self-hosting/databases">
    ClickHouse alongside Postgres, and backups.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.