> ## Documentation Index
> Fetch the complete documentation index at: https://docs.swarmd.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Upgrading to 0.4

> Release notes and the upgrade checklist for moving a 0.3.x install to 0.4.1: what is new, what behaves differently, and the one migration that makes a rollback need a restore.

# Upgrading from 0.3 to 0.4

These notes cover **0.3.3 → 0.4.1** in one jump. 0.4.0 was an intermediate
release. Everything in it is in 0.4.1, so if you are on 0.3.3, go straight to
0.4.1.

The upgrade itself is the usual single `helm upgrade` (see
[Upgrades](/self-hosting/upgrades)). Four things are different this time:

1. **Use the new values file.** The 0.4 file turns on `notification-service`
   (governance notices and alerts depend on it), ClickHouse (monitors, alerts,
   LLM traffic figures and agent telemetry depend on it) and agent telemetry.
   With the 0.3.3 file the upgrade succeeds, but those features quietly do
   nothing.
2. **Tell the relay where your agents are.** If your agents or MCP servers run
   in your cluster or on a private network, list them in
   `services.relay.ssrfAllowedHosts`. [Details](#before-you-upgrade).
3. **Plan for 12 GiB.** Three new pods (`swarmd-notification`,
   `swarmd-clickhouse`, `swarmd-telemetry`) take memory requests to about
   6 GiB, plus a 20 GiB ClickHouse volume.
4. **A rollback to 0.3.3 needs a database restore.** [Details](#rolling-back).

***

## What is new

### A redesigned platform UI

The whole UI has been rebuilt around a persistent sidebar grouped by what you
are doing. The dark theme is the default; light is still available.

| Was (0.3.3) | Now |
| - | - |
| Dashboards | Observe › Overview |
| Activity › Logs / Traces / Tasks | Observe › Logs & events / Trace explorer / Agent tasks |
| Activity › Monitors / Alerts | Observe › Monitors / Alerts |
| Policies, Approvals, Evaluations | Govern › Policies / Approvals / Evaluation sets |
| — | Govern › Policy bindings, AI register, Incidents, Immutable proof, Compliance reports |
| Agents › Inventory, Network, MCP Servers, LLMs | Your estate › Agents / Network / MCP servers / LLM gateways / LLM providers / Channels |
| Agents › Marketplace, MCP Servers › Marketplace | Discover › Agentic Marketplace |
| Settings › … | Manage › Users / Groups / Sign-in providers / Notifications / Integrations / Governance policy / Emergency controls |

* **Ctrl/Cmd-K** opens "Go to page" from anywhere.
* **Directories open beside the list.** Selecting an agent, MCP server,
  gateway, provider, user, group, channel or sign-in provider opens its full
  workspace next to the directory. A direct link still opens the full page,
  and old `?inspect=` links redirect.
* **One status filter**, with counts, on every directory.
* **Compact mode**: a toggle in the app bar for denser workspaces. It is
  remembered per browser.
* **Immutable proof** (Govern) shows hash-chain status and lets you verify it.
* **Policy bindings** (Govern) shows where each policy applies, including what
  is inherited, and lets you reorder.
* **Trace explorer** has a split inspector with a single breakdown timeline.
* **Manage › Notifications** has Channels and Destinations tabs. The old
  Overview tab is gone.
* Edit and delete controls now follow the exact permission the API enforces.
  `ADMIN` on its own no longer shows `WRITE` or `DELETE` controls. A custom
  group that relied on that needs `WRITE`/`DELETE` granted explicitly; the API
  never accepted those calls anyway.

### AI governance (EU AI Act)

Every agent, MCP server and LLM gateway now has a governance record: a risk
classification, accountable people, linked compliance documents and a
lifecycle. The **AI register** lists them across the tenant, and a **serious
incidents** centre tracks Art. 73 reporting deadlines.

**Nothing changes for running agents when you upgrade.** Every existing system
starts *Unclassified* and in *Draft*, every gate is in *Warn* mode, and
governance state never blocks relay traffic. Start with
[AI governance](/governance/overview).

### Art. 50 transparency

Per agent, you can now switch on an **AI disclosure** notice ("You are talking
to an AI…") and **synthetic content marking**. Both are off until you set them
on the agent's governance record. If you have built your own channel
integration, read [AI disclosure](/sdks/typescript/ai-disclosure) before you
switch disclosure on. The notice arrives as its own message, and an
integration that only shows the reply text will not display it.

### Evaluation sets run for real

In 0.3.3 an evaluation run simulated its verdicts and could not fail. Now each
evaluation sends a real message through the production policy chain and reads
the verdicts back from the audit trail. That means:

* **Target agents really execute** when you run a set.
* A hop that could not complete is **Error**, not **Failed**.
* Disabled evaluations show as **Skipped**.
* Rate limits stand aside for evaluation runs, so runs do not use your tenant's
  budget.
* Only the first (entry) hop of a chain can be asserted reliably for now.

A run that waits on slow agents can be given longer with
`services.relay.settings.EVALUATION_DISPATCH_TIMEOUT_SECONDS` (default 60).
The PrismForce values file sets 150.

### LLM gateway and provider traffic

Gateways and providers show trailing-24-hour requests, error rate and average
latency, and every directory can sort by **Most requests**.

<Warning>
  Gateway calls now record their outcome. **Failed LLM calls count as errors**:
  the dashboard's call total includes them, and monitors on `ERROR` events will
  start firing on LLM provider failures (including a failover that ran out of
  providers). Review those monitors' thresholds after upgrading.
</Warning>

### Agent telemetry (preview)

Agents can now send their own OpenTelemetry log lines and steps to Swarmd, and
the trace view shows them beside the audit events. The service is on in the
0.4 values file; each agent opts in with `SWARMD_TELEMETRY_ENABLED` and
`SWARMD_TELEMETRY_STEPS`, so nothing is collected until you change an agent.
It stays in your cluster and runs alongside an existing Datadog or Grafana
exporter. See [Agent telemetry](/self-hosting/agent-telemetry).

### Notifications

Monitor alerts, serious-incident deadline reminders and governance notices go
out through notification channels (**Manage › Notifications**). A channel
delivers by **e-mail** (needs [SMTP](/self-hosting/configuration#smtp)),
**Slack** or **Microsoft Teams**. A *Webhook* destination can be saved but is
not delivered in 0.4. Monitors need ClickHouse, which the 0.4 values file
turns on.

### Smaller changes

* **Clearer relay refusals.** A message to a deregistered or frozen agent now
  says so, instead of "Not authorized". If anything of yours matches on that
  error text, check it.
* **Sign-in and sign-out.** Sign-out always lands on the login page, a
  signed-in user who opens `/login` goes straight on, and SSO no longer needs a
  second attempt after an expired session. The login button cannot be
  submitted twice.
* **Python SDK 0.4.** `LOG_LEVEL` now sets the root logger's level, so your
  agents' own `logger.info()` lines appear. Agent telemetry support has been
  added (off by default).

### Security

* **Credentials at rest use your install's key.** The platform encrypts the
  credentials it stores: LLM provider keys, the API keys, bearer tokens and
  OAuth client secrets you give it for calling agents and MCP servers, MCP
  OAuth tokens, and agent webhook secrets. On 0.3.x installs from this chart they were encrypted with
  a built-in default key instead of the key generated for your install. From
  0.4.1 everything new is encrypted with your install's key, and existing
  values still decrypt. **Re-save stored credentials after upgrading** (see
  [After the upgrade](#after-the-upgrade)) so they are re-encrypted with it.
* **Service accounts have their own secrets.** Each platform service's
  Keycloak client now gets a secret generated for your install, replacing a
  built-in default. This is applied automatically during the upgrade.
* **Agents and services carry the audience the platform checks.** On 0.3.x
  installs from this chart, agent tokens and some service-to-service tokens
  lacked the audience that registry and relay require, so agents calling the
  platform API were refused. Existing agents and channels are fixed
  automatically when registry starts.
* **Session isolation in the web UI.** Concurrent token refreshes could, in a
  narrow race, hand one signed-in user another user's session. The refresh
  lock is now per refresh token. Upgrade every platform-ui replica; we also
  recommend asking users to sign in again after the upgrade.
* **Image pulls are scoped.** Licence-issued registry credentials can now pull
  only released images and charts. Installs from a release tarball are not
  affected.

***

## Before you upgrade

<Steps>
  <Step title="Back up Postgres">
    For the single-Postgres layout used by `values-prismforce-local.yaml`:

    ```bash theme={null}
    NS=swarmd
    PGUSER=$(kubectl -n $NS get secret swarmd-generated-credentials -o jsonpath='{.data.postgres-username}' | base64 -d)
    PGPASS=$(kubectl -n $NS get secret swarmd-generated-credentials -o jsonpath='{.data.postgres-password}' | base64 -d)

    for db in swarmd keycloak; do
      kubectl -n $NS exec deploy/swarmd-postgres -- \
        env PGPASSWORD="$PGPASS" pg_dump -U "$PGUSER" -Fc "$db" > "$db-0.3.3.dump"
    done
    ls -lh *.dump   # both files should be non-empty
    ```

    Other layouts: see [Databases › Backups](/self-hosting/databases).
  </Step>

  <Step title="Back up the two keys">
    ```bash theme={null}
    kubectl -n swarmd get secret swarmd-generated-credentials -o json \
      | jq '{"encryption-key": .data["encryption-key"], "audit-hashchain-signing-key": .data["audit-hashchain-signing-key"]}' \
      > swarmd-keys-0.3.3.json
    ```

    Keep this file outside the cluster. A database restore is useless without
    the matching `encryption-key`. See [Upgrades](/self-hosting/upgrades#the-one-that-bricks-installs).
  </Step>

  <Step title="Bring your values file forward">
    Start from the 0.4 `values-prismforce-local.yaml` we sent with the chart,
    and re-apply anything you changed in your 0.3.3 copy. Compare them:

    ```bash theme={null}
    diff my-0.3.3-values.yaml values-prismforce-local.yaml
    ```

    The 0.4 file turns on `clickhouse` (with a 20 GiB volume),
    `services.notification` and `services.telemetry`, sets
    `services.relay.settings.EVALUATION_DISPATCH_TIMEOUT_SECONDS: "150"`, and
    adds `services.relay.ssrfAllowedHosts`.
  </Step>

  <Step title="List where your agents and MCP servers run">
    The relay refuses private and in-cluster addresses unless the host is
    listed, so messages to an agent running in your cluster fail with
    *Blocked request to private/internal address*. Set
    `services.relay.ssrfAllowedHosts` to the hosts, or a `*.` pattern for a
    whole namespace:

    ```yaml theme={null}
    services:
      relay:
        ssrfAllowedHosts: "*.prismforce-agents.svc.cluster.local,crm-mcp.corp.internal"
    ```

    Agents and MCP servers on public addresses need no entry. A pattern never
    opens loopback or cloud-metadata addresses.
  </Step>

  <Step title="Check headroom">
    Three new pods (notification, ClickHouse, telemetry): plan for about
    6 GiB of memory requests in total, a 12 GiB node, and 20 GiB more disk for
    ClickHouse.
  </Step>

  <Step title="Preview">
    ```bash theme={null}
    helm diff upgrade swarmd ./swarmd-0.4.1.tgz -n swarmd -f values-prismforce-local.yaml
    ```

    You should see:

    * new Deployments and Services for `swarmd-notification`,
      `swarmd-clickhouse` and `swarmd-telemetry`, a ClickHouse PVC and
      ConfigMap;
    * new keys in `swarmd-generated-credentials` (`encryption-keys` and one
      `service-client-secret-*` per service). Existing keys are unchanged;
    * new image tags on every Swarmd workload, and new environment variables
      on the services;
    * a `wait-for-clickhouse` init container on audit, registry and telemetry.

    No selectors change.
  </Step>
</Steps>

***

## Upgrade

```bash theme={null}
helm upgrade swarmd ./swarmd-0.4.1.tgz -n swarmd \
  -f values-prismforce-local.yaml \
  --wait --timeout 15m
```

With the PrismForce values file every service runs as a single container, so
**there are no migration Jobs**. Each service runs its own migrations as it
starts. Expect pods to sit in `Init` while they wait for Keycloak (and audit,
registry and telemetry for ClickHouse), then take a little longer than usual on
their first start. Zero restarts is still the expected outcome.

### Schema changes in this release

| Service | Migrations | Effect |
| - | - | - |
| registry | V058–V065 | New governance tables and views (classification, assignments, documents, lifecycle, retirement, incidents, notification outbox). Additive. |
| relay | V053, V054 | Evaluation runs record their execution mode. **V053 drops `verdicts_simulated`.** V054 adds a nullable column to HITL resolutions. |
| audit | V027, V028 | Tenant retention settings and checkpoint requests. Additive. |
| notification | V001–V005 | New `notification` schema, created on first start. |
| ClickHouse | audit, registry, telemetry | ClickHouse is new to this install, so audit and registry create their tables, and telemetry creates `telemetry_logs` and `telemetry_spans`. Audit history in ClickHouse starts from the upgrade; Postgres keeps the full, authoritative record. |

All of them are quick at normal data sizes. Nothing is backfilled.

Keycloak needs no manual step. Service clients and roles are updated by the
services themselves at startup, and your realm and its customisations are left
alone.

***

## After the upgrade

<Steps>
  <Step title="Check the install">
    ```bash theme={null}
    helm -n swarmd list                 # chart swarmd-0.4.1
    kubectl -n swarmd get pods          # all Running, 1/1, RESTARTS 0 — including notification, clickhouse, telemetry
    kubectl -n swarmd logs deploy/swarmd-registry | grep -i 'Successfully applied\|up to date'
    ```
  </Step>

  <Step title="Exercise a real path">
    Sign in, open **Your estate › Agents** and send a message to an agent from
    **Conversation**. Then open the trace from **Observe › Trace explorer**.
    Until the agent opts in to telemetry, the trace notes that the agent sent
    no log lines, which is expected.
  </Step>

  <Step title="Open the AI register">
    **Govern › AI register** should list every agent, MCP server and gateway
    as *Unclassified* with an unconfirmed owner. That is the correct starting
    point.
  </Step>

  <Step title="Re-save stored credentials">
    Re-enter each LLM provider's key (**Your estate › LLM providers**), the
    authentication settings of agents and MCP servers that call out with an
    API key, bearer token or OAuth client secret, and reconnect MCP servers
    that use OAuth. Saving re-encrypts them with your install's own key.
    Values you don't re-save keep working. See [Security](#security).
  </Step>

  <Step title="Ask users to sign in again">
    See [Security](#security).
  </Step>

  <Step title="Review ERROR monitors">
    See [LLM gateway and provider traffic](#llm-gateway-and-provider-traffic).
  </Step>
</Steps>

***

## Rolling back

`helm rollback` to 0.3.3 restores the 0.3.3 images, but **not the schema**.
Relay migration V053 dropped the `verdicts_simulated` column, which the 0.3.3
relay still maps, so after a plain rollback every evaluation-run read and write
in 0.3.3 fails. Everything else keeps working.

To roll back fully:

```bash theme={null}
helm rollback swarmd -n swarmd --wait --timeout 15m
# then restore the dumps you took before upgrading
kubectl -n swarmd scale deploy -l app.kubernetes.io/instance=swarmd --replicas=0
kubectl -n swarmd scale deploy/swarmd-postgres --replicas=1
for db in swarmd keycloak; do
  kubectl -n swarmd exec -i deploy/swarmd-postgres -- \
    env PGPASSWORD="$PGPASS" pg_restore -U "$PGUSER" --clean --if-exists -d "$db" < "$db-0.3.3.dump"
done
kubectl -n swarmd scale deploy -l app.kubernetes.io/instance=swarmd --replicas=1
```

Anything created after the upgrade is lost by the restore, which is why it is
worth upgrading in a quiet window and checking the result straight away. The
restore matters for credentials too: values saved after the upgrade are
encrypted with your install's key, which 0.3.3 does not use.

***

## Next

<CardGroup cols={2}>
  <Card title="AI governance" icon="scale-balanced" href="/governance/overview">
    What the AI register is and where to start.
  </Card>

  <Card title="Agent telemetry" icon="wave-pulse" href="/self-hosting/agent-telemetry">
    What agents send, and how to read it.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.