Skip to main content

Upgrading from 0.3 to 0.4

These notes cover 0.3.3 → 0.4.1 in one jump. 0.4.0 was an intermediate release. Everything in it is in 0.4.1, so if you are on 0.3.3, go straight to 0.4.1. The upgrade itself is the usual single helm upgrade (see Upgrades). Four things are different this time:
  1. Use the new values file. The 0.4 file turns on notification-service (governance notices and alerts depend on it), ClickHouse (monitors, alerts, LLM traffic figures and agent telemetry depend on it) and agent telemetry. With the 0.3.3 file the upgrade succeeds, but those features quietly do nothing.
  2. Tell the relay where your agents are. If your agents or MCP servers run in your cluster or on a private network, list them in services.relay.ssrfAllowedHosts. Details.
  3. Plan for 12 GiB. Three new pods (swarmd-notification, swarmd-clickhouse, swarmd-telemetry) take memory requests to about 6 GiB, plus a 20 GiB ClickHouse volume.
  4. A rollback to 0.3.3 needs a database restore. Details.

What is new

A redesigned platform UI

The whole UI has been rebuilt around a persistent sidebar grouped by what you are doing. The dark theme is the default; light is still available.
  • Ctrl/Cmd-K opens “Go to page” from anywhere.
  • Directories open beside the list. Selecting an agent, MCP server, gateway, provider, user, group, channel or sign-in provider opens its full workspace next to the directory. A direct link still opens the full page, and old ?inspect= links redirect.
  • One status filter, with counts, on every directory.
  • Compact mode: a toggle in the app bar for denser workspaces. It is remembered per browser.
  • Immutable proof (Govern) shows hash-chain status and lets you verify it.
  • Policy bindings (Govern) shows where each policy applies, including what is inherited, and lets you reorder.
  • Trace explorer has a split inspector with a single breakdown timeline.
  • Manage › Notifications has Channels and Destinations tabs. The old Overview tab is gone.
  • Edit and delete controls now follow the exact permission the API enforces. ADMIN on its own no longer shows WRITE or DELETE controls. A custom group that relied on that needs WRITE/DELETE granted explicitly; the API never accepted those calls anyway.

AI governance (EU AI Act)

Every agent, MCP server and LLM gateway now has a governance record: a risk classification, accountable people, linked compliance documents and a lifecycle. The AI register lists them across the tenant, and a serious incidents centre tracks Art. 73 reporting deadlines. Nothing changes for running agents when you upgrade. Every existing system starts Unclassified and in Draft, every gate is in Warn mode, and governance state never blocks relay traffic. Start with AI governance.

Art. 50 transparency

Per agent, you can now switch on an AI disclosure notice (“You are talking to an AI…”) and synthetic content marking. Both are off until you set them on the agent’s governance record. If you have built your own channel integration, read AI disclosure before you switch disclosure on. The notice arrives as its own message, and an integration that only shows the reply text will not display it.

Evaluation sets run for real

In 0.3.3 an evaluation run simulated its verdicts and could not fail. Now each evaluation sends a real message through the production policy chain and reads the verdicts back from the audit trail. That means:
  • Target agents really execute when you run a set.
  • A hop that could not complete is Error, not Failed.
  • Disabled evaluations show as Skipped.
  • Rate limits stand aside for evaluation runs, so runs do not use your tenant’s budget.
  • Only the first (entry) hop of a chain can be asserted reliably for now.
A run that waits on slow agents can be given longer with services.relay.settings.EVALUATION_DISPATCH_TIMEOUT_SECONDS (default 60). The PrismForce values file sets 150.

LLM gateway and provider traffic

Gateways and providers show trailing-24-hour requests, error rate and average latency, and every directory can sort by Most requests.
Gateway calls now record their outcome. Failed LLM calls count as errors: the dashboard’s call total includes them, and monitors on ERROR events will start firing on LLM provider failures (including a failover that ran out of providers). Review those monitors’ thresholds after upgrading.

Agent telemetry (preview)

Agents can now send their own OpenTelemetry log lines and steps to Swarmd, and the trace view shows them beside the audit events. The service is on in the 0.4 values file; each agent opts in with SWARMD_TELEMETRY_ENABLED and SWARMD_TELEMETRY_STEPS, so nothing is collected until you change an agent. It stays in your cluster and runs alongside an existing Datadog or Grafana exporter. See Agent telemetry.

Notifications

Monitor alerts, serious-incident deadline reminders and governance notices go out through notification channels (Manage › Notifications). A channel delivers by e-mail (needs SMTP), Slack or Microsoft Teams. A Webhook destination can be saved but is not delivered in 0.4. Monitors need ClickHouse, which the 0.4 values file turns on.

Smaller changes

  • Clearer relay refusals. A message to a deregistered or frozen agent now says so, instead of “Not authorized”. If anything of yours matches on that error text, check it.
  • Sign-in and sign-out. Sign-out always lands on the login page, a signed-in user who opens /login goes straight on, and SSO no longer needs a second attempt after an expired session. The login button cannot be submitted twice.
  • Python SDK 0.4. LOG_LEVEL now sets the root logger’s level, so your agents’ own logger.info() lines appear. Agent telemetry support has been added (off by default).

Security

  • Credentials at rest use your install’s key. The platform encrypts the credentials it stores: LLM provider keys, the API keys, bearer tokens and OAuth client secrets you give it for calling agents and MCP servers, MCP OAuth tokens, and agent webhook secrets. On 0.3.x installs from this chart they were encrypted with a built-in default key instead of the key generated for your install. From 0.4.1 everything new is encrypted with your install’s key, and existing values still decrypt. Re-save stored credentials after upgrading (see After the upgrade) so they are re-encrypted with it.
  • Service accounts have their own secrets. Each platform service’s Keycloak client now gets a secret generated for your install, replacing a built-in default. This is applied automatically during the upgrade.
  • Agents and services carry the audience the platform checks. On 0.3.x installs from this chart, agent tokens and some service-to-service tokens lacked the audience that registry and relay require, so agents calling the platform API were refused. Existing agents and channels are fixed automatically when registry starts.
  • Session isolation in the web UI. Concurrent token refreshes could, in a narrow race, hand one signed-in user another user’s session. The refresh lock is now per refresh token. Upgrade every platform-ui replica; we also recommend asking users to sign in again after the upgrade.
  • Image pulls are scoped. Licence-issued registry credentials can now pull only released images and charts. Installs from a release tarball are not affected.

Before you upgrade

1

Back up Postgres

For the single-Postgres layout used by values-prismforce-local.yaml:
Other layouts: see Databases › Backups.
2

Back up the two keys

Keep this file outside the cluster. A database restore is useless without the matching encryption-key. See Upgrades.
3

Bring your values file forward

Start from the 0.4 values-prismforce-local.yaml we sent with the chart, and re-apply anything you changed in your 0.3.3 copy. Compare them:
The 0.4 file turns on clickhouse (with a 20 GiB volume), services.notification and services.telemetry, sets services.relay.settings.EVALUATION_DISPATCH_TIMEOUT_SECONDS: "150", and adds services.relay.ssrfAllowedHosts.
4

List where your agents and MCP servers run

The relay refuses private and in-cluster addresses unless the host is listed, so messages to an agent running in your cluster fail with Blocked request to private/internal address. Set services.relay.ssrfAllowedHosts to the hosts, or a *. pattern for a whole namespace:
Agents and MCP servers on public addresses need no entry. A pattern never opens loopback or cloud-metadata addresses.
5

Check headroom

Three new pods (notification, ClickHouse, telemetry): plan for about 6 GiB of memory requests in total, a 12 GiB node, and 20 GiB more disk for ClickHouse.
6

Preview

You should see:
  • new Deployments and Services for swarmd-notification, swarmd-clickhouse and swarmd-telemetry, a ClickHouse PVC and ConfigMap;
  • new keys in swarmd-generated-credentials (encryption-keys and one service-client-secret-* per service). Existing keys are unchanged;
  • new image tags on every Swarmd workload, and new environment variables on the services;
  • a wait-for-clickhouse init container on audit, registry and telemetry.
No selectors change.

Upgrade

With the PrismForce values file every service runs as a single container, so there are no migration Jobs. Each service runs its own migrations as it starts. Expect pods to sit in Init while they wait for Keycloak (and audit, registry and telemetry for ClickHouse), then take a little longer than usual on their first start. Zero restarts is still the expected outcome.

Schema changes in this release

All of them are quick at normal data sizes. Nothing is backfilled. Keycloak needs no manual step. Service clients and roles are updated by the services themselves at startup, and your realm and its customisations are left alone.

After the upgrade

1

Check the install

2

Exercise a real path

Sign in, open Your estate › Agents and send a message to an agent from Conversation. Then open the trace from Observe › Trace explorer. Until the agent opts in to telemetry, the trace notes that the agent sent no log lines, which is expected.
3

Open the AI register

Govern › AI register should list every agent, MCP server and gateway as Unclassified with an unconfirmed owner. That is the correct starting point.
4

Re-save stored credentials

Re-enter each LLM provider’s key (Your estate › LLM providers), the authentication settings of agents and MCP servers that call out with an API key, bearer token or OAuth client secret, and reconnect MCP servers that use OAuth. Saving re-encrypts them with your install’s own key. Values you don’t re-save keep working. See Security.
5

Ask users to sign in again

6

Review ERROR monitors


Rolling back

helm rollback to 0.3.3 restores the 0.3.3 images, but not the schema. Relay migration V053 dropped the verdicts_simulated column, which the 0.3.3 relay still maps, so after a plain rollback every evaluation-run read and write in 0.3.3 fails. Everything else keeps working. To roll back fully:
Anything created after the upgrade is lost by the restore, which is why it is worth upgrading in a quiet window and checking the result straight away. The restore matters for credentials too: values saved after the upgrade are encrypted with your install’s key, which 0.3.3 does not use.

Next

AI governance

What the AI register is and where to start.

Agent telemetry

What agents send, and how to read it.