> ## Documentation Index
> Fetch the complete documentation index at: https://docs.swarmd.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluation sets

> Prove your policies do what you think: send real test messages through the production policy chain and check what happened at each hop.

# Evaluation sets

An **evaluation set** (an *evaluation pack* in the API) tests one promise you
make about your agents, such as "agents must not leak PII" or "consequential
decisions need a human". Each **evaluation** in the set sends one test message
along one route through your estate. It then checks what the platform actually
did at each hop: which action it took, which policies matched, what was
masked, and what the next hop received.

You find them under **Govern › Evaluation sets**.

<Note>
  **Since 0.4, runs are real.** A run relays a real message through the same
  policy chain production traffic uses, under a fresh correlation id, and reads
  the verdicts back from the audit trail. In 0.3.x, runs were simulated and could
  not fail, so re-run your sets after upgrading before you rely on them.
</Note>

***

## What an evaluation checks

| Part | Meaning |
| - | - |
| **Source** | Who sends the test message: a user, a channel or an agent. MCP servers and LLM gateways cannot be sources. |
| **Target** | The agent the message is addressed to. |
| **Input** | The message text. |
| **Expected hops** | An ordered list of what should happen along the way. |

For each hop you can assert:

* **Expected action**: Allow (`LOG`), `WARN`, `MASK`, `HUMAN_REVIEW_REQUIRED`
  or `BLOCK`. When several actions happen on one hop, the most severe one
  counts, in this order from least to most severe: `LOG`, `WARN`, `MASK`,
  `HUMAN_REVIEW_REQUIRED`, `BLOCK`.
* **Expected policies**: policies that must have *matched* on that hop.
* **Masked fields**: the entity types the PII detectors must have found and
  masked (for example `EMAIL_ADDRESS`).
* **Patterns**: regular expressions that the payload **as delivered** must or
  must not match. They are checked after masking, never against the agent's
  reply.

***

## Running a set

Open a set and choose **Run pack**, or select evaluations and choose **Run
selected**.
Select a case to open it beside the list. It shows:

* expected and observed connections,
* the outcome of each check,
* the input, the delivered payload and the response,
* a link to the trace.

| Status | Meaning |
| - | - |
| Passed | Every check on every asserted hop held. |
| Failed | The platform did something other than what you expected. This is the result to act on. |
| Error | The run could not tell. The relay could not complete the message, or a hop was never reached. Fix the route, not the policy. |
| Skipped | The evaluation is disabled. |

<Warning>
  **Target agents really execute.** A run sends a real message, so the agent
  does whatever it does with a real message: calls tools, calls other agents,
  calls the LLM. Point evaluations at agents and tools where that is safe.
  Rate-limit policies stand aside for evaluation traffic, so runs do not use up
  your tenant's budget.
</Warning>

***

## Limits today

* **Only the entry hop is reliably assertable.** Hops are matched by their
  position in the route, not by participant. Assertions on deeper hops can
  match the wrong hop when a route branches, so keep them disabled until
  matching is by participant.
* **Runs are started by a person.** There is no scheduled or API-triggered run
  yet.
* **Slow routes need a longer timeout.** One evaluation may take up to
  `EVALUATION_DISPATCH_TIMEOUT_SECONDS` (default 60) end to end, after which it
  is an Error. On a self-hosted install, raise it with
  `services.relay.settings.EVALUATION_DISPATCH_TIMEOUT_SECONDS`. The
  `values-prismforce-local.yaml` preset sets 150.

***

## Permissions

| To | Needs |
| - | - |
| View sets and results | `TENANT` READ |
| Create, edit and run | `TENANT` WRITE |
| Delete | `TENANT` DELETE |

Evaluation results also count towards an AI system's governance readiness. A
passed evaluation within the freshness window (30 days by default) satisfies
the testing obligation. See [Lifecycle and policy](/governance/lifecycle-and-policy).

***

## API

Everything in the UI is available under `/relay/v1/evaluation-packs`. See the
**Evaluation Packs** group in the [API reference](/api-reference/overview).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.