Skip to main content

Evaluation sets

An evaluation set (an evaluation pack in the API) tests one promise you make about your agents, such as “agents must not leak PII” or “consequential decisions need a human”. Each evaluation in the set sends one test message along one route through your estate. It then checks what the platform actually did at each hop: which action it took, which policies matched, what was masked, and what the next hop received. You find them under Govern › Evaluation sets.
Since 0.4, runs are real. A run relays a real message through the same policy chain production traffic uses, under a fresh correlation id, and reads the verdicts back from the audit trail. In 0.3.x, runs were simulated and could not fail, so re-run your sets after upgrading before you rely on them.

What an evaluation checks

For each hop you can assert:
  • Expected action: Allow (LOG), WARN, MASK, HUMAN_REVIEW_REQUIRED or BLOCK. When several actions happen on one hop, the most severe one counts, in this order from least to most severe: LOG, WARN, MASK, HUMAN_REVIEW_REQUIRED, BLOCK.
  • Expected policies: policies that must have matched on that hop.
  • Masked fields: the entity types the PII detectors must have found and masked (for example EMAIL_ADDRESS).
  • Patterns: regular expressions that the payload as delivered must or must not match. They are checked after masking, never against the agent’s reply.

Running a set

Open a set and choose Run pack, or select evaluations and choose Run selected. Select a case to open it beside the list. It shows:
  • expected and observed connections,
  • the outcome of each check,
  • the input, the delivered payload and the response,
  • a link to the trace.
Target agents really execute. A run sends a real message, so the agent does whatever it does with a real message: calls tools, calls other agents, calls the LLM. Point evaluations at agents and tools where that is safe. Rate-limit policies stand aside for evaluation traffic, so runs do not use up your tenant’s budget.

Limits today

  • Only the entry hop is reliably assertable. Hops are matched by their position in the route, not by participant. Assertions on deeper hops can match the wrong hop when a route branches, so keep them disabled until matching is by participant.
  • Runs are started by a person. There is no scheduled or API-triggered run yet.
  • Slow routes need a longer timeout. One evaluation may take up to EVALUATION_DISPATCH_TIMEOUT_SECONDS (default 60) end to end, after which it is an Error. On a self-hosted install, raise it with services.relay.settings.EVALUATION_DISPATCH_TIMEOUT_SECONDS. The values-prismforce-local.yaml preset sets 150.

Permissions

Evaluation results also count towards an AI system’s governance readiness. A passed evaluation within the freshness window (30 days by default) satisfies the testing obligation. See Lifecycle and policy.

API

Everything in the UI is available under /relay/v1/evaluation-packs. See the Evaluation Packs group in the API reference.