Evaluation sets
An evaluation set (an evaluation pack in the API) tests one promise you make about your agents, such as “agents must not leak PII” or “consequential decisions need a human”. Each evaluation in the set sends one test message along one route through your estate. It then checks what the platform actually did at each hop: which action it took, which policies matched, what was masked, and what the next hop received. You find them under Govern › Evaluation sets.Since 0.4, runs are real. A run relays a real message through the same
policy chain production traffic uses, under a fresh correlation id, and reads
the verdicts back from the audit trail. In 0.3.x, runs were simulated and could
not fail, so re-run your sets after upgrading before you rely on them.
What an evaluation checks
For each hop you can assert:
- Expected action: Allow (
LOG),WARN,MASK,HUMAN_REVIEW_REQUIREDorBLOCK. When several actions happen on one hop, the most severe one counts, in this order from least to most severe:LOG,WARN,MASK,HUMAN_REVIEW_REQUIRED,BLOCK. - Expected policies: policies that must have matched on that hop.
- Masked fields: the entity types the PII detectors must have found and
masked (for example
EMAIL_ADDRESS). - Patterns: regular expressions that the payload as delivered must or must not match. They are checked after masking, never against the agent’s reply.
Running a set
Open a set and choose Run pack, or select evaluations and choose Run selected. Select a case to open it beside the list. It shows:- expected and observed connections,
- the outcome of each check,
- the input, the delivered payload and the response,
- a link to the trace.
Limits today
- Only the entry hop is reliably assertable. Hops are matched by their position in the route, not by participant. Assertions on deeper hops can match the wrong hop when a route branches, so keep them disabled until matching is by participant.
- Runs are started by a person. There is no scheduled or API-triggered run yet.
- Slow routes need a longer timeout. One evaluation may take up to
EVALUATION_DISPATCH_TIMEOUT_SECONDS(default 60) end to end, after which it is an Error. On a self-hosted install, raise it withservices.relay.settings.EVALUATION_DISPATCH_TIMEOUT_SECONDS. Thevalues-prismforce-local.yamlpreset sets 150.
Permissions
Evaluation results also count towards an AI system’s governance readiness. A
passed evaluation within the freshness window (30 days by default) satisfies
the testing obligation. See Lifecycle and policy.
API
Everything in the UI is available under/relay/v1/evaluation-packs. See the
Evaluation Packs group in the API reference.