Proof
How often the evaluators themselves are wrong
Every eval tool tells you your agent’s score. This page tells you how much to trust the scorer: each built-in rule’s precision, recall and F1, with the uncertainty shown rather than hidden, on a corpus you can download, from a command you can run.
Generated September 6, 2026 for v0.10.0 · corpus 02fb6dff0a73 · schema v2
- Rules measured
- 18 / 18
- Labelled cases
- 777
- Corpus
- 02fb6dff0a73
- Generated
- 2026-09-06
Bar = 95% confidence interval · tick = point estimate · scale 0 to 1. n splits into labelled violations (+) and labelled clean outputs (−); tp, fp, fn, tn are the confusion counts the three figures are computed from. PPV at 1% · 5% is what a fire is worth when one output in a hundred, or five, is a violation: the corpus is about half positive, and a deployment rarely is, so the precision column overstates a fire at field prevalence.
completeness
| Rule | n | Precision | Recall | F1 | PPV at 1% · 5% |
|---|---|---|---|---|---|
| min_output_length | 29 = 13 + / 16 − tp 13 · fp 0 · fn 0 · tn 16 | 1.00 0.77–1.00 | 1.00 0.77–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
| non_empty_output | 29 = 12 + / 17 − tp 10 · fp 0 · fn 2 · tn 17 | 1.00 0.72–1.00 | 0.83 0.55–0.95 | 0.91 0.74–1.00 | 1.00 · 1.00 |
| sentence_count | 30 = 16 + / 14 − tp 16 · fp 0 · fn 0 · tn 14 | 1.00 0.81–1.00 | 1.00 0.81–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
| expected_coverage | 29 = 14 + / 15 − tp 14 · fp 0 · fn 0 · tn 15 | 1.00 0.78–1.00 | 1.00 0.78–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
| valid_tool_arguments | 33 = 15 + / 18 − tp 15 · fp 0 · fn 0 · tn 18 | 1.00 0.80–1.00 | 1.00 0.80–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
| ask_coverage | 31 = 14 + / 17 − tp 12 · fp 4 · fn 2 · tn 13 | 0.75 0.51–0.90 | 0.86 0.60–0.96 | 0.80 0.61–0.93 | 0.04 · 0.16 |
relevance
| Rule | n | Precision | Recall | F1 | PPV at 1% · 5% |
|---|---|---|---|---|---|
| keyword_overlap | 30 = 10 + / 20 − tp 10 · fp 0 · fn 0 · tn 20 | 1.00 0.72–1.00 | 1.00 0.72–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
| topic_consistency | 31 = 12 + / 19 − tp 11 · fp 0 · fn 1 · tn 19 | 1.00 0.74–1.00 | 0.92 0.65–0.99 | 0.96 0.83–1.00 | 1.00 · 1.00 |
safety
| Rule | n | Precision | Recall | F1 | PPV at 1% · 5% |
|---|---|---|---|---|---|
| no_pii | 93 = 45 + / 48 − tp 34 · fp 5 · fn 11 · tn 43 | 0.87 0.73–0.94 | 0.76 0.61–0.86 | 0.81 0.71–0.89 | 0.07 · 0.28 |
| no_blocklist_words | 32 = 15 + / 17 − tp 14 · fp 1 · fn 1 · tn 16 | 0.93 0.70–0.99 | 0.93 0.70–0.99 | 0.93 0.81–1.00 | 0.14 · 0.46 |
| no_injection_patterns | 90 = 42 + / 48 − tp 41 · fp 0 · fn 1 · tn 48 | 1.00 0.91–1.00 | 0.98 0.88–1.00 | 0.99 0.96–1.00 | 1.00 · 1.00 |
| no_stub_output | 89 = 42 + / 47 − tp 30 · fp 5 · fn 12 · tn 42 | 0.86 0.71–0.94 | 0.71 0.56–0.83 | 0.78 0.67–0.87 | 0.06 · 0.26 |
| no_hallucination_markers | 90 = 46 + / 44 − tp 34 · fp 0 · fn 12 · tn 44 | 1.00 0.90–1.00 | 0.74 0.60–0.84 | 0.85 0.76–0.92 | 1.00 · 1.00 |
| no_silent_tool_failure | 30 = 14 + / 16 − tp 13 · fp 0 · fn 1 · tn 16 | 1.00 0.77–1.00 | 0.93 0.69–0.99 | 0.96 0.87–1.00 | 1.00 · 1.00 |
| grounded_in_reads | 32 = 16 + / 16 − tp 16 · fp 0 · fn 0 · tn 16 | 1.00 0.81–1.00 | 1.00 0.81–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
cost
| Rule | n | Precision | Recall | F1 | PPV at 1% · 5% |
|---|---|---|---|---|---|
| cost_under_threshold | 26 = 10 + / 16 − tp 10 · fp 0 · fn 0 · tn 16 | 1.00 0.72–1.00 | 1.00 0.72–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
| verbosity_ratio | 25 = 9 + / 16 − tp 9 · fp 0 · fn 0 · tn 16 | 1.00 0.70–1.00 | 1.00 0.70–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
| no_tool_loop | 28 = 12 + / 16 − tp 12 · fp 0 · fn 0 · tn 16 | 1.00 0.76–1.00 | 1.00 0.76–1.00 | 1.00 1.00–1.00 | 1.00 · 1.00 |
How to read it
Each rule is run over a labelled corpus of agent outputs: precision is the share of outputs the rule flagged that were truly violations, recall is the share of true violations the rule flagged, and F1 is their harmonic mean. The intervals are 95% confidence intervals (wilson-95 for precision and recall, bootstrap-percentile-2000-mulberry32-seed-proof-f1-bootstrap-v1 for F1); a wide bar means a small sample, not a bad rule, and a rule whose interval reaches low is one whose verdict you should not gate a deploy on alone.
Two kinds of rule, two kinds of number. Where a rule’s documented definition is a formula (an output length, a coverage count, a keyword overlap, a cost, a token ratio), the corpus checks that the implementation matches its definition, so a precision and recall near 1.0 there mean “implemented as documented”, not “catches what goes wrong in the wild”. Where the definition names a judgement (PII, prompt injection, hallucination signals, stub answers, blocklisted phrases, staying on topic), the corpus measures detection against cases written to evade it, and those are the numbers to weigh before trusting a verdict.
Where the corpus came from
The disclosures that matter, stated before the numbers are read:
- Synthetic. The cases are constructed, not sampled from production traffic. They cover the shapes the rules are meant to catch; they do not tell you the base rate in your agent.
- LLM-authored. A language model wrote the cases, so they share one author’s idea of what a violation looks like.
- Same-model-labelled. The labels were produced by the same model family that authored the cases, which means the labels share the author’s blind spots. A rule that agrees with those labels agrees with that model, not yet with a human.
- Human blind label: pending. founder blind label of a 140-case stratified sample (twenty per judgment family); until then the labels are same-model dual annotation (see proof/README.md)
The verdict, measured
A gate does not key on a rule; it keys on passed, the verdict the composer makes from every rule that ran. The composite corpus scores that verdict: the 24 real transcripts, promoted as they are, plus cases composed by splicing a rule family’s case into a clean transcript so the failure classes present are true by construction. It scores the risk composer that has decided passed since 0.10.0 beside the arithmetic that decided it before, on the same rule results, so the change is readable rather than asserted.
| Split | Composer | Accuracy vs should-ship (95% CI) | False blocks on clean | Missed blocks |
|---|---|---|---|---|
| test | legacy | 39.3% [23.6, 57.6] n=28 | 11.1% n=9 | 84.2% n=19 |
| test | risk, per-output prior | 67.9% [49.3, 82.1] n=28 | 11.1% n=9 | 42.1% n=19 |
| real transcripts | legacy | 45.8% [27.9, 64.9] n=24 | 0.0% n=6 | 72.2% n=18 |
| real transcripts | risk, per-output prior | 70.8% [50.8, 85.1] n=24 | 0.0% n=6 | 38.9% n=18 |
Difference from legacy on the test split, risk composer with the per-output prior: 28.6 points [2.5, 49.8] (Newcombe 95%). An interval that straddles zero says the corpus cannot yet tell the two apart. Read per class, the plan’s prior blocks nearly every clean case; read per output it holds the legacy false-block rate — both readings are in the file, and which one ships is a decision the numbers inform. Cases: 119 (24 real transcripts, 95 composed; 91 dev / 28 test by a hash of the id). Composite ecc52d7452d9 · proof/COMPOSITE.md · proof/composite-results.json · npm run proof -- --composite.
The evasions a leak arrives in
The three critical rules match text. For every positive a rule caught with a span into the raw output, the text inside the span is transformed — a zero-width space, Cyrillic homoglyphs, fullwidth forms, a no-break space, a tab, a line break, swapped case — and the rule is run again. Recall per transform is the share still caught; a case the rule missed in the clear is not counted, and a transform that does not apply to a span (no letters to swap, no space to replace) is not counted for that case.
| Rule | zero width | homoglyph | fullwidth | nbsp | tab | linebreak | case |
|---|---|---|---|---|---|---|---|
| no_pii | 100.0% n=34 | 100.0% n=27 | 100.0% n=34 | 100.0% n=11 | 52.9% n=34 | 38.2% n=34 | 55.6% n=27 |
| no_injection_patterns | 100.0% n=41 | 100.0% n=41 | 100.0% n=41 | 100.0% n=36 | 43.9% n=41 | 36.6% n=41 | 97.6% n=41 |
| no_blocklist_words | 100.0% n=14 | 100.0% n=14 | 100.0% n=14 | 100.0% n=13 | 7.1% n=14 | 0.0% n=14 | 100.0% n=14 |
Per-transform intervals and the dropped case ids are in proof/RESULTS.md (results.json → transforms). The table measures the rules as shipped; a normalisation pass that changes it will change these numbers.
What no_pii finds, by entity
Every positive in the PII family names what it contains — by the case author, never by the detector. Per entity: present (cases containing it), caught (the rule failed the case for any reason), named (the rule’s evidence named this entity). The vocabulary includes things the rule’s definition does not cover, so a gap shows as a row with named 0 rather than as silence, and a case caught for another reason shows as the difference between caught and named.
| Entity | Present | Caught | Named | Recall (95% CI) |
|---|---|---|---|---|
| ssn | 3 | 3 | 3 | 100.0% [43.9, 100.0] |
| credit_card | 2 | 2 | 2 | 100.0% [34.2, 100.0] |
| iban | 1 | 1 | 0 | 0.0% [0.0, 79.3] |
| phone | 7 | 7 | 7 | 100.0% [64.6, 100.0] |
| 11 | 11 | 11 | 100.0% [74.1, 100.0] | |
| dob | 1 | 1 | 1 | 100.0% [20.6, 100.0] |
| private_key | 2 | 2 | 2 | 100.0% [34.2, 100.0] |
| seed_phrase | 1 | 1 | 1 | 100.0% [20.6, 100.0] |
| api_key | 17 | 10 | 10 | 58.8% [36.0, 78.4] |
| password | 2 | 2 | 0 | 0.0% [0.0, 65.8] |
| address | 6 | 3 | 0 | 0.0% [0.0, 39.0] |
| url_token | 1 | 0 | 0 | 0.0% [0.0, 79.3] |
Custom rule types — conformance
A custom rule is your own constraint, so its accuracy is whether the type does what its documented definition says under a declared config. One family per type, run through the same factory custom_rules and deployed rules go through; a disagreement here is a rule defect or a definition error, never an opinion.
| Type | Config | n | Skipped | Precision (95% CI) | Recall (95% CI) |
|---|---|---|---|---|---|
| contains_keywords | {"keywords":["refund","policy","days"],"threshold":1} | 24 | 0 | 100.0% [78.5, 100.0] | 100.0% [78.5, 100.0] |
| cost_threshold | {"max_cost":0.05} | 24 | 2 | 100.0% [75.8, 100.0] | 100.0% [75.8, 100.0] |
| excludes_keywords | {"keywords":["guarantee","risk-free"]} | 24 | 0 | 100.0% [74.1, 100.0] | 100.0% [74.1, 100.0] |
| json_schema | {"schema":{"additionalProperties":false,"properties":{"count":{"type":"integer"},"id":{"type":"string"},"nested":{"properties":{"ok":{"type":"boolean"}},"required":["ok"],"type":"object"},"tags":{"items":{"type":"string"},"type":"array"}},"required":["id","count"],"type":"object"}} | 27 | 0 | 100.0% [79.6, 100.0] | 100.0% [79.6, 100.0] |
| max_length | {"max_length":120} | 24 | 0 | 100.0% [78.5, 100.0] | 100.0% [78.5, 100.0] |
| min_length | {"min_length":80} | 24 | 0 | 100.0% [77.2, 100.0] | 100.0% [77.2, 100.0] |
| regex_match | {"pattern":"^Ticket #[0-9]{6}\\b"} | 24 | 0 | 100.0% [78.5, 100.0] | 100.0% [78.5, 100.0] |
| regex_no_match | {"pattern":"\\bTODO\\b"} | 24 | 0 | 100.0% [75.8, 100.0] | 100.0% [75.8, 100.0] |
Families: proof/corpus/custom/<type>.json. The json_schema family’s definition says the schema is not consulted in this version; its cases are labelled by that.
Evaluator of evaluators
Thirteen trust questions — does it work, when does it fail, what does it measure, is it calibrated, can it be gamed, can it produce false confidence, and seven more — asked of every evaluator Iris ships: the built-in rules, the custom rule types, the judge templates, the citation verifier and the verdict composer. Every cell is derived from the committed proof files by a generator, never typed; a cell reads measured only when a number for it exists and the evidence names the file.
- Evaluators
- 33
- Questions each
- 13
- ≥ 3 questions measured
- 27 / 33
- the built-in rules: 18 of 18 with three or more questions measured.
- the custom rule types: 8 of 8 with three or more questions measured.
- the LLM-judge templates: 0 of 5 with three or more questions measured — measurable, pending a judge key that you or the maintainer supplies.
- the citation verifier: 0 of 1 with three or more questions measured — measurable, pending a judge key that you or the maintainer supplies.
- the verdict composer: 1 of 1 with three or more questions measured.
The full matrix with the evidence per cell: docs/evaluators.md. Regenerated by npm run claims:generate and npm run llms:render at every release.
Related: the capability map says what Iris can judge; this page says how well.
What is not in the table
The LLM-judge path (evaluate_with_llm_judge) and the citation verifier (verify_citations) are semantic, model-backed, and not deterministic; they are measured separately, on an adversarial set, under a cost cap. Status: pending. run npm run proof:judge with a judge API key, or dispatch the proof-judge workflow
Run it yourself
The corpus, the runner and the results file live in the repository. The numbers on this page are the committed results; regenerate them and diff:
git clone https://github.com/iris-eval/mcp-server.git cd mcp-server npm ci npm run proof # writes proof/results.json git diff proof/results.json
No network and no API key are needed: the built-in rules are deterministic, so the run completes offline and identical inputs produce identical figures. The page is rendered from the same results file through the truthbase (.claims.json), and the hardcoded-claim scanner refuses any public sentence that says “measured” without linking here.
Related
- Security — the controls in place, and measured issue-close times.
- Roadmap — Track 1 is this page; what comes after it.
- docs/roadmap.md — the measurement commitments in full, including the ones not yet met.