v0.10.0The verdict

Proof

How often the evaluators themselves are wrong

Every eval tool tells you your agent’s score. This page tells you how much to trust the scorer: each built-in rule’s precision, recall and F1, with the uncertainty shown rather than hidden, on a corpus you can download, from a command you can run.

Generated September 6, 2026 for v0.10.0 · corpus 02fb6dff0a73 · schema v2

Rules measured
18 / 18
Labelled cases
777
Corpus
02fb6dff0a73
Generated
2026-09-06

Bar = 95% confidence interval · tick = point estimate · scale 0 to 1. n splits into labelled violations (+) and labelled clean outputs (−); tp, fp, fn, tn are the confusion counts the three figures are computed from. PPV at 1% · 5% is what a fire is worth when one output in a hundred, or five, is a violation: the corpus is about half positive, and a deployment rarely is, so the precision column overstates a fire at field prevalence.

completeness

RulenPrecisionRecallF1PPV at 1% · 5%
min_output_length
29 = 13 + / 16
tp 13 · fp 0 · fn 0 · tn 16
1.00
precision 1.00, 95% interval 0.77 to 1.000.771.00
1.00
recall 1.00, 95% interval 0.77 to 1.000.771.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00
non_empty_output
29 = 12 + / 17
tp 10 · fp 0 · fn 2 · tn 17
1.00
precision 1.00, 95% interval 0.72 to 1.000.721.00
0.83
recall 0.83, 95% interval 0.55 to 0.950.550.95
0.91
F1 0.91, 95% interval 0.74 to 1.000.741.00
1.00 · 1.00
sentence_count
30 = 16 + / 14
tp 16 · fp 0 · fn 0 · tn 14
1.00
precision 1.00, 95% interval 0.81 to 1.000.811.00
1.00
recall 1.00, 95% interval 0.81 to 1.000.811.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00
expected_coverage
29 = 14 + / 15
tp 14 · fp 0 · fn 0 · tn 15
1.00
precision 1.00, 95% interval 0.78 to 1.000.781.00
1.00
recall 1.00, 95% interval 0.78 to 1.000.781.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00
valid_tool_arguments
33 = 15 + / 18
tp 15 · fp 0 · fn 0 · tn 18
1.00
precision 1.00, 95% interval 0.80 to 1.000.801.00
1.00
recall 1.00, 95% interval 0.80 to 1.000.801.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00
ask_coverage
31 = 14 + / 17
tp 12 · fp 4 · fn 2 · tn 13
0.75
precision 0.75, 95% interval 0.51 to 0.900.510.90
0.86
recall 0.86, 95% interval 0.60 to 0.960.600.96
0.80
F1 0.80, 95% interval 0.61 to 0.930.610.93
0.04 · 0.16

relevance

RulenPrecisionRecallF1PPV at 1% · 5%
keyword_overlap
30 = 10 + / 20
tp 10 · fp 0 · fn 0 · tn 20
1.00
precision 1.00, 95% interval 0.72 to 1.000.721.00
1.00
recall 1.00, 95% interval 0.72 to 1.000.721.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00
topic_consistency
31 = 12 + / 19
tp 11 · fp 0 · fn 1 · tn 19
1.00
precision 1.00, 95% interval 0.74 to 1.000.741.00
0.92
recall 0.92, 95% interval 0.65 to 0.990.650.99
0.96
F1 0.96, 95% interval 0.83 to 1.000.831.00
1.00 · 1.00

safety

RulenPrecisionRecallF1PPV at 1% · 5%
no_pii
93 = 45 + / 48
tp 34 · fp 5 · fn 11 · tn 43
0.87
precision 0.87, 95% interval 0.73 to 0.940.730.94
0.76
recall 0.76, 95% interval 0.61 to 0.860.610.86
0.81
F1 0.81, 95% interval 0.71 to 0.890.710.89
0.07 · 0.28
no_blocklist_words
32 = 15 + / 17
tp 14 · fp 1 · fn 1 · tn 16
0.93
precision 0.93, 95% interval 0.70 to 0.990.700.99
0.93
recall 0.93, 95% interval 0.70 to 0.990.700.99
0.93
F1 0.93, 95% interval 0.81 to 1.000.811.00
0.14 · 0.46
no_injection_patterns
90 = 42 + / 48
tp 41 · fp 0 · fn 1 · tn 48
1.00
precision 1.00, 95% interval 0.91 to 1.000.911.00
0.98
recall 0.98, 95% interval 0.88 to 1.000.881.00
0.99
F1 0.99, 95% interval 0.96 to 1.000.961.00
1.00 · 1.00
no_stub_output
89 = 42 + / 47
tp 30 · fp 5 · fn 12 · tn 42
0.86
precision 0.86, 95% interval 0.71 to 0.940.710.94
0.71
recall 0.71, 95% interval 0.56 to 0.830.560.83
0.78
F1 0.78, 95% interval 0.67 to 0.870.670.87
0.06 · 0.26
no_hallucination_markers
90 = 46 + / 44
tp 34 · fp 0 · fn 12 · tn 44
1.00
precision 1.00, 95% interval 0.90 to 1.000.901.00
0.74
recall 0.74, 95% interval 0.60 to 0.840.600.84
0.85
F1 0.85, 95% interval 0.76 to 0.920.760.92
1.00 · 1.00
no_silent_tool_failure
30 = 14 + / 16
tp 13 · fp 0 · fn 1 · tn 16
1.00
precision 1.00, 95% interval 0.77 to 1.000.771.00
0.93
recall 0.93, 95% interval 0.69 to 0.990.690.99
0.96
F1 0.96, 95% interval 0.87 to 1.000.871.00
1.00 · 1.00
grounded_in_reads
32 = 16 + / 16
tp 16 · fp 0 · fn 0 · tn 16
1.00
precision 1.00, 95% interval 0.81 to 1.000.811.00
1.00
recall 1.00, 95% interval 0.81 to 1.000.811.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00

cost

RulenPrecisionRecallF1PPV at 1% · 5%
cost_under_threshold
26 = 10 + / 16
tp 10 · fp 0 · fn 0 · tn 16
1.00
precision 1.00, 95% interval 0.72 to 1.000.721.00
1.00
recall 1.00, 95% interval 0.72 to 1.000.721.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00
verbosity_ratio
25 = 9 + / 16
tp 9 · fp 0 · fn 0 · tn 16
1.00
precision 1.00, 95% interval 0.70 to 1.000.701.00
1.00
recall 1.00, 95% interval 0.70 to 1.000.701.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00
no_tool_loop
28 = 12 + / 16
tp 12 · fp 0 · fn 0 · tn 16
1.00
precision 1.00, 95% interval 0.76 to 1.000.761.00
1.00
recall 1.00, 95% interval 0.76 to 1.000.761.00
1.00
F1 1.00, 95% interval 1.00 to 1.001.001.00
1.00 · 1.00

How to read it

Each rule is run over a labelled corpus of agent outputs: precision is the share of outputs the rule flagged that were truly violations, recall is the share of true violations the rule flagged, and F1 is their harmonic mean. The intervals are 95% confidence intervals (wilson-95 for precision and recall, bootstrap-percentile-2000-mulberry32-seed-proof-f1-bootstrap-v1 for F1); a wide bar means a small sample, not a bad rule, and a rule whose interval reaches low is one whose verdict you should not gate a deploy on alone.

Two kinds of rule, two kinds of number. Where a rule’s documented definition is a formula (an output length, a coverage count, a keyword overlap, a cost, a token ratio), the corpus checks that the implementation matches its definition, so a precision and recall near 1.0 there mean “implemented as documented”, not “catches what goes wrong in the wild”. Where the definition names a judgement (PII, prompt injection, hallucination signals, stub answers, blocklisted phrases, staying on topic), the corpus measures detection against cases written to evade it, and those are the numbers to weigh before trusting a verdict.

Where the corpus came from

The disclosures that matter, stated before the numbers are read:

  • Synthetic. The cases are constructed, not sampled from production traffic. They cover the shapes the rules are meant to catch; they do not tell you the base rate in your agent.
  • LLM-authored. A language model wrote the cases, so they share one author’s idea of what a violation looks like.
  • Same-model-labelled. The labels were produced by the same model family that authored the cases, which means the labels share the author’s blind spots. A rule that agrees with those labels agrees with that model, not yet with a human.
  • Human blind label: pending. founder blind label of a 140-case stratified sample (twenty per judgment family); until then the labels are same-model dual annotation (see proof/README.md)

The verdict, measured

A gate does not key on a rule; it keys on passed, the verdict the composer makes from every rule that ran. The composite corpus scores that verdict: the 24 real transcripts, promoted as they are, plus cases composed by splicing a rule family’s case into a clean transcript so the failure classes present are true by construction. It scores the risk composer that has decided passed since 0.10.0 beside the arithmetic that decided it before, on the same rule results, so the change is readable rather than asserted.

SplitComposerAccuracy vs should-ship (95% CI)False blocks on cleanMissed blocks
testlegacy39.3% [23.6, 57.6] n=2811.1% n=984.2% n=19
testrisk, per-output prior67.9% [49.3, 82.1] n=2811.1% n=942.1% n=19
real transcriptslegacy45.8% [27.9, 64.9] n=240.0% n=672.2% n=18
real transcriptsrisk, per-output prior70.8% [50.8, 85.1] n=240.0% n=638.9% n=18

Difference from legacy on the test split, risk composer with the per-output prior: 28.6 points [2.5, 49.8] (Newcombe 95%). An interval that straddles zero says the corpus cannot yet tell the two apart. Read per class, the plan’s prior blocks nearly every clean case; read per output it holds the legacy false-block rate — both readings are in the file, and which one ships is a decision the numbers inform. Cases: 119 (24 real transcripts, 95 composed; 91 dev / 28 test by a hash of the id). Composite ecc52d7452d9 · proof/COMPOSITE.md · proof/composite-results.json · npm run proof -- --composite.

The evasions a leak arrives in

The three critical rules match text. For every positive a rule caught with a span into the raw output, the text inside the span is transformed — a zero-width space, Cyrillic homoglyphs, fullwidth forms, a no-break space, a tab, a line break, swapped case — and the rule is run again. Recall per transform is the share still caught; a case the rule missed in the clear is not counted, and a transform that does not apply to a span (no letters to swap, no space to replace) is not counted for that case.

Rulezero widthhomoglyphfullwidthnbsptablinebreakcase
no_pii100.0% n=34100.0% n=27100.0% n=34100.0% n=1152.9% n=3438.2% n=3455.6% n=27
no_injection_patterns100.0% n=41100.0% n=41100.0% n=41100.0% n=3643.9% n=4136.6% n=4197.6% n=41
no_blocklist_words100.0% n=14100.0% n=14100.0% n=14100.0% n=137.1% n=140.0% n=14100.0% n=14

Per-transform intervals and the dropped case ids are in proof/RESULTS.md (results.json → transforms). The table measures the rules as shipped; a normalisation pass that changes it will change these numbers.

What no_pii finds, by entity

Every positive in the PII family names what it contains — by the case author, never by the detector. Per entity: present (cases containing it), caught (the rule failed the case for any reason), named (the rule’s evidence named this entity). The vocabulary includes things the rule’s definition does not cover, so a gap shows as a row with named 0 rather than as silence, and a case caught for another reason shows as the difference between caught and named.

EntityPresentCaughtNamedRecall (95% CI)
ssn333100.0% [43.9, 100.0]
credit_card222100.0% [34.2, 100.0]
iban1100.0% [0.0, 79.3]
phone777100.0% [64.6, 100.0]
email111111100.0% [74.1, 100.0]
dob111100.0% [20.6, 100.0]
private_key222100.0% [34.2, 100.0]
seed_phrase111100.0% [20.6, 100.0]
api_key17101058.8% [36.0, 78.4]
password2200.0% [0.0, 65.8]
address6300.0% [0.0, 39.0]
url_token1000.0% [0.0, 79.3]

Custom rule types — conformance

A custom rule is your own constraint, so its accuracy is whether the type does what its documented definition says under a declared config. One family per type, run through the same factory custom_rules and deployed rules go through; a disagreement here is a rule defect or a definition error, never an opinion.

TypeConfignSkippedPrecision (95% CI)Recall (95% CI)
contains_keywords{"keywords":["refund","policy","days"],"threshold":1}240100.0% [78.5, 100.0]100.0% [78.5, 100.0]
cost_threshold{"max_cost":0.05}242100.0% [75.8, 100.0]100.0% [75.8, 100.0]
excludes_keywords{"keywords":["guarantee","risk-free"]}240100.0% [74.1, 100.0]100.0% [74.1, 100.0]
json_schema{"schema":{"additionalProperties":false,"properties":{"count":{"type":"integer"},"id":{"type":"string"},"nested":{"properties":{"ok":{"type":"boolean"}},"required":["ok"],"type":"object"},"tags":{"items":{"type":"string"},"type":"array"}},"required":["id","count"],"type":"object"}}270100.0% [79.6, 100.0]100.0% [79.6, 100.0]
max_length{"max_length":120}240100.0% [78.5, 100.0]100.0% [78.5, 100.0]
min_length{"min_length":80}240100.0% [77.2, 100.0]100.0% [77.2, 100.0]
regex_match{"pattern":"^Ticket #[0-9]{6}\\b"}240100.0% [78.5, 100.0]100.0% [78.5, 100.0]
regex_no_match{"pattern":"\\bTODO\\b"}240100.0% [75.8, 100.0]100.0% [75.8, 100.0]

Families: proof/corpus/custom/<type>.json. The json_schema family’s definition says the schema is not consulted in this version; its cases are labelled by that.

Evaluator of evaluators

Thirteen trust questions — does it work, when does it fail, what does it measure, is it calibrated, can it be gamed, can it produce false confidence, and seven more — asked of every evaluator Iris ships: the built-in rules, the custom rule types, the judge templates, the citation verifier and the verdict composer. Every cell is derived from the committed proof files by a generator, never typed; a cell reads measured only when a number for it exists and the evidence names the file.

Evaluators
33
Questions each
13
≥ 3 questions measured
27 / 33
  • the built-in rules: 18 of 18 with three or more questions measured.
  • the custom rule types: 8 of 8 with three or more questions measured.
  • the LLM-judge templates: 0 of 5 with three or more questions measured — measurable, pending a judge key that you or the maintainer supplies.
  • the citation verifier: 0 of 1 with three or more questions measured — measurable, pending a judge key that you or the maintainer supplies.
  • the verdict composer: 1 of 1 with three or more questions measured.

The full matrix with the evidence per cell: docs/evaluators.md. Regenerated by npm run claims:generate and npm run llms:render at every release.

Related: the capability map says what Iris can judge; this page says how well.

What is not in the table

The LLM-judge path (evaluate_with_llm_judge) and the citation verifier (verify_citations) are semantic, model-backed, and not deterministic; they are measured separately, on an adversarial set, under a cost cap. Status: pending. run npm run proof:judge with a judge API key, or dispatch the proof-judge workflow

Run it yourself

The corpus, the runner and the results file live in the repository. The numbers on this page are the committed results; regenerate them and diff:

git clone https://github.com/iris-eval/mcp-server.git
cd mcp-server
npm ci
npm run proof          # writes proof/results.json
git diff proof/results.json

No network and no API key are needed: the built-in rules are deterministic, so the run completes offline and identical inputs produce identical figures. The page is rendered from the same results file through the truthbase (.claims.json), and the hardcoded-claim scanner refuses any public sentence that says “measured” without linking here.

Related

  • Security — the controls in place, and measured issue-close times.
  • Roadmap — Track 1 is this page; what comes after it.
  • docs/roadmap.md — the measurement commitments in full, including the ones not yet met.