v0.19.0Verdicts that say how sure they are, detectors that see through disguises, and one command to set up any client→

Self-Calibrating Eval

Eval rules that know when their own thresholds are wrong — and tell you how to fix them.

Definition#

Definition

Self-calibrating eval is the pattern where evaluation rules monitor their own scoring distribution and recommend threshold adjustments. Instead of manually tuning thresholds based on intuition, the system observes real output distributions and suggests when thresholds should tighten or loosen. Adjustments are always human-approved.

Why Static Thresholds Break#

A completeness threshold set at 0.7 might be perfect today — passing genuinely good outputs and catching bad ones. Three weeks later, the same threshold passes everything (because model quality improved) or fails everything (because input patterns shifted). The threshold didn't change. The world around it did. This is eval drift at the threshold level.

Key Data

A 100% pass rate is not a sign of quality — it's a sign your thresholds need tightening. A 100% fail rate is not a sign of failure — it's a sign your thresholds need loosening. Both are calibration problems.

The Pattern#

1

Monitor

Track the distribution of scores for each eval rule over time.

2

Detect

Flag when pass rates hit extremes (100% or 0%) or distributions shift significantly.

3

Recommend

Suggest adjusted thresholds based on observed data. Human approves or rejects.

How Iris Helps#

Iris provides the scoring data that this pattern needs. Every output you send it is scored with the same rules, building the distribution data needed to see calibration issues. The dashboard's trend and drift views show pass rates over time — when a rule passes 100% of outputs, it's visible immediately.

Definition

What Iris does today, precisely. Iris does not monitor its own thresholds or recommend adjustments — the Detect and Recommend steps above are yours. Every rule threshold is a setting you choose (eval.ruleThresholds in config.json, documented in the API reference); the dashboard shows the pass rate over time so you can see when a threshold has stopped discriminating. The measured accuracy of each rule, with intervals, is on the proof page.

Read the deep dive: Self-Calibrating Eval →

Frequently Asked Questions#