Output Quality Score
One number that tells you if your agent's output is good enough.
Definition#
Definition
Four Dimensions#
Completeness
Did the agent produce a full answer? In Iris: a minimum length, a non-empty output, a sentence count, and coverage of an expected answer when you supply one.
Relevance
Is the response on-topic? In Iris: keyword overlap and topic consistency against the input — lexical checks, not semantic ones; semantic judgment is the LLM judge's job, with your key.
Safety
Is the output safe to show to users? In Iris: PII patterns, prompt-injection patterns, hallucination signals grounded in the input you pass, stub markers, and your blocklist.
Cost
Did the output cost what you allow? In Iris: a cost threshold on the USD you report per trace and a completion-to-prompt token ratio, plus a tool-loop check when you pass the agent's tool calls.
Why One Number Matters#
Individual eval rules tell you what's wrong. The OQS tells you whether you should care. A response might score 0.9 on completeness but 0.3 on relevance — the OQS captures that it's a detailed answer to the wrong question. It's the signal you monitor on a dashboard, set alerts on, and report to stakeholders.
Key Data
passed: false no matter how high the weighted score is. Output carrying a real SSN still scores 0.765 and still returns passed: false, with no_pii named in critical_failures. The score stays a quality gradient; passed is the verdict. That veto, not a zeroed score, is why you can't average away a safety violation.How Iris Helps#
Iris scores every output across all four dimensions. The dashboard shows individual rule results and aggregate quality trends. The composite signal makes it easy to spot when overall quality is declining — even when individual dimensions look acceptable in isolation.
Definition
score field evaluate_output returns: a weighted mean over the rules that ran (rules with nothing to judge skip and are excluded), with fixed per-rule weights. passed is the verdict, decided by the threshold and the critical-rule veto. Iris computes no other quality index. How often each rule is right is measured and published, with intervals, on the proof page; how the composite verdict itself performs is not yet measured, and the page says so.Read the deep dive: Output Quality Score →
Related Concepts#
Self-Calibrating Eval
Individual dimension thresholds need calibration — which affects the composite OQS.
Eval Coverage
OQS is only meaningful with 100% coverage — a composite score on sampled data misleads.
The Eval Tax
The OQS quantifies what you're losing — low scores show the tax in real-time.
Agent Eval
The complete guide to evaluating AI agent outputs.