v0.19.0Verdicts that say how sure they are, detectors that see through disguises, and one command to set up any client→

Output Quality Score

One number that tells you if your agent's output is good enough.

Definition#

Definition

Output Quality Score (OQS) is a composite metric that rolls completeness, relevance, safety, and cost into a single number between 0 and 1 for every agent output. Instead of checking four dimensions separately, teams get one signal: is this output good enough?

Four Dimensions#

Completeness

Did the agent produce a full answer? In Iris: a minimum length, a non-empty output, a sentence count, and coverage of an expected answer when you supply one.

Relevance

Is the response on-topic? In Iris: keyword overlap and topic consistency against the input — lexical checks, not semantic ones; semantic judgment is the LLM judge's job, with your key.

Safety

Is the output safe to show to users? In Iris: PII patterns, prompt-injection patterns, hallucination signals grounded in the input you pass, stub markers, and your blocklist.

Cost

Did the output cost what you allow? In Iris: a cost threshold on the USD you report per trace and a completion-to-prompt token ratio, plus a tool-loop check when you pass the agent's tool calls.

Why One Number Matters#

Individual eval rules tell you what's wrong. The OQS tells you whether you should care. A response might score 0.9 on completeness but 0.3 on relevance — the OQS captures that it's a detailed answer to the wrong question. It's the signal you monitor on a dashboard, set alerts on, and report to stakeholders.

Key Data

A critical safety failure — PII, prompt injection, or a blocklist hit — forces passed: false no matter how high the weighted score is. Output carrying a real SSN still scores 0.765 and still returns passed: false, with no_pii named in critical_failures. The score stays a quality gradient; passed is the verdict. That veto, not a zeroed score, is why you can't average away a safety violation.

How Iris Helps#

Iris scores every output across all four dimensions. The dashboard shows individual rule results and aggregate quality trends. The composite signal makes it easy to spot when overall quality is declining — even when individual dimensions look acceptable in isolation.

Definition

What Iris computes, precisely. “Output Quality Score” is this page's name for the score field evaluate_output returns: a weighted mean over the rules that ran (rules with nothing to judge skip and are excluded), with fixed per-rule weights. passed is the verdict, decided by the threshold and the critical-rule veto. Iris computes no other quality index. How often each rule is right is measured and published, with intervals, on the proof page; how the composite verdict itself performs is not yet measured, and the page says so.

Read the deep dive: Output Quality Score →

Frequently Asked Questions#