v0.19.0Verdicts that say how sure they are, detectors that see through disguises, and one command to set up any client→

Comparison · Evaluation

Iris vs Patronus AI

MCP-Native Agent Eval vs Eval Models and Simulation Infrastructure.

TL;DR

Iris is an MCP server your agent discovers and uses on connect — one config block, no SDK, one SQLite file, every rule's precision and recall published at iris-eval.com/proof. Patronus AI — from its own pages, read 2026-09-21: Patronus AI today leads with simulation: 'simulation research and infrastructure to accelerate progress toward human-aligned AGI', built on Digital World Models. Its evaluation line remains: Lynx (SOTA hallucination detection model), GLIDER (3B evaluator scoring text on user-defined criteria), Percival (20+ agent failure modes), a usage-priced eval API ($10 per 1k small evaluator calls), toxicity and PII evaluators, and an Enterprise plan with on-prem/VPC deployment and SOC, TISAX and HIPAA badges. (source)

Iris grades what an agent did with its tools — the trace, the answer, the cost — not whether an MCP server honours its own contract; a server test harness answers that question, and Iris runs beside it. For the method, see the agent eval guide.

Feature comparison

Side by side.

13 features, the same 13 on every comparison. Every Patronus AI cell links the page it was read from and the date; where the sentence it was read from was found verbatim on a plain download of that page (checked 2026-09-22), the link carries it as its title, and where it was not, the cell says so. The highlighted cells are Iris's own call on which side is stronger for a team running MCP agents — 4 to Iris, 3 to Patronus AI — not a measurement.

FeatureIrisPatronus AI
Integration methodOne block in the MCP config, no code — the agent discovers Iris and its tools on connectREST API + SDKPatronus AI's page · read 2026-09-21
Self-hostingOne process, one SQLite file; Docker image with a health checkVendor-hosted API by default; Enterprise plan offers on-prem / dedicated VPC deploymentPatronus AI's page · read 2026-09-21
Where it runsNothing in the agent's process — Iris is a separate server the agent callsEach evaluation is a remote evaluator API call (RemoteEvaluator.evaluate); tracing via @traced decorator + OpenInferencePatronus AI's page · read 2026-09-21
Evaluation25 built-in deterministic rules and 9 custom-rule types, in-process; 7 judge templates on a key you supply; every rule's precision and recall publishedFine-tuned eval models (Lynx, Glider) and a judge evaluator run against criteria defined in the platform console, called through RemoteEvaluatorPatronus AI's page · read 2026-09-21
Cost trackingPer-trace USD cost and tokens; a cost spike judged against the agent's own historyNot stated in the vendor's pages read on 2026-09-21Patronus AI's page · read 2026-09-21
MCP supportProtocol-native — Iris is an MCP server with 12 tools; OTLP traces inPatronus MCP server (evaluate, batch_evaluate, custom_evaluate, run_experiment, create_criteria); API key requiredPatronus AI's page · read 2026-09-21
LicenseMIT, the whole packageProprietary platform with usage-priced API; Lynx and GLIDER model weights released as open source / open weightsPatronus AI's page · read 2026-09-21
OwnershipIndependent and founder-ledIndependent, venture-backed; $50M Series B led by Greenfield Partners with Lightspeed, Notable Capital, Datadog, SamsungPatronus AI's page · read 2026-09-21
DashboardA local dashboard on its own port: traces, moments, regressions, five viewsWeb dashboard for API logs, traces, experiments, comparisons and datasets; Percival agent-trace debuggerPatronus AI's page · read 2026-09-21
Framework supportAny MCP client (2 verified, 8 claimed — see /clients); OTLP/HTTP from anything elseOpenAI, Anthropic, smolagents, Pydantic AI, OpenAI Agents SDK, LangGraph, CrewAI, LangChain via SDK; OpenTelemetryPatronus AI's page · read 2026-09-21
Prompt managementNot includedNot a documented standalone feature; Percival proposes prompt rewrites as remediesPatronus AI's page · read 2026-09-21
Enterprise and complianceSelf-hosted. Nothing leaves your machine unless you set IRIS_OTEL_ENDPOINT, which exports traces to the collector you name, or enable the LLM judge with your own key. No compliance certification is claimed before it is heldAICPA SOC, TISAX and HIPAA badges on pricing page; Enterprise: on-prem / dedicated VPC, custom data retention, SSOPatronus AI's page · read 2026-09-21
Cost to runFree — MIT, one process on your machine; the only spend is a judge call on a key you supply, when you opt inDeveloper plan free with no credit card (2 projects; logs and traces for the last 2 weeks); the API starts with $10 in free credits; Enterprise by salesPatronus AI's page · read 2026-09-21

Decision guide

Which one fits your stack?

When to choose Iris

  • You are building with MCP-compatible agents and want the integration to be one config block
  • You want the evaluation to be deterministic and local — no model call, nothing leaving the machine
  • You want to read what each rule is worth before you trust it: every rule's precision and recall is published
  • You want self-hosting to be one process and one file
  • You want a fully permissive MIT license on the whole package

When to choose Patronus AI

  • You need deep semantic hallucination detection with purpose-built models source
  • You're in a regulated industry requiring enterprise-grade safety scoring source
  • You need custom fine-tuned evaluation models for your specific domain source
  • You want managed infrastructure with SOC, TISAX and HIPAA attestations (badges on the vendor's pricing page) source
  • You need toxicity and PII evaluators and custom policies beyond heuristic rules source

FAQ

The questions buyers ask.

What is the difference between Iris and Patronus AI?
Iris is an MCP-native agent eval server that needs no SDK. Add the config block, restart your client, and every session lists Iris's tools on connect. Iris never intercepts: it runs when your agent calls one of its tools, when a host hook or `iris-eval ingest` hands it a trace, or when you POST one to its HTTP API. Patronus AI now describes itself as 'Simulating the World's Intelligence' — developing simulation research and infrastructure (Digital World Models) toward human-aligned AGI — while still offering its evaluation models Lynx (hallucination detection) and GLIDER, Percival for agent-trace debugging, and a usage-priced evaluation API with a Python SDK and an MCP server. source
Is Iris or Patronus AI better for detecting hallucinations?
Patronus AI's Lynx is a fine-tuned hallucination detection model that the vendor says 'outperforms GPT-4o, Claude-3-Sonnet and closed and open-source LLM-as-a-judge models'; GLIDER scores text against user-defined criteria with reasoning chains. source
How much does Patronus AI cost to run?
Iris: Free — MIT, one process on your machine; the only spend is a judge call on a key you supply, when you opt in. Patronus AI, from its own page read 2026-09-21: Developer plan free with no credit card (2 projects; logs and traces for the last 2 weeks); the API starts with $10 in free credits; Enterprise by sales. source
Can I self-host Patronus AI?
Iris: One process, one SQLite file; Docker image with a health check. Patronus AI, from its own page read 2026-09-21: Vendor-hosted API by default; Enterprise plan offers on-prem / dedicated VPC deployment. source
Does Patronus AI work with MCP agents?
Iris: Protocol-native — Iris is an MCP server with 12 tools; OTLP traces in. Patronus AI, from its own page read 2026-09-21: Patronus MCP server (evaluate, batch_evaluate, custom_evaluate, run_experiment, create_criteria); API key required. source

Sources

Every Patronus AI statement on this page was read from one of these pages on the date shown. The file behind this page is website/src/lib/compare/patronus-ai.json.

Last verified: 2026-09-21. This comparison is based on publicly available documentation and may not reflect recent changes to Patronus AI. We aim to keep this page accurate and fair.

See something outdated or incorrect? Report an inaccuracy — we review and update within 48 hours.

Ready to see what your agents are doing?

Add Iris to your MCP config. First trace in 60 seconds. No SDK, no signup, no infrastructure.