v0.19.0Verdicts that say how sure they are, detectors that see through disguises, and one command to set up any client→

Comparison · Evaluation

Iris vs DeepEval

MCP-Native Heuristic Eval vs LLM-as-Judge Framework.

TL;DR

Iris is an MCP server your agent discovers and uses on connect — one config block, no SDK, one SQLite file, every rule's precision and recall published at iris-eval.com/proof. DeepEval — from its own pages, read 2026-09-21: DeepEval is the open-source LLM evaluation framework from the Confident AI team: eval files run the same way pytest would, almost all predefined metrics are LLM-as-a-judge (QAG, DAG, G-Eval), and it is built to run as a gate in a CI/CD pipeline. (source)

Iris grades what an agent did with its tools — the trace, the answer, the cost — not whether an MCP server honours its own contract; a server test harness answers that question, and Iris runs beside it. For the method, see the agent eval guide.

Feature comparison

Side by side.

13 features, the same 13 on every comparison. Every DeepEval cell links the page it was read from and the date; where the sentence it was read from was found verbatim on a plain download of that page (checked 2026-09-22), the link carries it as its title, and where it was not, the cell says so. The highlighted cells are Iris's own call on which side is stronger for a team running MCP agents — 2 to Iris, 2 to DeepEval — not a measurement.

FeatureIrisDeepEval
Integration methodOne block in the MCP config, no code — the agent discovers Iris and its tools on connectPython library (TypeScript SDK in beta); `deepeval test run` collects and runs eval files the way pytest wouldDeepEval's page · read 2026-09-21
Self-hostingOne process, one SQLite file; Docker image with a health checkpip install, local executionDeepEval's page · read 2026-09-21
Where it runsNothing in the agent's process — Iris is a separate server the agent callsAlmost all metrics are LLM-as-a-judge calls to a model provider; tracing via the @observe decorator in-processDeepEval's page · read 2026-09-21
Evaluation25 built-in deterministic rules and 9 custom-rule types, in-process; 7 judge templates on a key you supply; every rule's precision and recall published50+ ready-to-use metrics, almost all LLM-as-a-judge (QAG, DAG, G-Eval); Tool Correctness and JSON Correctness need no judge; custom via G-Eval or BaseMetricDeepEval's page · read 2026-09-21
Cost trackingPer-trace USD cost and tokens; a cost spike judged against the agent's own historyToken cost fields on traced spansDeepEval's page · read 2026-09-21
MCP supportProtocol-native — Iris is an MCP server with 12 tools; OTLP traces inEvaluates MCP use (MCPUseMetric, MultiTurnMCPUseMetric); MCPServer objects via mcp_servers; not itself an MCP serverDeepEval's page · read 2026-09-21
LicenseMIT, the whole packageApache 2.0DeepEval's page · read 2026-09-21
OwnershipIndependent and founder-ledCreated and maintained by Confident AI (independent, Y Combinator-backed; $2.2M seed round)DeepEval's page · read 2026-09-21
DashboardA local dashboard on its own port: traces, moments, regressions, five viewsConfident AI cloud dashboard (separate product)DeepEval's page · read 2026-09-21
Framework supportAny MCP client (2 verified, 8 claimed — see /clients); OTLP/HTTP from anything elseLangChain, LangGraph, LlamaIndex, CrewAI, Pydantic AI, OpenAI Agents, Google ADK, Strands; TS: Mastra, Vercel AI SDKDeepEval's page · read 2026-09-21
Prompt managementNot includedPrompt class works locally from files or code; versioned store with pull() by alias, version or label via Confident AIDeepEval's page · read 2026-09-21
Enterprise and complianceSelf-hosted. Nothing leaves your machine unless you set IRIS_OTEL_ENDPOINT, which exports traces to the collector you name, or enable the LLM judge with your own key. No compliance certification is claimed before it is heldNone in the library; RBAC, SOC2, SSO (Team) and on-prem, HIPAA (Enterprise) come from the Confident AI platformDeepEval's page · read 2026-09-21
Cost to runFree — MIT, one process on your machine; the only spend is a judge call on a key you supply, when you opt inThe library is free to run under Apache 2.0; each metric is a call to the judge model you configure, billed by that provider; Confident AI cloud has a free tierDeepEval's page · read 2026-09-21

Decision guide

Which one fits your stack?

When to choose Iris

  • You are building with MCP-compatible agents and want the integration to be one config block
  • You want the evaluation to be deterministic and local — no model call, nothing leaving the machine
  • You want to read what each rule is worth before you trust it: every rule's precision and recall is published
  • You want self-hosting to be one process and one file
  • You want a fully permissive MIT license on the whole package

When to choose DeepEval

  • You're building in Python (or TypeScript) and want pytest-style eval workflows source
  • You want to run eval suites in CI/CD pipelines before deployment source
  • You need a mature ecosystem with extensive documentation and community source
  • You want the Confident AI cloud platform for team collaboration source

FAQ

The questions buyers ask.

What is the difference between Iris and DeepEval?
Iris is an MCP-native agent eval server that needs no SDK. Add the config block, restart your client, and every session lists Iris's tools on connect. Iris never intercepts: it runs when your agent calls one of its tools, when a host hook or `iris-eval ingest` hands it a trace, or when you POST one to its HTTP API. DeepEval is the open-source LLM evaluation framework (Apache 2.0) whose `deepeval test run` collects and runs eval files the same way pytest would; almost all of its 50+ metrics are LLM-as-a-judge. It is Python-first with a TypeScript SDK in beta, and it can also trace apps with @observe and stream traces to Confident AI. source
Should I use Iris or DeepEval for agent evaluation?
Use DeepEval if you're building in Python (or TypeScript) and want pytest-style eval suites with LLM-as-a-judge metrics running in CI/CD. source
How much does DeepEval cost to run?
Iris: Free — MIT, one process on your machine; the only spend is a judge call on a key you supply, when you opt in. DeepEval, from its own page read 2026-09-21: The library is free to run under Apache 2.0; each metric is a call to the judge model you configure, billed by that provider; Confident AI cloud has a free tier. source
Can I self-host DeepEval?
Iris: One process, one SQLite file; Docker image with a health check. DeepEval, from its own page read 2026-09-21: pip install, local execution. source
Does DeepEval work with MCP agents?
Iris: Protocol-native — Iris is an MCP server with 12 tools; OTLP traces in. DeepEval, from its own page read 2026-09-21: Evaluates MCP use (MCPUseMetric, MultiTurnMCPUseMetric); MCPServer objects via mcp_servers; not itself an MCP server. source

Sources

Every DeepEval statement on this page was read from one of these pages on the date shown. The file behind this page is website/src/lib/compare/deepeval.json.

Last verified: 2026-09-21. This comparison is based on publicly available documentation and may not reflect recent changes to DeepEval. We aim to keep this page accurate and fair.

See something outdated or incorrect? Report an inaccuracy — we review and update within 48 hours.

Ready to see what your agents are doing?

Add Iris to your MCP config. First trace in 60 seconds. No SDK, no signup, no infrastructure.