v0.19.0Verdicts that say how sure they are, detectors that see through disguises, and one command to set up any client→

Comparison · Evaluation

Iris vs Confident AI

MCP-Native Eval vs Cloud Evaluation Platform.

TL;DR

Iris is an MCP server your agent discovers and uses on connect — one config block, no SDK, one SQLite file, every rule's precision and recall published at iris-eval.com/proof. Confident AI — from its own pages, read 2026-09-21: Confident AI is 'the cloud platform built on top of DeepEval', described by the vendor as an AI quality platform built for enterprise teams to standardize evals and observability: LLM-as-a-judge metrics, dashboards and alerting, CI/CD regression testing, dataset generation and curation from production, prompt versioning, and self-hosting on Enterprise plans. (source)

Iris grades what an agent did with its tools — the trace, the answer, the cost — not whether an MCP server honours its own contract; a server test harness answers that question, and Iris runs beside it. For the method, see the agent eval guide.

Feature comparison

Side by side.

13 features, the same 13 on every comparison. Every Confident AI cell links the page it was read from and the date; where the sentence it was read from was found verbatim on a plain download of that page (checked 2026-09-22), the link carries it as its title, and where it was not, the cell says so. The highlighted cells are Iris's own call on which side is stronger for a team running MCP agents — 3 to Iris, 3 to Confident AI — not a measurement.

FeatureIrisConfident AI
Integration methodOne block in the MCP config, no code — the agent discovers Iris and its tools on connectPython or TypeScript SDKs (deepeval + confident-trace) with a CONFIDENT_API_KEY; OpenTelemetry also supportedConfident AI's page · read 2026-09-21
Self-hostingOne process, one SQLite file; Docker image with a health checkManaged cloud; Enterprise plans can self-host into your own VPC/VNet via Terraform + Helm (AWS, GCP, Azure)Confident AI's page · read 2026-09-21
Where it runsNothing in the agent's process — Iris is a separate server the agent callsSpans via confident-trace decorators or OpenTelemetry in your app; evals run on Confident AI as traces are ingestedConfident AI's page · read 2026-09-21
Evaluation25 built-in deterministic rules and 9 custom-rule types, in-process; 7 judge templates on a key you supply; every rule's precision and recall publishedLLM-as-Judge metrics (semantic)Confident AI's page · read 2026-09-21
Cost trackingPer-trace USD cost and tokens; a cost spike judged against the agent's own historyAutomatic or manual token cost per LLM span, per-model pricing lookup; cost per user via user_idConfident AI's page · read 2026-09-21
MCP supportProtocol-native — Iris is an MCP server with 12 tools; OTLP traces inMCP server for coding agents (75 tools: run evals, evaluate traces, simulate conversations); needs Confident AI accountConfident AI's page · read 2026-09-21
LicenseMIT, the whole packageProprietary hosted platform on paid plans; the DeepEval framework underneath is Apache 2.0Confident AI's page · read 2026-09-21
OwnershipIndependent and founder-ledIndependent; Y Combinator-backed, $2.2M seed (YC, Flex Capital, Vermilion Cliffs, Liquid 2, January Capital, Rebel Fund)Confident AI's page · read 2026-09-21
DashboardA local dashboard on its own port: traces, moments, regressions, five viewsAgent graph view and user-level analytics dashboards; alerting to Slack, PagerDuty, emailConfident AI's page · read 2026-09-21
Framework supportAny MCP client (2 verified, 8 claimed — see /clients); OTLP/HTTP from anything elseLangChain, CrewAI, OpenAI Agents SDK, LlamaIndex and more; native Python and TypeScript SDKs plus OpenTelemetryConfident AI's page · read 2026-09-21
Prompt managementNot includedPrompt edits tracked as commits, promoted to versions, labelled (staging/production) and pulled in code by labelConfident AI's page · read 2026-09-21
Enterprise and complianceSelf-hosted. Nothing leaves your machine unless you set IRIS_OTEL_ENDPOINT, which exports traces to the collector you name, or enable the LLM judge with your own key. No compliance certification is claimed before it is heldTeam: custom RBAC, SOC2, SSO; Enterprise: dedicated on-prem deployment, HIPAA, custom data residency, 24x7 supportConfident AI's page · read 2026-09-21
Cost to runFree — MIT, one process on your machine; the only spend is a judge call on a key you supply, when you opt inFree Forever $0 (2 seats, 1 project, 5 test runs a week, 1 GB-month of spans); Starter $200/month; Team $2,000/month; Enterprise customConfident AI's page · read 2026-09-21

Decision guide

Which one fits your stack?

When to choose Iris

  • You are building with MCP-compatible agents and want the integration to be one config block
  • You want the evaluation to be deterministic and local — no model call, nothing leaving the machine
  • You want to read what each rule is worth before you trust it: every rule's precision and recall is published
  • You want self-hosting to be one process and one file
  • You want a fully permissive MIT license on the whole package

When to choose Confident AI

  • You need team collaboration with shared dashboards and experiments source
  • You want LLM-as-Judge evaluation for nuanced semantic scoring source
  • You need regression testing to compare model versions source
  • You want synthetic dataset generation for comprehensive test coverage source
  • You prefer a managed platform without infrastructure overhead source

FAQ

The questions buyers ask.

What is the difference between Iris and Confident AI?
Iris is an MCP-native agent eval server that needs no SDK. Add the config block, restart your client, and every session lists Iris's tools on connect. Iris never intercepts: it runs when your agent calls one of its tools, when a host hook or `iris-eval ingest` hands it a trace, or when you POST one to its HTTP API. Confident AI is the cloud platform built on top of DeepEval that adds centralized test management, observability, collaboration, and analytics. It is a commercial platform (Free, Starter, Team and Enterprise plans), with self-hosting into your own VPC available on Enterprise plans. source
Should I use Iris or Confident AI for my AI project?
Use Confident AI if you need a managed platform with team collaboration, CI/CD regression testing, dataset curation from production, and LLM-as-a-judge evaluation on every ingested trace. source
How much does Confident AI cost to run?
Iris: Free — MIT, one process on your machine; the only spend is a judge call on a key you supply, when you opt in. Confident AI, from its own page read 2026-09-21: Free Forever $0 (2 seats, 1 project, 5 test runs a week, 1 GB-month of spans); Starter $200/month; Team $2,000/month; Enterprise custom. source
Can I self-host Confident AI?
Iris: One process, one SQLite file; Docker image with a health check. Confident AI, from its own page read 2026-09-21: Managed cloud; Enterprise plans can self-host into your own VPC/VNet via Terraform + Helm (AWS, GCP, Azure). source
Does Confident AI work with MCP agents?
Iris: Protocol-native — Iris is an MCP server with 12 tools; OTLP traces in. Confident AI, from its own page read 2026-09-21: MCP server for coding agents (75 tools: run evals, evaluate traces, simulate conversations); needs Confident AI account. source

Sources

Every Confident AI statement on this page was read from one of these pages on the date shown. The file behind this page is website/src/lib/compare/confident-ai.json.

Last verified: 2026-09-21. This comparison is based on publicly available documentation and may not reflect recent changes to Confident AI. We aim to keep this page accurate and fair.

See something outdated or incorrect? Report an inaccuracy — we review and update within 48 hours.

Ready to see what your agents are doing?

Add Iris to your MCP config. First trace in 60 seconds. No SDK, no signup, no infrastructure.