Comparison · Evaluation
MCP-Native Eval vs Cloud Evaluation Platform.
TL;DR
Iris grades what an agent did with its tools — the trace, the answer, the cost — not whether an MCP server honours its own contract; a server test harness answers that question, and Iris runs beside it. For the method, see the agent eval guide.
Feature comparison
13 features, the same 13 on every comparison. Every Confident AI cell links the page it was read from and the date; where the sentence it was read from was found verbatim on a plain download of that page (checked 2026-09-22), the link carries it as its title, and where it was not, the cell says so. The highlighted cells are Iris's own call on which side is stronger for a team running MCP agents — 3 to Iris, 3 to Confident AI — not a measurement.
| Feature | Iris | Confident AI |
|---|---|---|
| Integration method | One block in the MCP config, no code — the agent discovers Iris and its tools on connect | Python or TypeScript SDKs (deepeval + confident-trace) with a CONFIDENT_API_KEY; OpenTelemetry also supportedConfident AI's page · read 2026-09-21 |
| Self-hosting | One process, one SQLite file; Docker image with a health check | Managed cloud; Enterprise plans can self-host into your own VPC/VNet via Terraform + Helm (AWS, GCP, Azure)Confident AI's page · read 2026-09-21 |
| Where it runs | Nothing in the agent's process — Iris is a separate server the agent calls | Spans via confident-trace decorators or OpenTelemetry in your app; evals run on Confident AI as traces are ingestedConfident AI's page · read 2026-09-21 |
| Evaluation | 25 built-in deterministic rules and 9 custom-rule types, in-process; 7 judge templates on a key you supply; every rule's precision and recall published | LLM-as-Judge metrics (semantic)Confident AI's page · read 2026-09-21 |
| Cost tracking | Per-trace USD cost and tokens; a cost spike judged against the agent's own history | Automatic or manual token cost per LLM span, per-model pricing lookup; cost per user via user_idConfident AI's page · read 2026-09-21 |
| MCP support | Protocol-native — Iris is an MCP server with 12 tools; OTLP traces in | MCP server for coding agents (75 tools: run evals, evaluate traces, simulate conversations); needs Confident AI accountConfident AI's page · read 2026-09-21 |
| License | MIT, the whole package | Proprietary hosted platform on paid plans; the DeepEval framework underneath is Apache 2.0Confident AI's page · read 2026-09-21 |
| Ownership | Independent and founder-led | Independent; Y Combinator-backed, $2.2M seed (YC, Flex Capital, Vermilion Cliffs, Liquid 2, January Capital, Rebel Fund)Confident AI's page · read 2026-09-21 |
| Dashboard | A local dashboard on its own port: traces, moments, regressions, five views | Agent graph view and user-level analytics dashboards; alerting to Slack, PagerDuty, emailConfident AI's page · read 2026-09-21 |
| Framework support | Any MCP client (2 verified, 8 claimed — see /clients); OTLP/HTTP from anything else | LangChain, CrewAI, OpenAI Agents SDK, LlamaIndex and more; native Python and TypeScript SDKs plus OpenTelemetryConfident AI's page · read 2026-09-21 |
| Prompt management | Not included | Prompt edits tracked as commits, promoted to versions, labelled (staging/production) and pulled in code by labelConfident AI's page · read 2026-09-21 |
| Enterprise and compliance | Self-hosted. Nothing leaves your machine unless you set IRIS_OTEL_ENDPOINT, which exports traces to the collector you name, or enable the LLM judge with your own key. No compliance certification is claimed before it is held | Team: custom RBAC, SOC2, SSO; Enterprise: dedicated on-prem deployment, HIPAA, custom data residency, 24x7 supportConfident AI's page · read 2026-09-21 |
| Cost to run | Free — MIT, one process on your machine; the only spend is a judge call on a key you supply, when you opt in | Free Forever $0 (2 seats, 1 project, 5 test runs a week, 1 GB-month of spans); Starter $200/month; Team $2,000/month; Enterprise customConfident AI's page · read 2026-09-21 |
Decision guide
FAQ
Every Confident AI statement on this page was read from one of these pages on the date shown. The file behind this page is website/src/lib/compare/confident-ai.json.
Last verified: 2026-09-21. This comparison is based on publicly available documentation and may not reflect recent changes to Confident AI. We aim to keep this page accurate and fair.
See something outdated or incorrect? Report an inaccuracy — we review and update within 48 hours.
Add Iris to your MCP config. First trace in 60 seconds. No SDK, no signup, no infrastructure.