Compare Iris
Iris is agent eval for MCP — built so you stop shipping agents on vibes. See how it compares to other evaluation and observability platforms — feature by feature, with no vendor lock-in. The same 13 features on every page; every cell about the other product links the page it was read from, with the date.
Iris grades what an agent did with its tools — the trace, the answer, the cost — not whether an MCP server honours its own contract; a server test harness answers that question, and Iris runs beside it.
Iris vs Langfuse
ObservabilityMCP-Native Agent Eval vs SDK-Based Tracing
Iris vs LangSmith
ObservabilityMCP-Native Eval vs LangChain Ecosystem Tracing
Iris vs Helicone
ObservabilityMCP-Native Agent Eval vs API Gateway Observability
Iris vs Braintrust
EvaluationMCP-Native Eval vs Experiment-Driven Evaluation
Iris vs Arize
ObservabilityMCP-Native Eval vs Enterprise ML Observability
Iris vs DeepEval
EvaluationMCP-Native Heuristic Eval vs LLM-as-Judge Framework
Iris vs Confident AI
EvaluationMCP-Native Eval vs Cloud Evaluation Platform
Iris vs Patronus AI
EvaluationMCP-Native Agent Eval vs Eval Models and Simulation Infrastructure
Iris vs Promptfoo
TestingMCP-Native Agent Eval vs Config-Driven CLI Eval and Red-Teaming
Iris vs Galileo
ObservabilityMCP-Native Agent Eval vs SDK-Instrumented Enterprise AI Observability
Iris vs Opik
ObservabilityMCP-Native Agent Eval vs Open-Source SDK-Traced LLM Observability
Iris vs Weave
ObservabilityMCP-Native Agent Eval vs Decorator-Instrumented Agent Observability Platform
Iris vs Judgment Labs
EvaluationMCP-Native Agent Eval vs SDK-Traced Hosted Agent Judge Platform
Iris vs Latitude
ObservabilityMCP-Native Agent Eval vs Open-Source Agent Observability Platform
Why Iris is different
- MCP-native — Iris runs as an MCP server. No SDK, no code changes, no vendor lock-in.
- Heuristic-first — Deterministic rules run on every output. No LLM-as-judge costs or latency.
- Quality + Safety + Cost — Three dimensions scored together. Not just quality, not just safety.
- Self-hosted — Nothing leaves your machine unless you turn one of these on: an OpenTelemetry endpoint (IRIS_OTEL_ENDPOINT), which exports traces to the collector you name; the LLM judge with your own key, which sends the text it judges to that provider, and whose citation check fetches the pages an output cites; or a webhook, which posts ids, the verdict and rule names, never the text, to the address you set. Free and MIT licensed: the open-source core, with no usage limits.