v0.20.0Iris says what it checked. A verdict tells a pass from a check that never ran, names who recorded the evidence it judged, and cannot be replaced by the agent it judges→

Compare Iris

Iris is agent eval for MCP — built so you stop shipping agents on vibes. See how it compares to other evaluation and observability platforms — feature by feature, with no vendor lock-in. The same 13 features on every page; every cell about the other product links the page it was read from, with the date.

Iris grades what an agent did with its tools — the trace, the answer, the cost — not whether an MCP server honours its own contract; a server test harness answers that question, and Iris runs beside it.

Iris vs Langfuse

Observability

MCP-Native Agent Eval vs SDK-Based Tracing

Iris vs LangSmith

Observability

MCP-Native Eval vs LangChain Ecosystem Tracing

Iris vs Helicone

Observability

MCP-Native Agent Eval vs API Gateway Observability

Iris vs Braintrust

Evaluation

MCP-Native Eval vs Experiment-Driven Evaluation

Iris vs Arize

Observability

MCP-Native Eval vs Enterprise ML Observability

Iris vs DeepEval

Evaluation

MCP-Native Heuristic Eval vs LLM-as-Judge Framework

Iris vs Confident AI

Evaluation

MCP-Native Eval vs Cloud Evaluation Platform

Iris vs Patronus AI

Evaluation

MCP-Native Agent Eval vs Eval Models and Simulation Infrastructure

Iris vs Promptfoo

Testing

MCP-Native Agent Eval vs Config-Driven CLI Eval and Red-Teaming

Iris vs Galileo

Observability

MCP-Native Agent Eval vs SDK-Instrumented Enterprise AI Observability

Iris vs Opik

Observability

MCP-Native Agent Eval vs Open-Source SDK-Traced LLM Observability

Iris vs Weave

Observability

MCP-Native Agent Eval vs Decorator-Instrumented Agent Observability Platform

Iris vs Judgment Labs

Evaluation

MCP-Native Agent Eval vs SDK-Traced Hosted Agent Judge Platform

Iris vs Latitude

Observability

MCP-Native Agent Eval vs Open-Source Agent Observability Platform

Why Iris is different