v0.19.0Verdicts that say how sure they are, detectors that see through disguises, and one command to set up any client→
All posts

Iris vs Langfuse vs Phoenix vs Promptfoo: where each wins, where each loses

·Ian Parent
agent-evalcomparisonlangfusephoenixpromptfoomcp
Written October 2026. Any product counts below describe Iris as it was then. It now ships 25 built-in rules across 4 categories and 12 MCP tools. The live numbers, each with the measurement behind it, are on the proof page.

Four tools, four different answers to the same question: how do you know what an AI agent did, and whether it was any good? Langfuse is an open-source AI engineering platform, part of ClickHouse since January 2026 (announcement). Arize Phoenix is the open-source half of Arize, whose acquisition by Dynatrace was announced in August 2026 (announcement). Promptfoo is an open-source CLI for evaluating and red-teaming LLM apps, now part of OpenAI (repository). Iris is an MCP server that evaluates agent traces with deterministic rules and publishes each rule's precision and recall.

They are not four flavours of one thing. Two are observability platforms that added evaluation, one is a test runner, and one is an evaluation server that speaks the agent's own protocol. The differences that matter to a team choosing between them are the boring ones: how it gets into the stack, where the evaluation runs, what it costs, and what happens when the agent uses tools.

Every vendor statement below links the vendor's own page, read on 2026-10-01; the compare pages carry the same cells with the sentence each was read from and whether that sentence was found on the page.

How it gets into your stack

Langfuse is an SDK: its Python and JS SDKs wrap your functions with an observe decorator or wrapper, which "is an easy way to automatically capture inputs, outputs, timings, and errors of a wrapped function" (Langfuse SDK docs). Beyond the SDKs it lists 100+ library and framework integrations and OpenTelemetry (Langfuse docs).

Phoenix is OpenTelemetry: you instrument the app with the OpenTelemetry SDK plus OpenInference auto-instrumentation, and export spans to the Phoenix collector (Phoenix tracing docs). If your framework is already among its tracing integrations, that is a few lines (Phoenix integrations).

Promptfoo is a test runner: declarative test cases run from the CLI or as a library, locally or as a CI step, calling the model providers directly (Promptfoo intro). Output assertions need nothing inside the app; its trajectory assertions read OpenTelemetry spans the app sends to Promptfoo (tracing).

Iris is one block in the MCP client's config. The agent connects, discovers Iris's tools, and logs traces or asks for verdicts through them; frameworks that are not MCP clients send OpenTelemetry traces to Iris's OTLP door instead (clients).

Where the evaluation runs

Langfuse evaluates with LLM-as-a-judge, Jev as a judge (a decision model that "does not sample text, so the same state and question return the same verdict"), human annotation, and custom scores through the API and SDK (evaluation overview, Jev as a judge); when "you want Langfuse to run deterministic Python or TypeScript logic for you, use code evaluators" (scores via API/SDK). LLM-as-a-judge calls are model calls you pay for.

Phoenix runs evaluators on the server: "LLM-as-a-judge evaluators backed by Phoenix-managed prompts", and code evaluators whose local backends ship with Phoenix, so they run on a self-hosted deployment (server evals, code evaluators). The managed Arize AX adds agent-as-a-judge (Arize AX docs).

Promptfoo goes furthest into the agent's trajectory: its assertion library includes trajectory:* and tool-call assertions alongside model-graded rubrics and custom JavaScript or Python (assertions reference). The assertions run on your machine against what the run produced.

Iris runs its built-in rules in-process, on the trace, with no model call, and publishes every built-in rule's precision and recall on a labelled corpus at iris-eval.com/proof, regenerated from the code at each release. Judge templates exist for the cases a rule cannot decide, on a key you supply.

What it costs to run

Prices are the vendors' own, read on 2026-10-01; the compare pages carry the sentence each was read from.

Self-hosting

Langfuse self-hosts as web and worker containers backed by PostgreSQL, ClickHouse, Redis or Valkey, and S3 or blob storage (self-hosting). It is a real deployment.

Phoenix is pip install arize-phoenix and phoenix serve (terminal); "by default Phoenix starts with a file-based SQLite database in a temporary folder", with PostgreSQL as the other database (configuration), and Docker and Kubernetes among its deployment options.

Promptfoo runs locally; a Docker image hosts a results server, and the vendor's own page says self-hosting "is not recommended for production use cases" (self-hosting).

Iris is one process and one SQLite file, or the Docker image with a health check (README).

What each does with MCP

Langfuse offers a hosted MCP server that can query observations, metrics and datasets and create scores (changelog), and create evaluators and evaluation rules (changelog) — a way for a coding agent to drive Langfuse.

Phoenix builds a remote MCP server into Phoenix 19 and later. "The operation catalog is generated from the Phoenix REST API", so an agent can work with projects, traces, datasets, experiments, prompts and annotations (remote MCP), including SQL over traces (Arize blog) — again, a way for a coding agent to drive Phoenix.

Promptfoo goes the other way: its mcp provider "calls Model Context Protocol (MCP) tools directly, so you can test or red team the server itself" (MCP provider) — it tests MCP servers.

Iris is an MCP server. The agent discovers it on connect and logs and evaluates through its tools; Iris grades what the agent did with its tools, not whether a server honours its contract — a server test harness like Promptfoo's answers that question, and Iris runs beside it (capabilities).

Ownership and license

Where each wins, where each loses

Langfuse wins when you need prompt management with versions and labels (prompt management), broad framework coverage, and enterprise compliance on paper today (SOC 2, ISO 27001, HIPAA per its security page). It loses on weight: self-hosting means two application containers and four datastores (self-hosting).

Phoenix wins when your stack is already OpenTelemetry and you want an open-source tracing and evaluation UI that starts with one pip install (terminal). It loses for a team that needs an OSI-approved license: Phoenix is ELv2 rather than MIT (license), and agent-as-a-judge lives in the managed AX tier (Arize AX docs).

Promptfoo wins for pre-deployment testing: declarative cases, deterministic and trajectory assertions, red-teaming plugins, all from the CLI in CI (intro). It loses as an open-source production observer: the open-source tool tests before deployment, runtime protection is a separate commercial product (Guardrails), and the vendor does not recommend self-hosting its results server for production (self-hosting).

Iris wins when the agent speaks MCP and you want deterministic, local evaluation whose accuracy is published before you rely on it (proof) — one config block, one process, one file. It loses when you need prompt management, a compliance certificate today, or a hundred framework integrations; those are not what it is, and the compare pages say so in muted cells rather than pretending otherwise (compare).

How to read this

Every vendor statement above was read from the linked page on 2026-10-01. The compare pages carry the same cells with the sentence each was read from, the date, and whether a plain download of the page still carries that sentence; the file behind each page is in the repository under website/src/lib/compare/. If a vendor's page has changed, the cell is wrong, and the fix is a pull request.

See what your agents are actually doing

Add Iris to your MCP config. First trace in 60 seconds. No SDK, no signup.