v0.10.0The verdict

Capabilities · v0.10.0

What Iris can judge, and what it cannot yet

Ten evaluation questions against six subjects. has means at least one shipped, measured thing answers the question for that subject; partial means something answers it with a stated limit; gap means nothing does yet; n/a means the question does not apply. Every answered cell names its evidence, and each name resolves to something registered in this release — a test fails otherwise. A gap is stated as a gap in Iris, never as a claim about anyone else.

The same map is served to agents inside iris://capabilities, rendered to docs/capabilities.md, and cut from capability-map.json. Accuracy numbers behind the has cells are on the proof page.

has
18
partial
17
gap
21
n/a
4
Capability map: status per question and subject
Questionsingle outputoutput with input / contexttrajectory (tool calls)multi-run of one inputpopulation / dataset / baselinethe evaluator itself
is it safehashasgapgappartialhas
is it grounded / correctpartialhashasgapgappartial
is it completepartialpartialgapgappartialhas
is it on-taskgappartialgapgappartialpartial
did it complete the taskgaphasgapgapgapgap
did it act well (tool choice, arguments, efficiency)n/an/ahasgappartialhas
what did it costhaspartialpartialgappartialhas
is it better or worse than beforen/an/agapgappartialpartial
where and why does it failhashashasgappartialhas
can this verdict be trustedhashaspartialgapgaphas

is it safe

is it grounded / correct

is it complete

is it on-task

did it complete the task

did it act well (tool choice, arguments, efficiency)

what did it cost

is it better or worse than before

where and why does it fail

can this verdict be trusted