Phoenix

Bench results

Every run the bench has recorded, read straight from the append-only results files. This page adds no measurement of its own — it shows what each run stored, and states plainly what those numbers can and cannot support.

Built 2026-08-22T20:32:19.413Z from commit 9258c7f · 14 recorded run(s), 12 citable group(s)

What is being measured

Phoenix compiles a specification into a working application. The question a bench case asks is whether the application works — booted for real and driven over HTTP — and whether the pipeline is why. So the tooling is taken away in two steps, and the same oracle judges all three results.

phoenix
The full pipeline as shipped: spec → clauses → canonical graph → implementation units → generated code, with the architecture target and the pipeline's own internal retries. Driven through the compiled CLI, exactly as a user would.
baseline
The same specification text, the same model, one call, no pipeline. Not "a model without tooling" — a model handed a precise specification.
intent
One sentence and the runtime contract. No requirement list, no routes. What a sentence alone produces.

How to read a rate here. 28/30 [0.78, 0.98] is one number, not two. The bracket is a 95% Wilson score interval: the range of true rates consistent with what was actually drawn. A perfect score at five samples is consistent with a true rate of 57%, which is why the fraction never appears without the interval. When two intervals overlap, the honest statement is these runs do not distinguish these rates — never that one arm is better.

Citable runs, pooled

Pooled by case, fixture digest, Phoenix commit, arm and model. Never across models, never across digests, and never across commits — a different digest is a different question, and a different commit is a different answerer, which is the whole point when the pipeline is the thing under test. A superseded row was drawn against an older version of the case; it is kept because deleting a number that has been read is how a results log stops being one.

CaseModelArmWorksChecks passednOutcomesFixture / code
library-apianthropic/claude-sonnet-5 phoenix 0/5 [0.00, 0.43] 66/204 [0.26, 0.39] 5 0 working · 4 disagreed · 1 broke 9d802e59531c88f3
phoenix 14dd469
library-apianthropic/claude-sonnet-5 baseline 5/5 [0.57, 1.00] 255/255 [0.99, 1.00] 5 5 working · 0 disagreed · 0 broke 9d802e59531c88f3
phoenix 14dd469
library-apianthropic/claude-sonnet-5 intent 0/5 [0.00, 0.43] 88/255 [0.29, 0.41] 5 0 working · 5 disagreed · 0 broke 9d802e59531c88f3
phoenix 14dd469
todo-apianthropic/claude-sonnet-5 phoenix 0/5 [0.00, 0.43] 70/120 [0.49, 0.67] 5 0 working · 5 disagreed · 0 broke 9794771ec142f89a
phoenix 14dd469
todo-apianthropic/claude-sonnet-5 baseline 5/5 [0.57, 1.00] 120/120 [0.97, 1.00] 5 5 working · 0 disagreed · 0 broke 9794771ec142f89a
phoenix 14dd469
todo-apianthropic/claude-sonnet-5 intent 0/5 [0.00, 0.43] 49/120 [0.32, 0.50] 5 0 working · 5 disagreed · 0 broke 9794771ec142f89a
phoenix 14dd469
todo-apianthropic/claude-sonnet-5 phoenix 0/5 [0.00, 0.43] 25/120 [0.15, 0.29] 5 0 working · 5 disagreed · 0 broke 9fc6564bd1448cda
phoenix 3e996d2
superseded
todo-apianthropic/claude-sonnet-5 phoenix 0/5 [0.00, 0.43] 25/120 [0.15, 0.29] 5 0 working · 5 disagreed · 0 broke 9794771ec142f89a
phoenix db405ca
superseded
todo-apianthropic/claude-sonnet-5 baseline 5/5 [0.57, 1.00] 120/120 [0.97, 1.00] 5 5 working · 0 disagreed · 0 broke 9fc6564bd1448cda
phoenix 3e996d2
superseded
todo-apianthropic/claude-sonnet-5 baseline 5/5 [0.57, 1.00] 120/120 [0.97, 1.00] 5 5 working · 0 disagreed · 0 broke 9794771ec142f89a
phoenix db405ca
superseded
todo-apianthropic/claude-sonnet-5 intent 0/5 [0.00, 0.43] —/0 (no samples) 5 0 working · 0 disagreed · 5 broke 9fc6564bd1448cda
phoenix 3e996d2
superseded
todo-apianthropic/claude-sonnet-5 intent 0/5 [0.00, 0.43] 47/120 [0.31, 0.48] 5 0 working · 5 disagreed · 0 broke 9794771ec142f89a
phoenix db405ca
superseded

What the comparisons support

Every recorded run

Smoke, dirty and unstated runs stay visible — they are honest records of what happened — and are excluded from every aggregate and every comparison above.

WhenCaseModelArmResnWorksOther outcomesCommitFlags
2026-08-21 16:08 todo-api anthropic/claude-sonnet-5 intent coarse 5 0/5 5 disagreed · 0 broke 14dd469 citable
2026-08-21 16:05 todo-api anthropic/claude-sonnet-5 baseline coarse 5 5/5 0 disagreed · 0 broke 14dd469 citable
2026-08-21 16:04 todo-api anthropic/claude-sonnet-5 phoenix coarse 5 0/5 5 disagreed · 0 broke 14dd469 citable
2026-08-21 15:09 todo-api anthropic/claude-sonnet-5 phoenix unstated 1 0/1 1 disagreed · 0 broke e263cae dirty unstated
2026-08-20 13:39 todo-api anthropic/claude-sonnet-5 intent coarse 5 0/5 5 disagreed · 0 broke db405ca citable
2026-08-20 13:36 todo-api anthropic/claude-sonnet-5 baseline coarse 5 5/5 0 disagreed · 0 broke db405ca citable
2026-08-20 13:35 todo-api anthropic/claude-sonnet-5 phoenix coarse 5 0/5 5 disagreed · 0 broke db405ca citable
2026-08-20 13:25 todo-api anthropic/claude-sonnet-5 intent coarse 5 0/5 0 disagreed · 5 broke 3e996d2 citable
2026-08-20 13:17 todo-api anthropic/claude-sonnet-5 baseline coarse 5 5/5 0 disagreed · 0 broke 3e996d2 citable
2026-08-20 13:15 todo-api anthropic/claude-sonnet-5 phoenix coarse 5 0/5 5 disagreed · 0 broke 3e996d2 citable
2026-08-21 15:54 library-api anthropic/claude-sonnet-5 intent coarse 5 0/5 5 disagreed · 0 broke 14dd469 citable
2026-08-21 15:50 library-api anthropic/claude-sonnet-5 baseline coarse 5 5/5 0 disagreed · 0 broke 14dd469 citable
2026-08-21 15:46 library-api anthropic/claude-sonnet-5 phoenix coarse 5 0/5 4 disagreed · 1 broke 14dd469 citable
2026-08-21 15:13 library-api anthropic/claude-sonnet-5 baseline unstated 1 1/1 0 disagreed · 0 broke e263cae dirty unstated

Glossary

working · disagreed · broke
Kept apart rather than reduced to one rate. Working: it booted and every assertion held. Disagreed: it booted, answered, and failed at least one assertion. Broke: a phase died — nothing runnable was produced, or it never booted. A generator that emits nothing and one that emits something subtly wrong are different findings with different fixes.
unreachable
The call never reached the model. Excluded from every denominator, because an unreachable endpoint is not a producer that chose badly — and the count is always shown beside the rate that excluded it.
resolution
The question a sample size was drawn for: smoke — does the plumbing work at all?; coarse — differences that are enormous; fine — moving a rate that is already high; unstated — no question was declared.
fixture digest
A hash of the case: its spec, its checks and its runtime contract. Every fixture is vendored in the repository, so the commit pins the question and the code — everything except the model's sampling.
budget
Model calls and retries an arm was allowed. Recorded per entry, never averaged away.

Disclosures

What this page cannot show

It cannot show why a sample worked. The oracle asserts against HTTP behaviour, so a producer that returned the right responses for the wrong reasons scores the same as one that understood the spec. The sharpest form of this: an application that mounts nothing at all still passes every assertion of the form “a missing thing is missing”, because everything it answers is 404. Read checks passed next to works, never instead of it, and keep positive assertions in the majority when writing a case.

Phoenix's own provenance claims — selective invalidation, drift, the trust surface — are not measured here at all; they are measured by phoenix selftest, which is a different instrument answering a different question.

It also cannot show a rate that would be constant by construction. The intent arm is given no file list, so an "authorized paths" rate on that arm would be 1.00 forever; a perfect score that cannot be anything else ends a question instead of inviting one, so it is not drawn.