What is being measured
Phoenix compiles a specification into a working application. The question a bench case asks is whether the application works — booted for real and driven over HTTP — and whether the pipeline is why. So the tooling is taken away in two steps, and the same oracle judges all three results.
- phoenix
- The full pipeline as shipped: spec → clauses → canonical graph → implementation units → generated code, with the architecture target and the pipeline's own internal retries. Driven through the compiled CLI, exactly as a user would.
- baseline
- The same specification text, the same model, one call, no pipeline. Not "a model without tooling" — a model handed a precise specification.
- intent
- One sentence and the runtime contract. No requirement list, no routes. What a sentence alone produces.
How to read a rate here. 28/30 [0.78, 0.98] is one number, not two. The bracket is a
95% Wilson score interval: the range of true rates consistent with what was actually drawn. A perfect score at
five samples is consistent with a true rate of 57%, which is why the fraction never appears without the interval.
When two intervals overlap, the honest statement is these runs do not distinguish these rates — never that
one arm is better.
Citable runs, pooled
Pooled by case, fixture digest, Phoenix commit, arm and model. Never across models, never across digests, and never across commits — a different digest is a different question, and a different commit is a different answerer, which is the whole point when the pipeline is the thing under test. A superseded row was drawn against an older version of the case; it is kept because deleting a number that has been read is how a results log stops being one.
| Case | Model | Arm | Works | Checks passed | n | Outcomes | Fixture / code |
|---|---|---|---|---|---|---|---|
| library-api | anthropic/claude-sonnet-5 | phoenix | 0/5 [0.00, 0.43] | 66/204 [0.26, 0.39] | 5 | 0 working · 4 disagreed · 1 broke | 9d802e59531c88f3 phoenix 14dd469 |
| library-api | anthropic/claude-sonnet-5 | baseline | 5/5 [0.57, 1.00] | 255/255 [0.99, 1.00] | 5 | 5 working · 0 disagreed · 0 broke | 9d802e59531c88f3 phoenix 14dd469 |
| library-api | anthropic/claude-sonnet-5 | intent | 0/5 [0.00, 0.43] | 88/255 [0.29, 0.41] | 5 | 0 working · 5 disagreed · 0 broke | 9d802e59531c88f3 phoenix 14dd469 |
| todo-api | anthropic/claude-sonnet-5 | phoenix | 0/5 [0.00, 0.43] | 70/120 [0.49, 0.67] | 5 | 0 working · 5 disagreed · 0 broke | 9794771ec142f89a phoenix 14dd469 |
| todo-api | anthropic/claude-sonnet-5 | baseline | 5/5 [0.57, 1.00] | 120/120 [0.97, 1.00] | 5 | 5 working · 0 disagreed · 0 broke | 9794771ec142f89a phoenix 14dd469 |
| todo-api | anthropic/claude-sonnet-5 | intent | 0/5 [0.00, 0.43] | 49/120 [0.32, 0.50] | 5 | 0 working · 5 disagreed · 0 broke | 9794771ec142f89a phoenix 14dd469 |
| todo-api | anthropic/claude-sonnet-5 | phoenix | 0/5 [0.00, 0.43] | 25/120 [0.15, 0.29] | 5 | 0 working · 5 disagreed · 0 broke | 9fc6564bd1448cda phoenix 3e996d2 superseded |
| todo-api | anthropic/claude-sonnet-5 | phoenix | 0/5 [0.00, 0.43] | 25/120 [0.15, 0.29] | 5 | 0 working · 5 disagreed · 0 broke | 9794771ec142f89a phoenix db405ca superseded |
| todo-api | anthropic/claude-sonnet-5 | baseline | 5/5 [0.57, 1.00] | 120/120 [0.97, 1.00] | 5 | 5 working · 0 disagreed · 0 broke | 9fc6564bd1448cda phoenix 3e996d2 superseded |
| todo-api | anthropic/claude-sonnet-5 | baseline | 5/5 [0.57, 1.00] | 120/120 [0.97, 1.00] | 5 | 5 working · 0 disagreed · 0 broke | 9794771ec142f89a phoenix db405ca superseded |
| todo-api | anthropic/claude-sonnet-5 | intent | 0/5 [0.00, 0.43] | —/0 (no samples) | 5 | 0 working · 0 disagreed · 5 broke | 9fc6564bd1448cda phoenix 3e996d2 superseded |
| todo-api | anthropic/claude-sonnet-5 | intent | 0/5 [0.00, 0.43] | 47/120 [0.31, 0.48] | 5 | 0 working · 5 disagreed · 0 broke | 9794771ec142f89a phoenix db405ca superseded |
What the comparisons support
- library-api anthropic/claude-sonnet-5 · phoenix 14dd469
These runs distinguish phoenix 0/5 [0.00, 0.43] from baseline 5/5 [0.57, 1.00] (intervals disjoint). - library-api anthropic/claude-sonnet-5 · phoenix 14dd469
These runs do NOT distinguish phoenix 0/5 [0.00, 0.43] from intent 0/5 [0.00, 0.43]. - todo-api anthropic/claude-sonnet-5 · phoenix 3e996d2
These runs distinguish phoenix 0/5 [0.00, 0.43] from baseline 5/5 [0.57, 1.00] (intervals disjoint). - todo-api anthropic/claude-sonnet-5 · phoenix 3e996d2
These runs do NOT distinguish phoenix 0/5 [0.00, 0.43] from intent 0/5 [0.00, 0.43]. - todo-api anthropic/claude-sonnet-5 · phoenix db405ca
These runs distinguish phoenix 0/5 [0.00, 0.43] from baseline 5/5 [0.57, 1.00] (intervals disjoint). - todo-api anthropic/claude-sonnet-5 · phoenix db405ca
These runs do NOT distinguish phoenix 0/5 [0.00, 0.43] from intent 0/5 [0.00, 0.43]. - todo-api anthropic/claude-sonnet-5 · phoenix 14dd469
These runs distinguish phoenix 0/5 [0.00, 0.43] from baseline 5/5 [0.57, 1.00] (intervals disjoint). - todo-api anthropic/claude-sonnet-5 · phoenix 14dd469
These runs do NOT distinguish phoenix 0/5 [0.00, 0.43] from intent 0/5 [0.00, 0.43].
Every recorded run
Smoke, dirty and unstated runs stay visible — they are honest records of what happened — and are excluded from every aggregate and every comparison above.
| When | Case | Model | Arm | Res | n | Works | Other outcomes | Commit | Flags |
|---|---|---|---|---|---|---|---|---|---|
| 2026-08-21 16:08 | todo-api | anthropic/claude-sonnet-5 | intent | coarse | 5 | 0/5 | 5 disagreed · 0 broke | 14dd469 | citable |
| 2026-08-21 16:05 | todo-api | anthropic/claude-sonnet-5 | baseline | coarse | 5 | 5/5 | 0 disagreed · 0 broke | 14dd469 | citable |
| 2026-08-21 16:04 | todo-api | anthropic/claude-sonnet-5 | phoenix | coarse | 5 | 0/5 | 5 disagreed · 0 broke | 14dd469 | citable |
| 2026-08-21 15:09 | todo-api | anthropic/claude-sonnet-5 | phoenix | unstated | 1 | 0/1 | 1 disagreed · 0 broke | e263cae | dirty unstated |
| 2026-08-20 13:39 | todo-api | anthropic/claude-sonnet-5 | intent | coarse | 5 | 0/5 | 5 disagreed · 0 broke | db405ca | citable |
| 2026-08-20 13:36 | todo-api | anthropic/claude-sonnet-5 | baseline | coarse | 5 | 5/5 | 0 disagreed · 0 broke | db405ca | citable |
| 2026-08-20 13:35 | todo-api | anthropic/claude-sonnet-5 | phoenix | coarse | 5 | 0/5 | 5 disagreed · 0 broke | db405ca | citable |
| 2026-08-20 13:25 | todo-api | anthropic/claude-sonnet-5 | intent | coarse | 5 | 0/5 | 0 disagreed · 5 broke | 3e996d2 | citable |
| 2026-08-20 13:17 | todo-api | anthropic/claude-sonnet-5 | baseline | coarse | 5 | 5/5 | 0 disagreed · 0 broke | 3e996d2 | citable |
| 2026-08-20 13:15 | todo-api | anthropic/claude-sonnet-5 | phoenix | coarse | 5 | 0/5 | 5 disagreed · 0 broke | 3e996d2 | citable |
| 2026-08-21 15:54 | library-api | anthropic/claude-sonnet-5 | intent | coarse | 5 | 0/5 | 5 disagreed · 0 broke | 14dd469 | citable |
| 2026-08-21 15:50 | library-api | anthropic/claude-sonnet-5 | baseline | coarse | 5 | 5/5 | 0 disagreed · 0 broke | 14dd469 | citable |
| 2026-08-21 15:46 | library-api | anthropic/claude-sonnet-5 | phoenix | coarse | 5 | 0/5 | 4 disagreed · 1 broke | 14dd469 | citable |
| 2026-08-21 15:13 | library-api | anthropic/claude-sonnet-5 | baseline | unstated | 1 | 1/1 | 0 disagreed · 0 broke | e263cae | dirty unstated |
Glossary
- working · disagreed · broke
- Kept apart rather than reduced to one rate. Working: it booted and every assertion held. Disagreed: it booted, answered, and failed at least one assertion. Broke: a phase died — nothing runnable was produced, or it never booted. A generator that emits nothing and one that emits something subtly wrong are different findings with different fixes.
- unreachable
- The call never reached the model. Excluded from every denominator, because an unreachable endpoint is not a producer that chose badly — and the count is always shown beside the rate that excluded it.
- resolution
- The question a sample size was drawn for:
smoke— does the plumbing work at all?;coarse— differences that are enormous;fine— moving a rate that is already high;unstated— no question was declared. - fixture digest
- A hash of the case: its spec, its checks and its runtime contract. Every fixture is vendored in the repository, so the commit pins the question and the code — everything except the model's sampling.
- budget
- Model calls and retries an arm was allowed. Recorded per entry, never averaged away.
Disclosures
- The phoenix arm gets the pipeline's internal retries (typecheck-and-retry, repair loop). The baseline and intent arms get one call and no retry. That asymmetry is real, is recorded on every entry, and is not corrected for — "the pipeline minus its retry loop" is not a thing anyone can run.
- The intent arm is told the runtime contract but not the routes, so a check that fails because it invented a different URL is a true finding about one sentence of intent, not a scoring accident.
- Every arm is judged by the same oracle: boot the produced app, drive it over HTTP, assert. The oracle imports nothing from the pipeline and is not told which arm produced the code.
- Overlapping intervals mean these runs do not distinguish these rates. They never mean the rates are equal, and they never mean one arm is better.
What this page cannot show
It cannot show why a sample worked. The oracle asserts against HTTP behaviour, so a producer that returned the right responses for the wrong reasons scores the same as one that understood the spec. The sharpest form of this: an application that mounts nothing at all still passes every assertion of the form “a missing thing is missing”, because everything it answers is 404. Read checks passed next to works, never instead of it, and keep positive assertions in the majority when writing a case.
Phoenix's own provenance claims — selective invalidation, drift, the trust surface — are not measured here at
all; they are measured by phoenix selftest, which is a different instrument answering a different
question.
It also cannot show a rate that would be constant by construction. The intent arm is given no file list, so an "authorized paths" rate on that arm would be 1.00 forever; a perfect score that cannot be anything else ends a question instead of inviting one, so it is not drawn.