Every run, with the scripts.
One planner, one worker, with and without the QA gate, against raw Claude Opus on the same task — including the two runs that failed because of antiloki.
First edition — kept as history (2026-10-01). Every arm on this page ran on 2026-09-22 with Claude as the brain and a DeepSeek model reached through OpenCode Zen as the hands. The product no longer sells that configuration: it is five subscription seats now (Claude, Codex, Cursor, Gemini, Grok), Zen and the Chinese models are optional and off by default, QA is obligatory and its placeholder-claim fault described below is fixed. The same eight-module task, on the five seats and with the dates and sample sizes, is in the technical report, fourth edition (§11.4 and §11.7), and the front page's table is built from it.
A benchmark you cannot inspect is an advertisement. This is every arm we ran, in the order we ran it, with the losses kept in — the shape loses on small tasks, the QA gate measured nothing the day we took these numbers, and two of the biggest runs failed for reasons that were ours. If a number on the front page looks good, the run behind it is on this page.
01What was measured
Two tasks against ministore, a small but real Vite + React + TypeScript store.
S — one function changed in one file. A money formatter with a rounding rule. Small enough that the overhead of planning is visible.
L — eight independent feature modules, each in its own new file under src/features/, with
no changes to any existing file: fuzzy accent-insensitive product search; a stable non-mutating sort;
a useWishlist hook persisted to localStorage that survives a corrupt stored value; validated,
versioned cart storage; money formatting through Intl.NumberFormat with two discount codes;
a QuantityStepper component with bounds, keyboard steps and clamp-on-blur; checkout validation
with a Luhn card check and an expiry in the future; and a capped useRecentlyViewed. Strict
TypeScript, no any, JSDoc on every export, no new dependencies, and
tsc -b && vite build must pass at the end.
02How a run was scored
Three numbers per run, none of them a model's opinion.
Quality is a behavioural audit: 37 checks that import the finished modules and exercise them —
does the search match "blue mug" against "Mug, blue", does the sort leave the input untouched, does
applyDiscount refuse to go below zero, does the Luhn check reject a transposed digit. The same
script grades every arm, including raw Opus.
Cost is deterministic: every call the serve made for that arm — planner, worker, verifier, gate —
priced at published list rates, cache reads at a tenth. Raw Opus is the CLI's own
total_cost_usd. A subscription pays none of this; the dollar figure is the API-equivalent, which
is the only yardstick that compares two different arrangements of the same work.
Time is a wall clock around the whole arm, not the sum of the model calls.
03L — eight modules, the main result
| arm | wall | cost | calls | audit | note |
|---|---|---|---|---|---|
| raw Claude Opus, one session | 303 s | $1.28 | 26 turns | 37/37 | the day before, the same goal took it 112 s and $0.61 — same-day is the fair row |
| Opus plans · DeepSeek Flash writes · QA off | 249 s | $0.25 | 3 | 37/37 | worker 185 s, 24 tools, first pass |
| Sonnet plans · Flash writes · QA off | 266 s | $0.31 | 5 | 35/37 | worker rejected for missing JSDoc on exports → 55 s retry → pass |
| Sonnet · Flash · QA on | 179 s | $0.42 | 4 | 36/37 | first pass · gate 13 s / $0.18 |
| Opus · Flash · QA on | 311 s | $0.54 | 6 | 36/37 | worker retried once · gate 11 s / $0.16 |
The cheapest arm matches raw Opus on all 37 checks, 54 seconds faster, at a fifth of the price. One run per arm — not a best-of-N. Raw Opus moves between days by more than the gap between our arms, which is why the row above carries the previous day's figure instead of hiding it.
04S — one function, where the shape loses
| arm | wall | cost | calls | note |
|---|---|---|---|---|
| raw Claude Opus, one session | 24 s | $0.18 | 5 turns | |
| Sonnet plans · Flash writes · QA off | 44 s | $0.12 | 3 | |
| Sonnet · Flash · QA on | 212 s | $0.34 | 6 | the verifier caught 18.5 → "$18.5" and sent it back |
| Opus plans · Flash · QA off | 414 s | $0.12 | 3 | planner 12 s · worker stalled 393 s — OpenCode tail latency on a one-line edit |
| Opus · Flash · QA on | 48 s | $0.24 | 4 | gate 10 s / $0.12 |
A planner call is pure overhead on a one-line edit. This is the row the product had to obey rather than explain: on a one-file change the benchmark says don't merge, and the router sends the ask straight to one model. A benchmark that only ever agreed with the product would not be worth running.
05What the numbers actually say
The planner's model barely moves quality. Opus planning scored 37 and 36; Sonnet planning scored 35 and
36; raw Opus scored 37. The worker was the same cheap model in every antiloki row. What closes the gap is the
verifier — and both of its rejections were real: missing JSDoc, and a money format that printed
$18.5.
The bill is the planner. On the best L arm: planner ≈ $0.20, worker ≈ $0.02, verifier ≈ $0.03. Almost all of the saving comes from moving the writing off the frontier model, and almost all of the remaining cost is the thinking you kept there.
The QA gate measured nothing that day. Every gate run produced one placeholder claim and returned
unproven in 10–13 s for $0.12–0.18. The always-branch was discarding the planner's acceptance
lines, so the switch bought a bill and no verdict. It is on this page because it happened, and the QA-on rows
in the table above are not slower because of QA — that is the worker's retry noise.
The worker's variance is the runtime, not the model. Same worker, same one-line task: 30 s once, 393 s another time.
06The two runs that failed, and why they were ours
| arm (XXL, same day) | wall | cost | result |
|---|---|---|---|
| raw Opus | 795 s | $5.30 | 169 tests, build ✓ |
| Opus plans · Sonnet writes, one worker | 921 s | $2.64 | 112 tests, build ✓ |
| Sonnet · Haiku, a wave of 8 | failed 327 s | $1.30 | 3 of 8 — Haiku's id generator could not pass the verifier twice |
| Sonnet · Flash, a wave of 8, floor off | failed 1 884 s | $0.18 | 7 of 8 — the plan was a five-deep chain and nobody installed the test runner |
Neither failure was the models'. The first was a plan whose dependency chain the projection did not fold into its wave count; the second was an orchestrator that let a goal ask for tests without installing a test runner. Both are fixed, and both are the reason this section exists: the interesting number in a benchmark is usually the one that went wrong.
07Running it yourself
Every arm above is one shell script against a fresh clone, with its own serve on its own port and its own
ANTILOKI_HOME, so no arm inherits another's history or its learned rankings. The shape is always
the same four steps:
1 · clone ministore into a scratch directory · 2 · start a serve with an empty home
on a free port · 3 · post the goal with the arm's planner, worker and gate setting · 4 · when it finishes, run
audit.mjs over the result and record wall time, dollars and the audit score in one row.
The same measurement runs inside the product: every run you do is scored on your machine and filed in a table that grows with use, which the router reads when it picks a merge for the next job. Ours ships as the starting point; your runs adjust its scores, by at most a third, and none of it leaves the machine.
← more from the antiloki blog · the long version: Compiled, Not Trained →