antiloki · the benchmark

Every run, with the scripts.

One planner, one worker, with and without the QA gate, against raw Claude Opus on the same task — including the two runs that failed because of antiloki.

First edition — kept as history (2026-10-01). Every arm on this page ran on 2026-09-22 with Claude as the brain and a DeepSeek model reached through OpenCode Zen as the hands. The product no longer sells that configuration: it is five subscription seats now (Claude, Codex, Cursor, Gemini, Grok), Zen and the Chinese models are optional and off by default, QA is obligatory and its placeholder-claim fault described below is fixed. The same eight-module task, on the five seats and with the dates and sample sizes, is in the technical report, fourth edition (§11.4 and §11.7), and the front page's table is built from it.

why this page exists

A benchmark you cannot inspect is an advertisement. This is every arm we ran, in the order we ran it, with the losses kept in — the shape loses on small tasks, the QA gate measured nothing the day we took these numbers, and two of the biggest runs failed for reasons that were ours. If a number on the front page looks good, the run behind it is on this page.

01What was measured

Two tasks against ministore, a small but real Vite + React + TypeScript store.

S — one function changed in one file. A money formatter with a rounding rule. Small enough that the overhead of planning is visible.

L — eight independent feature modules, each in its own new file under src/features/, with no changes to any existing file: fuzzy accent-insensitive product search; a stable non-mutating sort; a useWishlist hook persisted to localStorage that survives a corrupt stored value; validated, versioned cart storage; money formatting through Intl.NumberFormat with two discount codes; a QuantityStepper component with bounds, keyboard steps and clamp-on-blur; checkout validation with a Luhn card check and an expiry in the future; and a capped useRecentlyViewed. Strict TypeScript, no any, JSDoc on every export, no new dependencies, and tsc -b && vite build must pass at the end.

02How a run was scored

Three numbers per run, none of them a model's opinion.

Quality is a behavioural audit: 37 checks that import the finished modules and exercise them — does the search match "blue mug" against "Mug, blue", does the sort leave the input untouched, does applyDiscount refuse to go below zero, does the Luhn check reject a transposed digit. The same script grades every arm, including raw Opus.

Cost is deterministic: every call the serve made for that arm — planner, worker, verifier, gate — priced at published list rates, cache reads at a tenth. Raw Opus is the CLI's own total_cost_usd. A subscription pays none of this; the dollar figure is the API-equivalent, which is the only yardstick that compares two different arrangements of the same work.

Time is a wall clock around the whole arm, not the sum of the model calls.

03L — eight modules, the main result

armwallcostcallsauditnote
raw Claude Opus, one session303 s$1.2826 turns37/37the day before, the same goal took it 112 s and $0.61 — same-day is the fair row
Opus plans · DeepSeek Flash writes · QA off249 s$0.25337/37worker 185 s, 24 tools, first pass
Sonnet plans · Flash writes · QA off266 s$0.31535/37worker rejected for missing JSDoc on exports → 55 s retry → pass
Sonnet · Flash · QA on179 s$0.42436/37first pass · gate 13 s / $0.18
Opus · Flash · QA on311 s$0.54636/37worker retried once · gate 11 s / $0.16

The cheapest arm matches raw Opus on all 37 checks, 54 seconds faster, at a fifth of the price. One run per arm — not a best-of-N. Raw Opus moves between days by more than the gap between our arms, which is why the row above carries the previous day's figure instead of hiding it.

04S — one function, where the shape loses

armwallcostcallsnote
raw Claude Opus, one session24 s$0.185 turns
Sonnet plans · Flash writes · QA off44 s$0.123
Sonnet · Flash · QA on212 s$0.346the verifier caught 18.5 → "$18.5" and sent it back
Opus plans · Flash · QA off414 s$0.123planner 12 s · worker stalled 393 s — OpenCode tail latency on a one-line edit
Opus · Flash · QA on48 s$0.244gate 10 s / $0.12

A planner call is pure overhead on a one-line edit. This is the row the product had to obey rather than explain: on a one-file change the benchmark says don't merge, and the router sends the ask straight to one model. A benchmark that only ever agreed with the product would not be worth running.

05What the numbers actually say

The planner's model barely moves quality. Opus planning scored 37 and 36; Sonnet planning scored 35 and 36; raw Opus scored 37. The worker was the same cheap model in every antiloki row. What closes the gap is the verifier — and both of its rejections were real: missing JSDoc, and a money format that printed $18.5.

The bill is the planner. On the best L arm: planner ≈ $0.20, worker ≈ $0.02, verifier ≈ $0.03. Almost all of the saving comes from moving the writing off the frontier model, and almost all of the remaining cost is the thinking you kept there.

The QA gate measured nothing that day. Every gate run produced one placeholder claim and returned unproven in 10–13 s for $0.12–0.18. The always-branch was discarding the planner's acceptance lines, so the switch bought a bill and no verdict. It is on this page because it happened, and the QA-on rows in the table above are not slower because of QA — that is the worker's retry noise.

The worker's variance is the runtime, not the model. Same worker, same one-line task: 30 s once, 393 s another time.

06The two runs that failed, and why they were ours

arm (XXL, same day)wallcostresult
raw Opus795 s$5.30169 tests, build ✓
Opus plans · Sonnet writes, one worker921 s$2.64112 tests, build ✓
Sonnet · Haiku, a wave of 8failed 327 s$1.303 of 8 — Haiku's id generator could not pass the verifier twice
Sonnet · Flash, a wave of 8, floor offfailed 1 884 s$0.187 of 8 — the plan was a five-deep chain and nobody installed the test runner

Neither failure was the models'. The first was a plan whose dependency chain the projection did not fold into its wave count; the second was an orchestrator that let a goal ask for tests without installing a test runner. Both are fixed, and both are the reason this section exists: the interesting number in a benchmark is usually the one that went wrong.

07Running it yourself

Every arm above is one shell script against a fresh clone, with its own serve on its own port and its own ANTILOKI_HOME, so no arm inherits another's history or its learned rankings. The shape is always the same four steps:

1 · clone ministore into a scratch directory · 2 · start a serve with an empty home on a free port · 3 · post the goal with the arm's planner, worker and gate setting · 4 · when it finishes, run audit.mjs over the result and record wall time, dollars and the audit score in one row.

The same measurement runs inside the product: every run you do is scored on your machine and filed in a table that grows with use, which the router reads when it picks a merge for the next job. Ours ships as the starting point; your runs adjust its scores, by at most a third, and none of it leaves the machine.

← more from the antiloki blog  ·  the long version: Compiled, Not Trained →