Heimdall, the merge and the X-ray adapter: how antiloki builds a specialised coding model out of measured parts.
Second edition, measured 28–29 September 2026 on the five subscription seats — Claude Code, Codex, Gemini, Grok and Cursor. The first edition (22–23 September) measured Chinese models through OpenCode Zen; those arms are withdrawn and kept on disk as history. §4.2, §4.3, §7.1, §8 and §9 are rewritten; the depth runs still to come are listed in §10.
Third edition, 1 October 2026 — adds §11: the one tolerance (a seat wins if it keeps 80 % of the best measured quality, then price, then time), the solo path for asks of size L and smaller, a screening call of its own, obligatory QA, the QA judges' false-accept rates on two defect sets, the cost against a raw Claude session cell by cell (the 3× the site carried at the time of the third edition), and three X-ray experiments. It supersedes the 0.10 absolute band of §7.1, the planner and screening rows of Table 1, the “on request” QA of §5, the judge paragraph of §8.1 and the open items of §10 that are now done. Measured 2026-09-28 to 10-01; no Claude arm was re-run in this edition, and the Claude references date from 2026-09-22 to 28.
Fourth-edition revision, 1 October 2026 — the site’s headline reference is now raw Claude Opus and not the cheapest Claude arm inside the tolerance, because Opus is the Claude model a subscriber on the top plan reaches for. The headline is 10× on the eight-module task: the best measured cell, 11.3× rounded down, with a range of 6×–11× over the days Opus was measured (§11.7). The seven-cell analysis of §11.4, which gave the 3×, stays as the Sonnet-and-Haiku-reference result and as context; it is no longer the headline. The weekly-limit figure is restated as a paired task result (§11.4, §11.7).
A coding agent's behaviour is decided by three things: which model answers each kind of question, what that model is shown about the codebase, and how the whole is judged. antiloki fixes all three without training a weight. A merge assigns every job of a coding run — planning, editing, verifying, judging, reading, reporting — to one side of a brain × hands pair; a deterministic adapter conditions each call — X-ray's facts computed from the repository, plus the skills, specialists and tools the role carries — rather than on what the model remembers; and a benchmark re-weights every combination from measured outcomes, so the assignment learns. The result behaves like a new model, with its own score, price and speed. The central result of this edition is about where the money goes. At small size the writer decides quality: every pairing writing with Cursor Composer scored 1.00 on a calibrated sixteen-case task whatever planned it, and a frontier planner only multiplied the price, by 1.6–2.7× (raw Claude Haiku, writing alone, scored 0.87 on the same checks — the grader separates writers). At large size one cheap seat planning and writing its own work stayed within the quality band of the best pairing in all six domains, at $0.08–0.12 a run. The cheap seats are also the better screeners — Gemini and Haiku caught 87.5 % of deliberately under-specified prompts, Claude Sonnet 68.8 % — and the better verifier: Composer judged every must-pass and mutant case correctly. What a merge sells is a checked result at a cheap model's price; the expensive model is paid for only where a measurement says it earns it. We give every rule, formula, price and number the system uses, the results, and where the analogy to a trained merge breaks.
"Model merging" in machine learning means arithmetic on the weights of checkpoints that share a base (task vectors, TIES, DARE; mergekit is the tool). antiloki does none of that: the models are frozen, hosted by their vendors, reached through the vendors' own CLIs on your own subscription. What antiloki merges is behaviour: the answer to "which model does which job, shown what, judged how" — and that answer is learned from outcomes.
The analogy we defend is narrower and, we think, exact in practice. A LoRA adapter specialises a frozen model for a task by adding a small learned component in front of it. X-ray is a small deterministic component in front of every call that supplies what a coding fine-tune would otherwise have to learn — the structure, laws, risks and history of this repository — and it is recomputed at each commit instead of going stale. The merge table plays the role of the training loop: an outcome-weighted selection over roles, updated by every benchmark row and every audited use. The difference from a trained merge is the ceiling: no role can exceed what its base model can do when shown the right evidence. For code, where the missing knowledge is almost always the codebase and not the craft, that ceiling has not been the binding constraint in any measurement below.
brain × hands. The brain plans, verifies, judges, reads and screens; the hands edit, write tests and report. Each side is one engine and, for OpenCode, one model.
X-ray: deterministic code intelligence — reviews scored axis by axis, the laws, the graph, blast radius, coverage gaps — packed into every call's context.
the benchmark: every run and every audited call becomes a row per job × size × planner × worker × QA, re-weighted by quality, cost and time.
Query routing between a strong and a weak model is well studied. RouteLLM (LMSYS, 2024) trains routers on Chatbot Arena preferences and reports cost reductions of over 85% on MT Bench at 95% of GPT-4's quality [1]. Commercial routers — Martian [2], Not Diamond [3], and the OpenRouter aggregator now processing 25 trillion tokens a week [4] — sell the same lever per request. Heimdall differs in what it routes over: not a request, but a job inside a run, with the job's own acceptance and its own measured winner.
The brain × hands split is aider's architect/editor mode [5]: a reasoning model describes the change, a cheaper model writes the edit. aider reported a new state of the art on its polyglot benchmark with DeepSeek R1 as architect and Claude Sonnet as editor at 14× less cost than the previous o1 result [6]. Cline and Roo Code let users bind different models to Plan and Act modes by hand [7]. antiloki generalises the split from two roles to ten jobs and makes the binding measured rather than chosen.
That a composition of models can behave as a new, better model was shown by Mixture-of-Agents (Together, 2024): layered open models reached 65.1% on AlpacaEval 2.0 against GPT-4o's 57.5% [8]. The "compound AI systems" position from Berkeley [9] argues the unit of capability is now the system. We take both as the frame: the merged agent is the model; its card is its benchmark row.
Every call antiloki makes on a coding run belongs to one of ten jobs. A job has a side (brain or hands), a prompt template, and — the point of the system — its own winner and fallback per engine and globally.
| job | side | what the call does | fires when |
|---|---|---|---|
| planner | brain | screens the ask and writes the plan: tasks that own disjoint files, each with acceptance lines | a goal of two or more enumerated tasks or over 600 characters; an ask of size L or smaller skips it (§11.2) |
| screening | brain | the blocking/assumable questions before a plan | before the planner, as a call of its own on the best screening seat (§11.2) |
| editor-s | hands | a change to one file, an ask under 400 characters | a task that owns one file |
| editor-l | hands | a change across files or a long ask | every other task |
| verifier | brain | checks a task's diff against its acceptance; may send it back once | after every task |
| test-writer | hands | writes the tests a contract's acceptance lines imply | blueprint mode |
| qa-gate | brain | checks each acceptance claim against the final diff; the isolated judge | every run — QA is obligatory (§11.2) |
| report | hands | consolidates the tasks' reports into one | a multi-task run (a single task's report stands) |
| lookup | brain | a factual question about the code: where, which, who | a chat turn the classifier reads as a lookup |
| synthesis | brain | why, how, compare, in which order — an answer that reasons | a chat turn the classifier reads as synthesis |
The classifier is deterministic (cli::classify_ask): a question mark or an interrogative first word makes a read; among reads, the words why, how, explain, describe, compare, difference, trade-off, should, in which order make it synthesis, else lookup. A write that names at most one file and is under 400 characters is editor-s, else editor-l. Sizes for the benchmark follow the plan, not the ask: two or more tasks is XXL; one task owning one file with a goal under 400 characters is S; else L.
For a job on a run, Heimdall produces a binding (engine, model) and a fallback, in this order:
job_binding_paired.job_binding_learned, merges::best_for(job, size, 3).which), else the run engine's own cell. job_binding_any.worker_under_planner walks the worker down its ladder until its list price is under the planner's, and logs why.run_job for every other job.| job | winner | fallback | the measurement behind it |
|---|---|---|---|
| lookup | opencode · deepseek-v4-pro | claude-code · haiku | readers /50: Sonnet 42, Pro 41 at $0.13, Qwen 38, Haiku 36 (strong on lookup, fabricates on synthesis) |
| synthesis | claude-code · sonnet | opencode · qwen3.6-plus | synthesis /20: Sonnet 16, Qwen 14, Pro 13 — outside the band, Sonnet keeps it |
| screening | opencode · deepseek-v4-flash | claude-code · haiku | part of the planner call; cheap by design |
| planner | opencode · deepseek-v4-pro | claude-code · sonnet | the planner's model did not move L quality (Opus 37, 36 · Sonnet 35, 36); Pro is within the band and 5× cheaper |
| editor-s | opencode · qwen3.6-plus | claude-code · sonnet | S ×5 raw: Qwen 5/5 · 34 s · $0.011; Sonnet 5/5 · $0.093; Pro 4/5; Flash 2/5 |
| editor-l | opencode · deepseek-v4-flash | opencode · deepseek-v4-pro | L one-worker: Flash under a planner 36–37/37 at $0.02 of worker cost |
| verifier | claude-code · sonnet | opencode · deepseek-v4-pro | both L rejections it made were real (missing JSDoc; a wrong money format) |
| test-writer | claude-code · sonnet | opencode · deepseek-v4-pro | blueprint runs — being measured |
| qa-gate | claude-code · opus | opencode · deepseek-v4-pro | the judge is the one place the dearest model is bound by default |
| report | opencode · deepseek-v4-flash | claude-code · haiku | a consolidation; cheap by design |
The three-point rule. Our rule for every table above: when two candidates are within three points of each other on the job's score, the cheaper one is the winner. Three points of fifty is 0.06 on the unit scale, and that number reappears as the band in §7.
Every cost antiloki reports is computed, never read from a vendor's estimate: input and output tokens times the model's list price per million, with cache reads at one tenth of input and cache writes at 1.25× (cli::price_tokens). The same table gives the ladder each engine's effort dial walks.
| model | input $ | output $ | model | input $ | output $ |
|---|---|---|---|---|---|
| claude opus | 5.00 | 25.00 | gemini pro | 1.25 | 10.00 |
| claude sonnet | 3.00 | 15.00 | gemini flash | 0.14 | 0.28 |
| claude haiku | 1.00 | 5.00 | grok 4.6 | 3.00 | 15.00 |
| codex / gpt-5 | 1.25 | 10.00 | grok code fast | 0.20 | 1.50 |
| cursor composer 2.5 | 0.50 | 1.50 |
Each call is also timed on the serve's own clock and its tokens counted from the engine's stream; the triple (cost, ms, tokens) lands on the pane's timeline as an audit line and in the ai_calls table, and — through merges::record_use — as a benchmark row of size call.
What "saved" is measured against (since 2026-09-28): the same tokens priced at the frontier model's list rate — Claude Opus, or Codex's frontier model on a Codex seat (cli::frontier_of) — against what the work cost at the seats that actually ran it, each task priced at its own engine and model. Before that the yardstick was the seat's own top rung, which for four of the five seats is the same rate as its cheap rung and printed "saved $0.00" on work five times cheaper than Opus. A row nobody priced is priced from list rates (merges::estimated_usd); a replayed plan is never counted as a free planner.
A merged agent in Studio is a recipe with a pair: a name, a brain side, a hands side, a theme, a scope, its specialists and skills. Since 2026-09-27 every side is one of the five subscription seats — Claude Code, Codex, Gemini, Grok and Cursor; the OpenCode/Zen pairs of the first edition (Claude × DeepSeek, Claude × Kimi, Codex × DeepSeek …) are withdrawn and kept on disk only as history. A seat may sit on both sides (Composer plans and writes), which is measured as a pairing like any other. Heimdall's chip on a pane shows the mode and the money the router saved on that pane, computed as API-equivalent dollars against one high-tier session doing the same tokens.
A goal run is the unit the benchmark measures. Its stages, and the model that answers each, are what the merge assigns.
The projection. Before a plan fans out, antiloki estimates whether a wave beats one worker. A single session doing k parts is projected at SERIAL_FACTOR = 1.5 times the parallel per-part time; the wave's concurrency is bounded by the depth of the plan's dependency chain (effective_concurrency = ceil(fills / chain), clamped to the configured cap); and a plan whose single-session projection is under the parallel floor — 900 s by default, ANTILOKI_PARALLEL_FLOOR_S, 0 = off — collapses to one worker. With no history the medians assumed are 120 s per fill and 150 s of orchestration. This is why every S and L row below is a one-worker run: the wave only pays above fifteen minutes of projected serial work.
The verifier's one send-back. A task the verifier rejects goes back to its worker once with the rejection as the brief; a second rejection fails the task. Both rejections observed on L were real defects a compile gate would not catch.
The QA gate. An isolated judge checks each acceptance claim against the final diff and answers met / unmet / unproven; on a one-worker run the always-branch kept a placeholder claim, so the gate reported unproven. That gap is closed in the third edition: QA is obligatory, runs the deterministic checks (§11.2), sends unmet claims back to their workers a bounded number of times and, when those are spent, turns the run into a question for the person.
A LoRA specialises a frozen model by adding a small learned component in front of it. antiloki's adapter is the whole conditioning layer around frozen models, and it is larger than the router: which model answers each job, the brain × hands split, the skills materialised into the agent's box, the specialists it may consult, the tools it may drive, and the deterministic intel it boots with. A Designer is not a different model from a Backender — it is a different adapter.
a merge binds every job to a side: the brain plans, verifies, judges, reads; the hands edit, test, report. Per job, per domain, measured.
skills materialised into its box, specialists it may consult, and the X-ray packs it reads before it explores.
the panes it may open — Frames and the design system for a Designer, X-ray and the constitution for a Reviewer, the runner for a Tester.
X-ray is antiloki's analysis of the repository. Nothing in it is generated by a model; all of it is computed from the files, the ASTs, the tests and the git history, and it is what every call is shown before it explores.
any, no var, no loose equality, no debugger, no @ts-ignore, no swallowed errors), with what compliance is worth.laws · folder · file · graph · tests · blast · doc items, derived from the pane's scope by default, under a character budget (200 000).chat_auto::compose_for picks from the workspace's intel and prior runs — the chat's own log excluded — under an 8 000-character cap.This is why the analogy in §1 holds for code specifically. What a repository-specific fine-tune would teach — where things are, what the conventions are, which files are dangerous — is the content of the pack, and the pack is true at the commit the agent is working on.
A job is a stage inside one run; nobody hires a verifier. A role is what a person hires — a domain of the codebase — and inside it every one of the ten jobs still happens. So a role is not an eleventh job; it is a profile over all ten, scoped to a domain, carrying its own adapter.
| role | domain | consults | drives |
|---|---|---|---|
| Backender — Database, API under it | backend | correctness · errors, types · architecture · dependencies · security · injection | code graph, code, diff |
| Frontender — Designer, Architect under it | frontend | framework · idioms · style · suppressions | components, code, diff |
| Frontend designer | design | category · style · axis · cohesion | Frames, design system, components, canvas |
| Tester | tests | testing · isolation · axis · tests | run, X-ray, diff |
| Reviewer | mixed | axis · defects, compliance, change · book · warden | X-ray, constitution, diff — no editor |
| Product owner | docs | book · definer, questioneer, synthesizer | intent, todo, kanban, context |
| Infra | infra | security · secrets · book · warden | terminal, run, git graph |
The domain is derived from the paths a task owns, deterministically (roles::domain_of): a test of the
API is test work; the styles folder beats the extension; a change touching six backend files and one stylesheet is
still backend, and an even split is honestly mixed. So every run tags itself, and runs that happened
before the dimension existed can be attributed without re-running them.
A role asks the table for the merge that won its kind of work. When its domain has no samples it reaches in order — its nearest neighbour (design looks at frontend; infra and docs at backend), then the shared pool, then everything pooled — so a role always has an answer and the sample count says how far it had to reach.
The same rows, three questions. Cheaper counts the bill above the score (0.35 / 0.55 / 0.10); Best is the table's own weight (0.70 / 0.20 / 0.10); Smarter counts only quality and allows no gap at all below the best. Each returns a different merge, and the three-point band still holds for the first two: a row ten points behind cannot be bought back at any price. They are the same three stops as the effort dial on the pane.
What "training" a merge means here. No weights change, so a new pair — Codex × Gemini, say — cannot be fine-tuned. Three things do adapt: it inherits its sides' measured priors rather than starting blank; it learns from its own use, since every audited call and run files a row and the learned table overrides the static one at three samples in its domain; and the handoff between the sides — the brief the brain writes, how much diff the verifier is shown, the acceptance lines it is judged on — is a set of deterministic parameters that can be measured and tuned per pair. That last one is prompt-level adaptation with frozen weights: the only honest sense of training available, and the one this system has not yet done.
The merge table (~/.antiloki/merges.sqlite, merge_scores) has one row per job × size × planner × worker × qa. Sizes are s · l · xxl for runs and call for audited single calls. A row accumulates:
where audit is the external behavioural score a bench posts after the run (L: 37 checks of the eight modules; S: a five-case check of the function the ask names) and source says who measured it: bench (antiloki's own runs, shipped as the seed), run (the user's goal runs), use (audited calls). A user's table is therefore the shipped base plus their own rows, weighted together; nothing is collected back.
Every record re-weights all rows of that job and size against each other (merges::reweigh, merges::rank):
Price first, speed last, quality as a gate. Third edition: the gate is now relative — a seat or pair wins if it keeps at least 80 % of the best measured quality (Objective::band() = 0.20), and measured quality outranks claimed; there is no second tolerance anywhere in the product (§11.1). The band disqualifies a materially worse answer outright; inside it, the bill decides. The first edition weighted 0.7 / 0.2 / 0.1 with a 0.06 band, which was narrower than the instrument — a ten-instance bench resolves quality in steps of 0.10 — and would pay 6.7× for one instance in ten. Speed was reduced from 0.30 to 0.10 on 2026-09-28, after the fastest frontier planner won a simple frontend task that a model a tenth of its price had measured the same on. A pairing measured on several suites is pooled by samples, never overwritten; a row nobody priced is priced from list rates; a planner row is never cheaper than its own list price.
A job's winner is its best sized row; a job the benches never sized falls back to its best audited-call row with at least two samples; else it has no winner yet, and the Studio says so. Each winner is named by composition and role — the brain's brand and the hands' suffix as one word, then the job's title: CodexComposer Builder, Composer Tinker, ClaudeComposer Architect — and can be made an agent in one click, carrying its pair. For solo runs the learned table outranks Table 2 once a row has three samples, which is how the static table is meant to retire.
Measured 28–29 September 2026 on the five subscription seats, one machine, a pinned fixture (ministore at 3de696c), every arm in a fresh clone with its own plan (plan replay cleared per arm). Quality is the deterministic grader's score for that task — calibrated before any arm ran, never a model's opinion; dollars are the API-equivalent of every call the serve made for the run (planner, writer, verifier, QA), priced from the tokens each seat reported (§4.2). The tables are generated from the record by bench/campaign/make-results.py; rows under three samples are marked and never ranked.
| planner × writer | n | quality | $/run | s/run | priced |
|---|---|---|---|---|---|
| Codex medium × Composer 2.5 * | 2 | 1.00 | 0.1299 | 256 | before QA · 09-26/28 |
| Gemini × Composer 2.5 * | 2 | 1.00 | 0.1384 | 132 | before QA · 09-26/28 |
| Claude Sonnet × Composer 2.5 * | 2 | 1.00 | 0.1500 | 118 | before QA · 09-26/28 |
| Composer 2.5 × Composer 2.5 | 3 | 0.97 | 0.1017 | 95 | with QA · 09-29 |
* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.
| planner × writer | n | quality | $/run | s/run | priced |
|---|---|---|---|---|---|
| Codex terra × Gemini Flash * | 2 | 1.00 | 0.0363 | 271 | before QA · 09-26/28 |
| — none (raw) × Grok 4.6 medium * | 2 | 1.00 | 0.0837 | 260 | before QA · 09-26/28 |
| Composer 2.5 × Composer 2.5 | 3 | 1.00 | 0.1133 | 157 | with QA · 09-29 |
| Claude Sonnet × Claude Haiku * | 2 | 0.98 | — | 144 | before QA · 09-26/28 |
| Grok 4.6 high × Composer 2.5 * | 2 | 0.98 | 0.0377 | 156 | before QA · 09-26/28 |
| Codex terra × Composer 2.5 * | 2 | 0.98 | 0.0378 | 168 | before QA · 09-26/28 |
| Claude Opus × Grok 4.6 medium * | 2 | 0.98 | 0.2766 | 147 | before QA · 09-26/28 |
| — none (raw) × Composer 2.5 * | 2 | 0.93 | 0.0319 | 107 | before QA · 09-26/28 |
| Claude Sonnet × Composer 2.5 * | 2 | 0.80 | 0.0318 | 137 | before QA · 09-26/28 |
* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.
| planner × writer | n | quality | $/run | s/run | priced |
|---|---|---|---|---|---|
| Codex terra × Gemini Flash * | 2 | 1.00 | 0.0143 | 103 | before QA · 09-26/28 |
| Codex terra × Composer 2.5 * | 2 | 1.00 | 0.0157 | 46 | before QA · 09-26/28 |
| Grok 4.6 high × Composer 2.5 * | 2 | 1.00 | 0.0158 | 87 | before QA · 09-26/28 |
| — none (raw) × Composer 2.5 * | 2 | 1.00 | 0.0167 | 37 | before QA · 09-26/28 |
| Claude Sonnet × Composer 2.5 * | 2 | 1.00 | 0.0211 | 74 | before QA · 09-26/28 |
| — none (raw) × Grok 4.6 medium * | 2 | 1.00 | 0.0264 | 78 | before QA · 09-26/28 |
| Claude Sonnet × Claude Haiku * | 2 | 1.00 | 0.0740 | 95 | before QA · 09-26/28 |
| Composer 2.5 × Composer 2.5 | 3 | 1.00 | 0.1173 | 93 | with QA · 09-29 |
| Claude Opus × Grok 4.6 medium * | 2 | 1.00 | 0.1248 | 57 | before QA · 09-26/28 |
* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.
| planner × writer | n | quality | $/run | s/run | priced |
|---|---|---|---|---|---|
| Codex terra × Gemini Flash * | 2 | 1.00 | 0.0209 | 57 | before QA · 09-26/28 |
| Codex terra × Composer 2.5 * | 2 | 0.98 | 0.0131 | 42 | before QA · 09-26/28 |
| Claude Opus × Grok 4.6 medium * | 2 | 0.98 | 0.1022 | 46 | before QA · 09-26/28 |
| Grok 4.6 high × Composer 2.5 * | 2 | 0.96 | 0.0124 | 54 | before QA · 09-26/28 |
| Claude Sonnet × Composer 2.5 * | 2 | 0.96 | 0.0135 | 55 | before QA · 09-26/28 |
| — none (raw) × Grok 4.6 medium * | 2 | 0.96 | 0.0305 | 65 | before QA · 09-26/28 |
| Claude Sonnet × Claude Haiku * | 2 | 0.96 | 0.0570 | 105 | before QA · 09-26/28 |
| Composer 2.5 × Composer 2.5 | 3 | 0.96 | 0.0810 | 83 | with QA · 09-29 |
| — none (raw) × Composer 2.5 * | 2 | 0.94 | 0.0112 | 27 | before QA · 09-26/28 |
* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.
| planner × writer | n | quality | $/run | s/run | priced |
|---|---|---|---|---|---|
| — none (raw) × Composer 2.5 * | 2 | 1.00 | 0.0070 | 47 | before QA · 09-26/28 |
| Claude Sonnet × Composer 2.5 * | 2 | 1.00 | 0.0076 | 44 | before QA · 09-26/28 |
| Codex terra × Composer 2.5 * | 2 | 1.00 | 0.0076 | 36 | before QA · 09-26/28 |
| Grok 4.6 high × Composer 2.5 * | 2 | 1.00 | 0.0081 | 51 | before QA · 09-26/28 |
| — none (raw) × Grok 4.6 medium * | 2 | 1.00 | 0.0129 | 32 | before QA · 09-26/28 |
| Codex terra × Gemini Flash * | 2 | 1.00 | 0.0206 | 50 | before QA · 09-26/28 |
| Claude Sonnet × Claude Haiku * | 2 | 1.00 | 0.0500 | 79 | before QA · 09-26/28 |
| Claude Opus × Grok 4.6 medium * | 2 | 1.00 | 0.0606 | 43 | before QA · 09-26/28 |
| Composer 2.5 × Composer 2.5 | 3 | 1.00 | 0.0800 | 68 | with QA · 09-29 |
* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.
| planner × writer | n | quality | $/run | s/run | priced |
|---|---|---|---|---|---|
| Claude Sonnet × Claude Haiku * | 2 | 0.94 | 0.0580 | 75 | before QA · 09-26/28 |
| — none (raw) × Composer 2.5 * | 2 | 0.88 | 0.0122 | 37 | before QA · 09-26/28 |
| Codex terra × Composer 2.5 * | 2 | 0.88 | 0.0136 | 49 | before QA · 09-26/28 |
| — none (raw) × Grok 4.6 medium * | 2 | 0.88 | 0.0180 | 50 | before QA · 09-26/28 |
| Grok 4.6 high × Composer 2.5 * | 2 | 0.88 | 0.0189 | 100 | before QA · 09-26/28 |
| Codex terra × Gemini Flash * | 2 | 0.88 | 0.0248 | 85 | before QA · 09-26/28 |
| Composer 2.5 × Composer 2.5 | 3 | 0.88 | 0.0830 | 105 | with QA · 09-29 |
| Claude Sonnet × Composer 2.5 * | 2 | 0.88 | 0.1030 | 63 | before QA · 09-26/28 |
| Claude Opus × Grok 4.6 medium * | 2 | 0.88 | 0.2668 | 71 | before QA · 09-26/28 |
* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.
| planner × writer | n | quality | $/run | s/run |
|---|---|---|---|---|
| — none (raw) × Gemini Flash | 4 | 1.00 | 0.0445 | 230 |
| — none (raw) × Composer 2.5 | 4 | 1.00 | 0.0488 | 25 |
| Composer 2.5 × Composer 2.5 | 8 | 1.00 | 0.0860 | 66 |
| Codex terra × Composer 2.5 | 8 | 1.00 | 0.1260 | 54 |
| Claude Sonnet × Composer 2.5 | 16 | 1.00 | 0.1400 | 48 |
| Grok 4.6 high × Composer 2.5 | 8 | 1.00 | 0.2320 | 104 |
| — none (raw) × Grok 4.6 medium | 4 | 1.00 | 0.2489 | 30 |
| — none (raw) × Claude Haiku | 4 | 0.87 | 0.0437 | 25 |
Calibrated before any arm ran: a correct reference 16/16, rounding the discount 14/16, case-sensitive codes 14/16, naive floats 7/16, no change 0/16.
| seat | screens | balanced accuracy | $/screen | s/screen |
|---|---|---|---|---|
| Claude Haiku | 36 | 0.94 | 0.0142 | 9 |
| Gemini | 36 | 0.94 | 0.0061 | 25 |
| Codex low | 36 | 0.91 | 0.0188 | 4 |
| Grok 4.6 medium | 35 | 0.87 | 0.0590 | 21 |
| Composer 2.5 | 34 | 0.86 | 0.0145 | 15 |
| Claude Sonnet | 36 | 0.84 | 0.0412 | 5 |
Balanced accuracy = (recall on damaged prompts + precision on clear ones) / 2. The per-seat recall and precision are in bench/screen-solo/out.
| job · domain · seat | cases | accuracy | $/call |
|---|---|---|---|
| qa-gate · backend · Composer 2.5 | 6 | 0.83 | 0.0457 |
| qa-gate · backend · Claude Haiku | 5 | 0.80 | 0.0461 |
| qa-gate · backend · Codex low | 6 | 0.50 | 0.0899 |
| qa-gate · docs · Codex low | 6 | 0.67 | 0.0823 |
| qa-gate · docs · Composer 2.5 | 6 | 0.67 | 0.1483 |
| qa-gate · infra · Claude Haiku | 4 | 1.00 | 0.0296 |
| qa-gate · infra · Composer 2.5 | 6 | 1.00 | 0.0474 |
| qa-gate · infra · Codex low | 6 | 0.50 | 0.0765 |
| verifier · backend · Composer 2.5 | 6 | 1.00 | 0.0140 |
| verifier · backend · Codex low | 6 | 0.67 | 0.0340 |
| verifier · backend · Claude Haiku | 6 | 0.50 | 0.0455 |
| verifier · docs · Codex low | 6 | 1.00 | 0.0266 |
| verifier · docs · Composer 2.5 | 6 | 1.00 | 0.0110 |
| verifier · docs · Claude Haiku | 6 | 0.67 | 0.0389 |
| verifier · infra · Claude Haiku | 6 | 1.00 | 0.0283 |
| verifier · infra · Codex low | 6 | 1.00 | 0.0238 |
| verifier · infra · Composer 2.5 | 6 | 1.00 | 0.0093 |
Cases: 3 must-pass artefacts and 3 mutants per domain (files the ask names cut in half, re-graded below 0.90). Frontend, design and mixed have no honest must-fail case yet — see §10.
| planner × writer | n | resolved | $/run | s/run |
|---|---|---|---|---|
| — none (raw) × Claude Sonnet | 10 | 0.90 | 0.3737 | 111 |
| Claude Sonnet × Composer 2.5 | 10 | 0.80 | 0.1187 | 161 |
| — none (raw) × Composer 2.5 | 14 | 0.79 | 0.2018 | 85 |
default model as a third-party one and silently re-seated a Gemini planner on Claude Sonnet; those runs are filed as Sonnet (which is what planned them), Gemini as an S planner is unmeasured, and a bench arm whose pair would be re-seated is now refused instead.run, use) and never averaged with benched ones: observation moves a benched score by at most a third, it never sets one.Everything here was measured or decided after the second edition's rows were written. The record is the same merge table (§7); a row is only counted once its quality has been audited by an executed check.
The rule, 2026-09-30: a seat or a pair wins if it keeps at least 80 % of the best measured quality in its job and size; among those, the cheapest wins; among equally cheap ones, the fastest. Measured quality outranks claimed quality. The earlier 0.10 absolute band (§7.1) is retired, and with it the per-job and per-role pickers in the app: the choice of AI for a job is automatic, and Settings keeps only the providers, the main provider and a switch for each seat. A seat that is switched off, signed out, or at 95 % or more of its quota window is skipped before the job starts and the work moves to the next seat in the measured ladder.
raw × editor. A planner is still used for an ask with two or more enumerated tasks or over 600 characters. At S and L the router compares measured solo cells against planner × worker pairs under the one tolerance. On a probe over a copy of the live record, solo Composer at editor-s kept 98 % of the best measured quality at $0.016 a run against $0.086 for the cheapest pair inside the band; documentation stays a pair. Only the backend domain is measured at editor-s, and no end-to-end run was watched binding a solo seat.A false accept lets a broken change through, so the QA gate is chosen on it, not on price. Two defect sets, each a broken run or a good run, graded by an executed check. The first (10 broken and 10 good per class, five classes): Opus 0 of 50 false accepts, Grok-4.6-high 0, Composer 2.5 0, Gemini 0, Codex 0 (but it rejects 39–41 of 50 good runs), Sonnet 8 of 50, Haiku 6 of 50. The first set is easy — five seats scored zero — so it does not separate them, and Gemini’s zero there does not carry over: on the second set it is the worst seat measured (below). The second (15 broken and 15 good per class, five classes) does:
| seat | false accepts, all classes | without the frontend class | false rejects |
|---|---|---|---|
| Claude Opus | 14 / 75 (19 %) | 0 / 60 | 26 / 75 (35 %) |
| Grok-4.6-high | 18 / 75 (24 %) | 4 / 60 | 8 / 71 (11 %) |
| Cursor Composer 2.5 | 20 / 75 (27 %) | 6 / 60 | 14 / 73 (19 %) |
| Claude Sonnet | 26 / 74 (35 %) | 12 / 59 | 18 / 75 (24 %) |
| Claude Haiku | 27 / 75 (36 %) | 19 / 60 | 21 / 74 (28 %) |
| Gemini flash (medium) · incomplete | 27 / 45 readable (60 %): backend 14 / 15, frontend 13 / 15 | not measured | 2 / 50 · 55 unread (docs 25, infra 30) |
| Gemini default · incomplete | 15 / 15 on backend (100 %); 113 unread | not measured | 0 / 22 |
Paired sign tests on the cases where two seats disagreed: Opus against Grok-high 0 : 4 (p = 0.125), Opus against Composer 0 : 6 (p = 0.031), Grok-high against Composer 4 : 6 (p = 0.75), Grok-high against Sonnet 2 : 10 (p = 0.039), Opus against Sonnet 0 : 12 (p = 0.0005); no correction for multiple tests. An earlier ranking in tiers and an unpaired test were withdrawn after this paired recount.
Revised the same day, after a fact-check. The first version of this section printed 10.7× and the site said “10×”. It compared each winner with the Claude arm that made the ratio largest, it counted a 95 % interval as if the four S cells were independent, and it set raw single-seat sessions, which carry no QA, against raw Claude while the product ends every run in QA. This is the corrected section; the site said 3× on this basis until the fourth-edition revision (§11.7).
A cell is one job × size × domain × suite that holds a raw Claude reference with audited quality and at least one non-Claude arm with at least two audited samples. The winner of each cell is picked by the router’s own rule (§11.1), five subscription seats only. The reference is the cheapest raw Claude arm (Haiku, Sonnet or Opus) that keeps at least 80 % of the cell’s best measured quality, so that the ratio is not made larger by choosing a dearer Claude; the ratio is the reference’s dollars per run divided by the winner’s, API-equivalent list prices.
| cell | reference | winner | quality | × cheaper |
|---|---|---|---|---|
| S backend · plain | Haiku $0.0364 (n=3) | Composer alone $0.0048 (n=3) | 1.00 vs 1.00 | 7.6 |
| S backend · edges | Haiku $0.0509 (n=3) | Composer alone $0.0057 (n=3) | 1.00 vs 1.00 | 8.9 |
| S backend · branches | Sonnet $0.0941 (n=3; Haiku scores 0.73, below 0.8 of 1.00) | Composer alone $0.0062 (n=3) | 0.87 vs 0.87 | 15.3 |
| S backend · contract | Haiku $0.0464 (n=3) | Composer alone $0.0090 (n=3) | 1.00 vs 1.00 | 5.2 |
| L backend · eight modules | Sonnet $0.330 (n=1) | Codex low alone $0.1215 (n=1, 2 audits) | 0.97 vs 0.97 | 2.7 |
| L backend · fix | Haiku $0.0577 (n=3) | Composer alone $0.0150 (n=3) | 0.93 vs 0.90 | 3.8 |
| L test-writer · SWT-bench | Sonnet $0.3737 (n=10) | Gemini × Composer $0.1041 (n=10) | 0.80 vs 0.90 | 3.6 |
| geometric mean, before the QA check | 5.7 | |||
| the same with the QA step added (estimate) | 3.2 |
What the number is. The median is 5.2× and the cells run from 2.7× to 15.3×; dropping any one cell moves the geometric mean between 4.9× and 6.5×. No interval is printed: the four S cells are one synthetic small-ask family (twelve instances) where every arm scores about 1.0, they are not independent, and a bootstrap that treats them as seven cells is narrower than the evidence (the first version printed 7.0–15.2×). Without the family there are three cells, 2.7×, 3.8× and 3.6×, geometric mean 3.3×. Weighted by what a bill would be (dollars summed over the runs of each cell) the saving is 3.9× over the six cells with n of at least 3.
What the number leaves out: QA. Six of the seven winners are one seat running alone, filed as raw with no QA check, and the Claude arm they are set against has none either; only SWT-bench uses an antiloki pair (Gemini × Composer, 3.6×). The product ends every run in QA, and QA costs money: on the S2 suite raw Composer is $0.0488 a run and Composer × Composer with its QA check $0.086, 1.76× more; on the eight-module task Composer alone is $0.0671 and Composer × Composer × QA $0.1017, 1.5×. Dividing the 5.7× by 1.76 gives 3.2×. That division is an estimate, not a measurement — no end-to-end run was watched binding a solo seat (§11.2, §11.6), and the QA overhead of the other cells was not measured. The one cell where the whole pipeline was priced is the eight-module task: $0.1017 against raw Sonnet at $0.330, 3.2×. The 3× the site carried on this basis was that estimate rounded down; it sat under the 5.7× of the QA-free comparison, and under the 3.3× of the three cells outside the synthetic family.
Limits. n is 3 per arm in five cells and 10 in SWT-bench; the eight-module cell is one Codex run (two audits) against one Sonnet run (the first version used Opus, mean of 2, $1.151, which is inside the tolerance but not the cheapest Claude arm that is). The same-quality swap on SWT-bench (Codex × Composer, 9 of 10 like raw Sonnet, $0.157) is 2.4×; the 3.6× above is Gemini × Composer, which scored 8 of 10 (merges::separated is false at n=10, so “same quality” means inside the tolerance and not separated at these samples, not identical). The 6.7× of the first headline was raw Qwen3.6-plus ($0.0558) against raw Sonnet ($0.3737) on that same SWT cell — not an antiloki arm, a metered model that is off by default — and is withdrawn; the five-seat equivalent is the 3.6× above. Composer, Gemini and Codex dollars are estimated from reported tokens at list rate (Cursor and Antigravity report tokens, not dollars); the Claude references ran 2026-09-22 to 28. The eight-module row prices Codex low alone ($0.1215, 2.7×), the winner when the table was derived; the five-seat re-run of the same task later measured Composer alone at $0.0671 (4.9× under raw Sonnet) and the QA pipeline at $0.1017, and with Composer alone in that row the geometric mean before QA would be 6.2×. S2 and the S cells also disagree: raw Haiku $0.0437, Composer alone $0.0488 and Composer × Composer with QA $0.086 on S2, while Composer alone is priced about ten times lower in the S cells for asks of similar size (cause not found; Composer dollars are token estimates at list price). Time is not a headline axis: matched cells are about equal, except the eight-module cell, where Composer alone took 61 s against Opus’s 278 s (the Opus time moves between 112 s and 303 s from day to day).
The weekly limit. The site states +112 %, and restates what it is (fourth-edition revision). It is the paired task result of §11.7 — on the eight-module task, $0.1554 for the arm where Claude plans and judges against $0.3297 for raw Sonnet, 2.12× — and not a measured percentage of anyone’s weekly limit. The earlier derivation, 1 / (1 − 0.528) for the share of one machine’s tokens that ran off the frontier seat, is withdrawn: an independent audit could not reproduce the share, and the cuts of the record differed (+471 % over the whole mix, +106 %, +118 % and +122 % over three others). By dollars the gain over the recorded mix is +33 % to +45 %, which is the range to quote for a mix of work and not for a task.
The eight-module task, re-run on the five seats (37 checks, 2026-09-28/29; the page’s table): Composer alone $0.067, 61 s, 36.3 of 37 (n=3); Composer × Composer with a Composer QA check $0.102, 95 s, 36.0 (n=3); Sonnet × Composer $0.155, 106 s, 36.8 (n=5); Codex × Composer $0.165, 137 s, 36.7 (n=3); Gemini × Composer $0.184, 134 s, 36.3 (n=5); Grok-high × Composer $0.292, 199 s, 36.3 (n=3); raw Opus $1.151, 278 s, 37 (n=2, the 2026-09-22/23 reference); raw Sonnet $0.330 (n=1, 0.973).
Each experiment crossed three agents (Cursor Composer 2.5, Codex, Gemini via Antigravity) with a small and a wide change and three scenarios — the repository only, the X-ray intel only with the agent told not to explore, or both. One run per cell; graded statically, nothing compiled; Gemini reports no tool-call telemetry.
How the headline is defined: the best measured antiloki result against raw Claude Opus on the same task and the same checks, not against the cheapest Claude model. The reasoning is that Opus is the Claude model a subscriber on the top plan reaches for. All figures below come from the merge record (merge_scores) and are the eight-module task of §11.4 (37 checks, ministore); the row ids are in the merge record.
| arm | n | checks | $/run | s/run | × under Opus | × under Sonnet |
|---|---|---|---|---|---|---|
| raw Claude Opus · 2026-09-22/23 | 2 | 37.0 / 37 | 1.151 | 278 | — | — |
| raw Claude Sonnet · 2026-09-23 | 1 | 36 / 37 | 0.3297 | 121 | 3.5 | — |
| Composer alone, no QA step | 3 | 36.3 / 37 | 0.0671 | 61 | 17.2 | 4.9 |
| Composer × Composer, with QA — the headline arm | 3 | 36.0 / 37 | 0.1017 | 95 | 11.3 | 3.2 |
| Sonnet plans, Composer writes, Sonnet judges | 5 | 36.8 / 37 | 0.1554 | 106 | 7.4 | 2.1 |
The headline is the pipeline with its QA step: 11.3×, rounded down to 10×. Its quality is 0.973 of 1.0, inside the one tolerance of §11.1 and not equal to raw Opus; it passes 36.0 of the same 37 checks on average. Opus is not stable from day to day: the ladder run of 2026-09-22 recorded $0.61 for the same task, and against that run the pipeline is about 6.0×, so the honest range is 6×–11× and the headline is the best measured cell. The weekly-limit figure, +112 %, is the Sonnet-planned arm against raw Sonnet, $0.1554 against $0.3297 = 2.12×; with n=5 against n=1 and list-price dollars standing in for an allowance the vendors do not publish, it is a paired task result and not a share of a quota window.
The headline arm is the nearest measured arm to the shipped path, not the shipped path itself: for an ask of size L the shipped path is solo (no planning call) plus QA, and that exact path has not been run end to end. What the headline does not include: the separate screening call the shipped path makes before the plan ($0.006 on Gemini, $0.014 on Haiku), and any end-to-end run of the solo path in real use (§11.2). On the small-ask suite S2 the picture is different: raw Haiku $0.0437, Composer alone $0.0488, Composer × Composer with QA $0.086, so on a small ask there is no saving against Haiku. That is why the S family is not the evidence for the headline.
bench/merges/run.sh L claude-deepseek false # one arm: pair · size · QA — clones ministore, drives /cockpit/goals/plan,
# audits with bench/ladder/audit-dir.sh, POSTs the audit to /settings/merges/record
bench/merges/matrix.sh # the built-in pairs × S · L, own serve on :41941, its own home
bench/merges/jobs.sh · jobs2.sh · jobs3.sh · jobs4.sh # QA-on runs · blueprint (tests) · a wave of 3 (report) · raw singles + six pairs
GET /api/v1/settings/merges # the table, weights included
POST /api/v1/settings/merges/record # an Outcome; audit_only joins an existing row
GET /api/v1/settings/audit # every call: model · ms · tokens · deterministic cost