← antiloki.com

antiloki · technical report · fourth-edition revision 2026-10-01 · third 2026-10-01 · first 2026-09-23, second 2026-09-29

Compiled, Not Trained

Heimdall, the merge and the X-ray adapter: how antiloki builds a specialised coding model out of measured parts.

Second edition, measured 28–29 September 2026 on the five subscription seats — Claude Code, Codex, Gemini, Grok and Cursor. The first edition (22–23 September) measured Chinese models through OpenCode Zen; those arms are withdrawn and kept on disk as history. §4.2, §4.3, §7.1, §8 and §9 are rewritten; the depth runs still to come are listed in §10.

Third edition, 1 October 2026 — adds §11: the one tolerance (a seat wins if it keeps 80 % of the best measured quality, then price, then time), the solo path for asks of size L and smaller, a screening call of its own, obligatory QA, the QA judges' false-accept rates on two defect sets, the cost against a raw Claude session cell by cell (the 3× the site carried at the time of the third edition), and three X-ray experiments. It supersedes the 0.10 absolute band of §7.1, the planner and screening rows of Table 1, the “on request” QA of §5, the judge paragraph of §8.1 and the open items of §10 that are now done. Measured 2026-09-28 to 10-01; no Claude arm was re-run in this edition, and the Claude references date from 2026-09-22 to 28.

Fourth-edition revision, 1 October 2026 — the site’s headline reference is now raw Claude Opus and not the cheapest Claude arm inside the tolerance, because Opus is the Claude model a subscriber on the top plan reaches for. The headline is 10× on the eight-module task: the best measured cell, 11.3× rounded down, with a range of 6×–11× over the days Opus was measured (§11.7). The seven-cell analysis of §11.4, which gave the 3×, stays as the Sonnet-and-Haiku-reference result and as context; it is no longer the headline. The weekly-limit figure is restated as a paired task result (§11.4, §11.7).

Abstract

A coding agent's behaviour is decided by three things: which model answers each kind of question, what that model is shown about the codebase, and how the whole is judged. antiloki fixes all three without training a weight. A merge assigns every job of a coding run — planning, editing, verifying, judging, reading, reporting — to one side of a brain × hands pair; a deterministic adapter conditions each call — X-ray's facts computed from the repository, plus the skills, specialists and tools the role carries — rather than on what the model remembers; and a benchmark re-weights every combination from measured outcomes, so the assignment learns. The result behaves like a new model, with its own score, price and speed. The central result of this edition is about where the money goes. At small size the writer decides quality: every pairing writing with Cursor Composer scored 1.00 on a calibrated sixteen-case task whatever planned it, and a frontier planner only multiplied the price, by 1.6–2.7× (raw Claude Haiku, writing alone, scored 0.87 on the same checks — the grader separates writers). At large size one cheap seat planning and writing its own work stayed within the quality band of the best pairing in all six domains, at $0.08–0.12 a run. The cheap seats are also the better screeners — Gemini and Haiku caught 87.5 % of deliberately under-specified prompts, Claude Sonnet 68.8 % — and the better verifier: Composer judged every must-pass and mutant case correctly. What a merge sells is a checked result at a cheap model's price; the expensive model is paid for only where a measurement says it earns it. We give every rule, formula, price and number the system uses, the results, and where the analogy to a trained merge breaks.

1The claim, stated carefully

"Model merging" in machine learning means arithmetic on the weights of checkpoints that share a base (task vectors, TIES, DARE; mergekit is the tool). antiloki does none of that: the models are frozen, hosted by their vendors, reached through the vendors' own CLIs on your own subscription. What antiloki merges is behaviour: the answer to "which model does which job, shown what, judged how" — and that answer is learned from outcomes.

The analogy we defend is narrower and, we think, exact in practice. A LoRA adapter specialises a frozen model for a task by adding a small learned component in front of it. X-ray is a small deterministic component in front of every call that supplies what a coding fine-tune would otherwise have to learn — the structure, laws, risks and history of this repository — and it is recomputed at each commit instead of going stale. The merge table plays the role of the training loop: an outcome-weighted selection over roles, updated by every benchmark row and every audited use. The difference from a trained merge is the ceiling: no role can exceed what its base model can do when shown the right evidence. For code, where the missing knowledge is almost always the codebase and not the craft, that ceiling has not been the binding constraint in any measurement below.

the merge

brain × hands. The brain plans, verifies, judges, reads and screens; the hands edit, write tests and report. Each side is one engine and, for OpenCode, one model.

the adapter

X-ray: deterministic code intelligence — reviews scored axis by axis, the laws, the graph, blast radius, coverage gaps — packed into every call's context.

the training loop

the benchmark: every run and every audited call becomes a row per job × size × planner × worker × QA, re-weighted by quality, cost and time.

2Related work

Query routing between a strong and a weak model is well studied. RouteLLM (LMSYS, 2024) trains routers on Chatbot Arena preferences and reports cost reductions of over 85% on MT Bench at 95% of GPT-4's quality [1]. Commercial routers — Martian [2], Not Diamond [3], and the OpenRouter aggregator now processing 25 trillion tokens a week [4] — sell the same lever per request. Heimdall differs in what it routes over: not a request, but a job inside a run, with the job's own acceptance and its own measured winner.

The brain × hands split is aider's architect/editor mode [5]: a reasoning model describes the change, a cheaper model writes the edit. aider reported a new state of the art on its polyglot benchmark with DeepSeek R1 as architect and Claude Sonnet as editor at 14× less cost than the previous o1 result [6]. Cline and Roo Code let users bind different models to Plan and Act modes by hand [7]. antiloki generalises the split from two roles to ten jobs and makes the binding measured rather than chosen.

That a composition of models can behave as a new, better model was shown by Mixture-of-Agents (Together, 2024): layered open models reached 65.1% on AlpacaEval 2.0 against GPT-4o's 57.5% [8]. The "compound AI systems" position from Berkeley [9] argues the unit of capability is now the system. We take both as the frame: the merged agent is the model; its card is its benchmark row.

3The jobs

Every call antiloki makes on a coding run belongs to one of ten jobs. A job has a side (brain or hands), a prompt template, and — the point of the system — its own winner and fallback per engine and globally.

Table 1 · the ten jobs and their sides
jobsidewhat the call doesfires when
plannerbrainscreens the ask and writes the plan: tasks that own disjoint files, each with acceptance linesa goal of two or more enumerated tasks or over 600 characters; an ask of size L or smaller skips it (§11.2)
screeningbrainthe blocking/assumable questions before a planbefore the planner, as a call of its own on the best screening seat (§11.2)
editor-shandsa change to one file, an ask under 400 charactersa task that owns one file
editor-lhandsa change across files or a long askevery other task
verifierbrainchecks a task's diff against its acceptance; may send it back onceafter every task
test-writerhandswrites the tests a contract's acceptance lines implyblueprint mode
qa-gatebrainchecks each acceptance claim against the final diff; the isolated judgeevery run — QA is obligatory (§11.2)
reporthandsconsolidates the tasks' reports into onea multi-task run (a single task's report stands)
lookupbraina factual question about the code: where, which, whoa chat turn the classifier reads as a lookup
synthesisbrainwhy, how, compare, in which order — an answer that reasonsa chat turn the classifier reads as synthesis

The classifier is deterministic (cli::classify_ask): a question mark or an interrogative first word makes a read; among reads, the words why, how, explain, describe, compare, difference, trade-off, should, in which order make it synthesis, else lookup. A write that names at most one file and is under 400 characters is editor-s, else editor-l. Sizes for the benchmark follow the plan, not the ask: two or more tasks is XXL; one task owning one file with a goal under 400 characters is S; else L.

4Heimdall, the router

4.1 Resolving a binding

For a job on a run, Heimdall produces a binding (engine, model) and a fallback, in this order:

  1. The merge. If the run carries a pair, the job's side decides the engine and, when the side names one, the model. job_binding_paired.
  2. The learned table. For a solo run, if the merge table holds a row for this job and size with at least three samples, its best row outranks the static table. job_binding_learned, merges::best_for(job, size, 3).
  3. The static table. The winner and fallback measured for the job (Table 2), taken globally if the winner's CLI is installed (which), else the run engine's own cell. job_binding_any.
  4. The effort dial. CHEAPER moves every role one rung down its engine's price ladder; SMARTER one rung up; BALANCED leaves the bound models. AUTO lets steps 1–3 stand; ON obeys the dial; OFF routes nothing.
  5. The price rule. A worker is never dearer than its planner: worker_under_planner walks the worker down its ladder until its list price is under the planner's, and logs why.
  6. Failure. A binding whose engine fails is retried once on the fallback, at the fill for workers and in run_job for every other job.
Table 2 · the global winner and fallback per job (the static table, FIRST EDITION 2026-09-23 — its winners are the withdrawn OpenCode Zen arms; the learned table of §7–§8 now decides, and this table is only the seed)
jobwinnerfallbackthe measurement behind it
lookupopencode · deepseek-v4-proclaude-code · haikureaders /50: Sonnet 42, Pro 41 at $0.13, Qwen 38, Haiku 36 (strong on lookup, fabricates on synthesis)
synthesisclaude-code · sonnetopencode · qwen3.6-plussynthesis /20: Sonnet 16, Qwen 14, Pro 13 — outside the band, Sonnet keeps it
screeningopencode · deepseek-v4-flashclaude-code · haikupart of the planner call; cheap by design
planneropencode · deepseek-v4-proclaude-code · sonnetthe planner's model did not move L quality (Opus 37, 36 · Sonnet 35, 36); Pro is within the band and 5× cheaper
editor-sopencode · qwen3.6-plusclaude-code · sonnetS ×5 raw: Qwen 5/5 · 34 s · $0.011; Sonnet 5/5 · $0.093; Pro 4/5; Flash 2/5
editor-lopencode · deepseek-v4-flashopencode · deepseek-v4-proL one-worker: Flash under a planner 36–37/37 at $0.02 of worker cost
verifierclaude-code · sonnetopencode · deepseek-v4-proboth L rejections it made were real (missing JSDoc; a wrong money format)
test-writerclaude-code · sonnetopencode · deepseek-v4-problueprint runs — being measured
qa-gateclaude-code · opusopencode · deepseek-v4-prothe judge is the one place the dearest model is bound by default
reportopencode · deepseek-v4-flashclaude-code · haikua consolidation; cheap by design

The three-point rule. Our rule for every table above: when two candidates are within three points of each other on the job's score, the cheaper one is the winner. Three points of fifty is 0.06 on the unit scale, and that number reappears as the band in §7.

4.2 Prices and the deterministic cost

Every cost antiloki reports is computed, never read from a vendor's estimate: input and output tokens times the model's list price per million, with cache reads at one tenth of input and cache writes at 1.25× (cli::price_tokens). The same table gives the ladder each engine's effort dial walks.

Table 3 · list prices per million tokens for the five seats (cli::prices_of) — an API-equivalent yardstick; the seats are flat subscriptions
modelinput $output $modelinput $output $
claude opus5.0025.00gemini pro1.2510.00
claude sonnet3.0015.00gemini flash0.140.28
claude haiku1.005.00grok 4.63.0015.00
codex / gpt-51.2510.00grok code fast0.201.50
cursor composer 2.50.501.50
cost = tin·p_in + tout·p_out + cache_read·0.1·p_in + cache_write·1.25·p_in (per 10⁶ tokens)

Each call is also timed on the serve's own clock and its tokens counted from the engine's stream; the triple (cost, ms, tokens) lands on the pane's timeline as an audit line and in the ai_calls table, and — through merges::record_use — as a benchmark row of size call.

What "saved" is measured against (since 2026-09-28): the same tokens priced at the frontier model's list rate — Claude Opus, or Codex's frontier model on a Codex seat (cli::frontier_of) — against what the work cost at the seats that actually ran it, each task priced at its own engine and model. Before that the yardstick was the seat's own top rung, which for four of the five seats is the same rate as its cheap rung and printed "saved $0.00" on work five times cheaper than Opus. A row nobody priced is priced from list rates (merges::estimated_usd); a replayed plan is never counted as a free planner.

4.3 Modes and the merge in Studio

A merged agent in Studio is a recipe with a pair: a name, a brain side, a hands side, a theme, a scope, its specialists and skills. Since 2026-09-27 every side is one of the five subscription seats — Claude Code, Codex, Gemini, Grok and Cursor; the OpenCode/Zen pairs of the first edition (Claude × DeepSeek, Claude × Kimi, Codex × DeepSeek …) are withdrawn and kept on disk only as history. A seat may sit on both sides (Composer plans and writes), which is measured as a pairing like any other. Heimdall's chip on a pane shows the mode and the money the router saved on that pane, computed as API-equivalent dollars against one high-tier session doing the same tokens.

5The pipeline a goal runs through

A goal run is the unit the benchmark measures. Its stages, and the model that answers each, are what the merge assigns.

ask → planner (screens: blocking / assumable questions; one answer round) → plan (tasks · owns · needs · acceptance) → ratify → projection (one worker, or a wave?) → fills, in waves → verify each task (once back at most) → join (a report call on multi-task runs) → QA gate (claims vs the final diff, with the deterministic expect, coverage and API checks; every run) → baseline · savings · the merge row

The projection. Before a plan fans out, antiloki estimates whether a wave beats one worker. A single session doing k parts is projected at SERIAL_FACTOR = 1.5 times the parallel per-part time; the wave's concurrency is bounded by the depth of the plan's dependency chain (effective_concurrency = ceil(fills / chain), clamped to the configured cap); and a plan whose single-session projection is under the parallel floor — 900 s by default, ANTILOKI_PARALLEL_FLOOR_S, 0 = off — collapses to one worker. With no history the medians assumed are 120 s per fill and 150 s of orchestration. This is why every S and L row below is a one-worker run: the wave only pays above fifteen minutes of projected serial work.

The verifier's one send-back. A task the verifier rejects goes back to its worker once with the rejection as the brief; a second rejection fails the task. Both rejections observed on L were real defects a compile gate would not catch.

The QA gate. An isolated judge checks each acceptance claim against the final diff and answers met / unmet / unproven; on a one-worker run the always-branch kept a placeholder claim, so the gate reported unproven. That gap is closed in the third edition: QA is obligatory, runs the deterministic checks (§11.2), sends unmet claims back to their workers a bounded number of times and, when those are spent, turns the run into a question for the person.

6The adapter

A LoRA specialises a frozen model by adding a small learned component in front of it. antiloki's adapter is the whole conditioning layer around frozen models, and it is larger than the router: which model answers each job, the brain × hands split, the skills materialised into the agent's box, the specialists it may consult, the tools it may drive, and the deterministic intel it boots with. A Designer is not a different model from a Backender — it is a different adapter.

the models

a merge binds every job to a side: the brain plans, verifies, judges, reads; the hands edit, test, report. Per job, per domain, measured.

what it knows

skills materialised into its box, specialists it may consult, and the X-ray packs it reads before it explores.

what it drives

the panes it may open — Frames and the design system for a Designer, X-ray and the constitution for a Reviewer, the runner for a Tester.

6.1 X-ray — the deterministic half

X-ray is antiloki's analysis of the repository. Nothing in it is generated by a model; all of it is computed from the files, the ASTs, the tests and the git history, and it is what every call is shown before it explores.

reviews
every file scored, and why not 100 axis by axis — coupling, structure, cohesion, defects, tests, change — with the comment at the line that carries the move.
laws
the rules the codebase lives by, checked at the line (for TypeScript: no any, no var, no loose equality, no debugger, no @ts-ignore, no swallowed errors), with what compliance is worth.
graph
every file a node, imports the edges: hubs, orphans, cycles, the most central and the most risky, per folder.
blast radius
what a change to a file would touch, before it lands.
coverage gaps
what the tests never reach, by file; the score's history over snapshots.
context pack
the bundle a pane boots with: laws · folder · file · graph · tests · blast · doc items, derived from the pane's scope by default, under a character budget (200 000).
the retriever
for a chat turn, chat_auto::compose_for picks from the workspace's intel and prior runs — the chat's own log excluded — under an 8 000-character cap.
MIN mode
exploration off: the agent navigates by the pack instead of reading the tree, and touches the disk only to edit inside its scope.

This is why the analogy in §1 holds for code specifically. What a repository-specific fine-tune would teach — where things are, what the conventions are, which files are dangerous — is the content of the pack, and the pack is true at the commit the agent is working on.

6.2 The role — the adapter as a unit a person hires

A job is a stage inside one run; nobody hires a verifier. A role is what a person hires — a domain of the codebase — and inside it every one of the ten jobs still happens. So a role is not an eleventh job; it is a profile over all ten, scoped to a domain, carrying its own adapter.

Table 8 · the roles, and what each one's adapter carries
roledomainconsultsdrives
Backender — Database, API under itbackendcorrectness · errors, types · architecture · dependencies · security · injectioncode graph, code, diff
Frontender — Designer, Architect under itfrontendframework · idioms · style · suppressionscomponents, code, diff
Frontend designerdesigncategory · style · axis · cohesionFrames, design system, components, canvas
Testerteststesting · isolation · axis · testsrun, X-ray, diff
Reviewermixedaxis · defects, compliance, change · book · wardenX-ray, constitution, diff — no editor
Product ownerdocsbook · definer, questioneer, synthesizerintent, todo, kanban, context
Infrainfrasecurity · secrets · book · wardenterminal, run, git graph

The domain is derived from the paths a task owns, deterministically (roles::domain_of): a test of the API is test work; the styles folder beats the extension; a change touching six backend files and one stylesheet is still backend, and an even split is honestly mixed. So every run tags itself, and runs that happened before the dimension existed can be attributed without re-running them.

A role asks the table for the merge that won its kind of work. When its domain has no samples it reaches in order — its nearest neighbour (design looks at frontend; infra and docs at backend), then the shared pool, then everything pooled — so a role always has an answer and the sample count says how far it had to reach.

6.3 Three of each

The same rows, three questions. Cheaper counts the bill above the score (0.35 / 0.55 / 0.10); Best is the table's own weight (0.70 / 0.20 / 0.10); Smarter counts only quality and allows no gap at all below the best. Each returns a different merge, and the three-point band still holds for the first two: a row ten points behind cannot be bought back at any price. They are the same three stops as the effort dial on the pane.

What "training" a merge means here. No weights change, so a new pair — Codex × Gemini, say — cannot be fine-tuned. Three things do adapt: it inherits its sides' measured priors rather than starting blank; it learns from its own use, since every audited call and run files a row and the learned table overrides the static one at three samples in its domain; and the handoff between the sides — the brief the brain writes, how much diff the verifier is shown, the acceptance lines it is judged on — is a set of deterministic parameters that can be measured and tuned per pair. That last one is prompt-level adaptation with frozen weights: the only honest sense of training available, and the one this system has not yet done.

7The benchmark, and how it learns

The merge table (~/.antiloki/merges.sqlite, merge_scores) has one row per job × size × planner × worker × qa. Sizes are s · l · xxl for runs and call for audited single calls. A row accumulates:

n · pass_first · retries · failed · gate_met · gate_unmet · audit_sum / audit_n · ms_total · usd_total · source

where audit is the external behavioural score a bench posts after the run (L: 37 checks of the eight modules; S: a five-case check of the function the ask names) and source says who measured it: bench (antiloki's own runs, shipped as the seed), run (the user's goal runs), use (audited calls). A user's table is therefore the shipped base plus their own rows, weighted together; nothing is collected back.

7.1 The weight

Every record re-weights all rows of that job and size against each other (merges::reweigh, merges::rank):

quality = audit (the deterministic grader's score) when one exists, else pass_first / tasks, × (1 − failed / tasks) band = quality ≥ 0.80 · best_quality (ONE tolerance, 2026-09-30 — third edition; the second edition: best_quality − 0.10) cost = band ? best_usd / usd : 0 time = band ? best_s / s : 0 raw = 0.15·quality + 0.75·cost + 0.10·time (Objective::Best, since 2026-09-28) weight = 0.5 + (raw − 0.5) · n / (n + 1) (n = 1 → halfway to 0.5)

Price first, speed last, quality as a gate. Third edition: the gate is now relative — a seat or pair wins if it keeps at least 80 % of the best measured quality (Objective::band() = 0.20), and measured quality outranks claimed; there is no second tolerance anywhere in the product (§11.1). The band disqualifies a materially worse answer outright; inside it, the bill decides. The first edition weighted 0.7 / 0.2 / 0.1 with a 0.06 band, which was narrower than the instrument — a ten-instance bench resolves quality in steps of 0.10 — and would pay 6.7× for one instance in ten. Speed was reduced from 0.30 to 0.10 on 2026-09-28, after the fastest frontier planner won a simple frontend task that a model a tenth of its price had measured the same on. A pairing measured on several suites is pooled by samples, never overwritten; a row nobody priced is priced from list rates; a planner row is never cheaper than its own list price.

7.2 Winners

A job's winner is its best sized row; a job the benches never sized falls back to its best audited-call row with at least two samples; else it has no winner yet, and the Studio says so. Each winner is named by composition and role — the brain's brand and the hands' suffix as one word, then the job's title: CodexComposer Builder, Composer Tinker, ClaudeComposer Architect — and can be made an agent in one click, carrying its pair. For solo runs the learned table outranks Table 2 once a row has three samples, which is how the static table is meant to retire.

8Results

Measured 28–29 September 2026 on the five subscription seats, one machine, a pinned fixture (ministore at 3de696c), every arm in a fresh clone with its own plan (plan replay cleared per arm). Quality is the deterministic grader's score for that task — calibrated before any arm ran, never a model's opinion; dollars are the API-equivalent of every call the serve made for the run (planner, writer, verifier, QA), priced from the tokens each seat reported (§4.2). The tables are generated from the record by bench/campaign/make-results.py; rows under three samples are marked and never ranked.

L · backend — eight feature modules, 37 behavioural checks
planner × writernquality$/runs/runpriced
Codex medium × Composer 2.5 *21.000.1299256before QA · 09-26/28
Gemini × Composer 2.5 *21.000.1384132before QA · 09-26/28
Claude Sonnet × Composer 2.5 *21.000.1500118before QA · 09-26/28
Composer 2.5 × Composer 2.530.970.101795with QA · 09-29

* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.

L · frontend — two components + their pure helpers
planner × writernquality$/runs/runpriced
Codex terra × Gemini Flash *21.000.0363271before QA · 09-26/28
— none (raw) × Grok 4.6 medium *21.000.0837260before QA · 09-26/28
Composer 2.5 × Composer 2.531.000.1133157with QA · 09-29
Claude Sonnet × Claude Haiku *20.98—144before QA · 09-26/28
Grok 4.6 high × Composer 2.5 *20.980.0377156before QA · 09-26/28
Codex terra × Composer 2.5 *20.980.0378168before QA · 09-26/28
Claude Opus × Grok 4.6 medium *20.980.2766147before QA · 09-26/28
— none (raw) × Composer 2.5 *20.930.0319107before QA · 09-26/28
Claude Sonnet × Composer 2.5 *20.800.0318137before QA · 09-26/28

* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.

L · design — harvest the design system, 25 checks
planner × writernquality$/runs/runpriced
Codex terra × Gemini Flash *21.000.0143103before QA · 09-26/28
Codex terra × Composer 2.5 *21.000.015746before QA · 09-26/28
Grok 4.6 high × Composer 2.5 *21.000.015887before QA · 09-26/28
— none (raw) × Composer 2.5 *21.000.016737before QA · 09-26/28
Claude Sonnet × Composer 2.5 *21.000.021174before QA · 09-26/28
— none (raw) × Grok 4.6 medium *21.000.026478before QA · 09-26/28
Claude Sonnet × Claude Haiku *21.000.074095before QA · 09-26/28
Composer 2.5 × Composer 2.531.000.117393with QA · 09-29
Claude Opus × Grok 4.6 medium *21.000.124857before QA · 09-26/28

* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.

L · docs — documentation from the code
planner × writernquality$/runs/runpriced
Codex terra × Gemini Flash *21.000.020957before QA · 09-26/28
Codex terra × Composer 2.5 *20.980.013142before QA · 09-26/28
Claude Opus × Grok 4.6 medium *20.980.102246before QA · 09-26/28
Grok 4.6 high × Composer 2.5 *20.960.012454before QA · 09-26/28
Claude Sonnet × Composer 2.5 *20.960.013555before QA · 09-26/28
— none (raw) × Grok 4.6 medium *20.960.030565before QA · 09-26/28
Claude Sonnet × Claude Haiku *20.960.0570105before QA · 09-26/28
Composer 2.5 × Composer 2.530.960.081083with QA · 09-29
— none (raw) × Composer 2.5 *20.940.011227before QA · 09-26/28

* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.

L · infra — Dockerfile, compose and CI, 31 checks
planner × writernquality$/runs/runpriced
— none (raw) × Composer 2.5 *21.000.007047before QA · 09-26/28
Claude Sonnet × Composer 2.5 *21.000.007644before QA · 09-26/28
Codex terra × Composer 2.5 *21.000.007636before QA · 09-26/28
Grok 4.6 high × Composer 2.5 *21.000.008151before QA · 09-26/28
— none (raw) × Grok 4.6 medium *21.000.012932before QA · 09-26/28
Codex terra × Gemini Flash *21.000.020650before QA · 09-26/28
Claude Sonnet × Claude Haiku *21.000.050079before QA · 09-26/28
Claude Opus × Grok 4.6 medium *21.000.060643before QA · 09-26/28
Composer 2.5 × Composer 2.531.000.080068with QA · 09-29

* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.

L · mixed — review cart.ts: 4 planted defects, 5 clean controls
planner × writernquality$/runs/runpriced
Claude Sonnet × Claude Haiku *20.940.058075before QA · 09-26/28
— none (raw) × Composer 2.5 *20.880.012237before QA · 09-26/28
Codex terra × Composer 2.5 *20.880.013649before QA · 09-26/28
— none (raw) × Grok 4.6 medium *20.880.018050before QA · 09-26/28
Grok 4.6 high × Composer 2.5 *20.880.0189100before QA · 09-26/28
Codex terra × Gemini Flash *20.880.024885before QA · 09-26/28
Composer 2.5 × Composer 2.530.880.0830105with QA · 09-29
Claude Sonnet × Composer 2.5 *20.880.103063before QA · 09-26/28
Claude Opus × Grok 4.6 medium *20.880.266871before QA · 09-26/28

* fewer than three samples — shown, never ranked. Prices from different columns are not comparable: the later ones include the QA check every run now pays.

S2 — cartTotal in whole cents, two discount codes: 16 executed cases + the build
planner × writernquality$/runs/run
— none (raw) × Gemini Flash41.000.0445230
— none (raw) × Composer 2.541.000.048825
Composer 2.5 × Composer 2.581.000.086066
Codex terra × Composer 2.581.000.126054
Claude Sonnet × Composer 2.5161.000.140048
Grok 4.6 high × Composer 2.581.000.2320104
— none (raw) × Grok 4.6 medium41.000.248930
— none (raw) × Claude Haiku40.870.043725

Calibrated before any arm ran: a correct reference 16/16, rounding the discount 14/16, case-sensitive codes 14/16, naive floats 7/16, no change 0/16.

Screening — HumanEvalComm, one call per screen, clear vs ambiguous+incomplete prompts
seatscreensbalanced accuracy$/screens/screen
Claude Haiku360.940.01429
Gemini360.940.006125
Codex low360.910.01884
Grok 4.6 medium350.870.059021
Composer 2.5340.860.014515
Claude Sonnet360.840.04125

Balanced accuracy = (recall on damaged prompts + precision on clear ones) / 2. The per-seat recall and precision are in bench/screen-solo/out.

The judges, per domain — verdicts scored against the deterministic grader (must-pass artefacts and re-graded mutants)
job · domain · seatcasesaccuracy$/call
qa-gate · backend · Composer 2.560.830.0457
qa-gate · backend · Claude Haiku50.800.0461
qa-gate · backend · Codex low60.500.0899
qa-gate · docs · Codex low60.670.0823
qa-gate · docs · Composer 2.560.670.1483
qa-gate · infra · Claude Haiku41.000.0296
qa-gate · infra · Composer 2.561.000.0474
qa-gate · infra · Codex low60.500.0765
verifier · backend · Composer 2.561.000.0140
verifier · backend · Codex low60.670.0340
verifier · backend · Claude Haiku60.500.0455
verifier · docs · Codex low61.000.0266
verifier · docs · Composer 2.561.000.0110
verifier · docs · Claude Haiku60.670.0389
verifier · infra · Claude Haiku61.000.0283
verifier · infra · Codex low61.000.0238
verifier · infra · Composer 2.561.000.0093

Cases: 3 must-pass artefacts and 3 mutants per domain (files the ask names cut in half, re-graded below 0.90). Frontend, design and mixed have no honest must-fail case yet — see §10.

Test writer — SWT-bench Verified, fail→pass reproduction (their harness, their images)
planner × writernresolved$/runs/run
— none (raw) × Claude Sonnet100.900.3737111
Claude Sonnet × Composer 2.5100.800.1187161
— none (raw) × Composer 2.5140.790.201885

8.1 What the numbers say

9Limitations and threats to validity

10What follows

11Third edition: what changed between 29 September and 1 October

Everything here was measured or decided after the second edition's rows were written. The record is the same merge table (§7); a row is only counted once its quality has been audited by an executed check.

11.1 One tolerance

The rule, 2026-09-30: a seat or a pair wins if it keeps at least 80 % of the best measured quality in its job and size; among those, the cheapest wins; among equally cheap ones, the fastest. Measured quality outranks claimed quality. The earlier 0.10 absolute band (§7.1) is retired, and with it the per-job and per-role pickers in the app: the choice of AI for a job is automatic, and Settings keeps only the providers, the main provider and a switch for each seat. A seat that is switched off, signed out, or at 95 % or more of its quota window is skipped before the job starts and the work moves to the next seat in the measured ladder.

11.2 The solo path, a screening call, and QA on every run

11.3 The QA judges, scored on false accepts

A false accept lets a broken change through, so the QA gate is chosen on it, not on price. Two defect sets, each a broken run or a good run, graded by an executed check. The first (10 broken and 10 good per class, five classes): Opus 0 of 50 false accepts, Grok-4.6-high 0, Composer 2.5 0, Gemini 0, Codex 0 (but it rejects 39–41 of 50 good runs), Sonnet 8 of 50, Haiku 6 of 50. The first set is easy — five seats scored zero — so it does not separate them, and Gemini’s zero there does not carry over: on the second set it is the worst seat measured (below). The second (15 broken and 15 good per class, five classes) does:

QA gate · second defect set · false accepts of broken runs, and false rejects of good runs
seatfalse accepts, all classeswithout the frontend classfalse rejects
Claude Opus14 / 75 (19 %)0 / 6026 / 75 (35 %)
Grok-4.6-high18 / 75 (24 %)4 / 608 / 71 (11 %)
Cursor Composer 2.520 / 75 (27 %)6 / 6014 / 73 (19 %)
Claude Sonnet26 / 74 (35 %)12 / 5918 / 75 (24 %)
Claude Haiku27 / 75 (36 %)19 / 6021 / 74 (28 %)
Gemini flash (medium) · incomplete27 / 45 readable (60 %): backend 14 / 15, frontend 13 / 15not measured2 / 50 · 55 unread (docs 25, infra 30)
Gemini default · incomplete15 / 15 on backend (100 %); 113 unreadnot measured0 / 22

Paired sign tests on the cases where two seats disagreed: Opus against Grok-high 0 : 4 (p = 0.125), Opus against Composer 0 : 6 (p = 0.031), Grok-high against Composer 4 : 6 (p = 0.75), Grok-high against Sonnet 2 : 10 (p = 0.039), Opus against Sonnet 0 : 12 (p = 0.0005); no correction for multiple tests. An earlier ranking in tiers and an unpaired test were withdrawn after this paired recount.

11.4 The cost against a raw Claude session, cell by cell

Revised the same day, after a fact-check. The first version of this section printed 10.7× and the site said “10×”. It compared each winner with the Claude arm that made the ratio largest, it counted a 95 % interval as if the four S cells were independent, and it set raw single-seat sessions, which carry no QA, against raw Claude while the product ends every run in QA. This is the corrected section; the site said 3× on this basis until the fourth-edition revision (§11.7).

A cell is one job × size × domain × suite that holds a raw Claude reference with audited quality and at least one non-Claude arm with at least two audited samples. The winner of each cell is picked by the router’s own rule (§11.1), five subscription seats only. The reference is the cheapest raw Claude arm (Haiku, Sonnet or Opus) that keeps at least 80 % of the cell’s best measured quality, so that the ratio is not made larger by choosing a dearer Claude; the ratio is the reference’s dollars per run divided by the winner’s, API-equivalent list prices.

Seven cells · reference → winner · quality is winner against reference
cellreferencewinnerquality× cheaper
S backend · plainHaiku $0.0364 (n=3)Composer alone $0.0048 (n=3)1.00 vs 1.007.6
S backend · edgesHaiku $0.0509 (n=3)Composer alone $0.0057 (n=3)1.00 vs 1.008.9
S backend · branchesSonnet $0.0941 (n=3; Haiku scores 0.73, below 0.8 of 1.00)Composer alone $0.0062 (n=3)0.87 vs 0.8715.3
S backend · contractHaiku $0.0464 (n=3)Composer alone $0.0090 (n=3)1.00 vs 1.005.2
L backend · eight modulesSonnet $0.330 (n=1)Codex low alone $0.1215 (n=1, 2 audits)0.97 vs 0.972.7
L backend · fixHaiku $0.0577 (n=3)Composer alone $0.0150 (n=3)0.93 vs 0.903.8
L test-writer · SWT-benchSonnet $0.3737 (n=10)Gemini × Composer $0.1041 (n=10)0.80 vs 0.903.6
geometric mean, before the QA check5.7
the same with the QA step added (estimate)3.2

What the number is. The median is 5.2× and the cells run from 2.7× to 15.3×; dropping any one cell moves the geometric mean between 4.9× and 6.5×. No interval is printed: the four S cells are one synthetic small-ask family (twelve instances) where every arm scores about 1.0, they are not independent, and a bootstrap that treats them as seven cells is narrower than the evidence (the first version printed 7.0–15.2×). Without the family there are three cells, 2.7×, 3.8× and 3.6×, geometric mean 3.3×. Weighted by what a bill would be (dollars summed over the runs of each cell) the saving is 3.9× over the six cells with n of at least 3.

What the number leaves out: QA. Six of the seven winners are one seat running alone, filed as raw with no QA check, and the Claude arm they are set against has none either; only SWT-bench uses an antiloki pair (Gemini × Composer, 3.6×). The product ends every run in QA, and QA costs money: on the S2 suite raw Composer is $0.0488 a run and Composer × Composer with its QA check $0.086, 1.76× more; on the eight-module task Composer alone is $0.0671 and Composer × Composer × QA $0.1017, 1.5×. Dividing the 5.7× by 1.76 gives 3.2×. That division is an estimate, not a measurement — no end-to-end run was watched binding a solo seat (§11.2, §11.6), and the QA overhead of the other cells was not measured. The one cell where the whole pipeline was priced is the eight-module task: $0.1017 against raw Sonnet at $0.330, 3.2×. The 3× the site carried on this basis was that estimate rounded down; it sat under the 5.7× of the QA-free comparison, and under the 3.3× of the three cells outside the synthetic family.

Limits. n is 3 per arm in five cells and 10 in SWT-bench; the eight-module cell is one Codex run (two audits) against one Sonnet run (the first version used Opus, mean of 2, $1.151, which is inside the tolerance but not the cheapest Claude arm that is). The same-quality swap on SWT-bench (Codex × Composer, 9 of 10 like raw Sonnet, $0.157) is 2.4×; the 3.6× above is Gemini × Composer, which scored 8 of 10 (merges::separated is false at n=10, so “same quality” means inside the tolerance and not separated at these samples, not identical). The 6.7× of the first headline was raw Qwen3.6-plus ($0.0558) against raw Sonnet ($0.3737) on that same SWT cell — not an antiloki arm, a metered model that is off by default — and is withdrawn; the five-seat equivalent is the 3.6× above. Composer, Gemini and Codex dollars are estimated from reported tokens at list rate (Cursor and Antigravity report tokens, not dollars); the Claude references ran 2026-09-22 to 28. The eight-module row prices Codex low alone ($0.1215, 2.7×), the winner when the table was derived; the five-seat re-run of the same task later measured Composer alone at $0.0671 (4.9× under raw Sonnet) and the QA pipeline at $0.1017, and with Composer alone in that row the geometric mean before QA would be 6.2×. S2 and the S cells also disagree: raw Haiku $0.0437, Composer alone $0.0488 and Composer × Composer with QA $0.086 on S2, while Composer alone is priced about ten times lower in the S cells for asks of similar size (cause not found; Composer dollars are token estimates at list price). Time is not a headline axis: matched cells are about equal, except the eight-module cell, where Composer alone took 61 s against Opus’s 278 s (the Opus time moves between 112 s and 303 s from day to day).

The weekly limit. The site states +112 %, and restates what it is (fourth-edition revision). It is the paired task result of §11.7 — on the eight-module task, $0.1554 for the arm where Claude plans and judges against $0.3297 for raw Sonnet, 2.12× — and not a measured percentage of anyone’s weekly limit. The earlier derivation, 1 / (1 − 0.528) for the share of one machine’s tokens that ran off the frontier seat, is withdrawn: an independent audit could not reproduce the share, and the cuts of the record differed (+471 % over the whole mix, +106 %, +118 % and +122 % over three others). By dollars the gain over the recorded mix is +33 % to +45 %, which is the range to quote for a mix of work and not for a task.

The eight-module task, re-run on the five seats (37 checks, 2026-09-28/29; the page’s table): Composer alone $0.067, 61 s, 36.3 of 37 (n=3); Composer × Composer with a Composer QA check $0.102, 95 s, 36.0 (n=3); Sonnet × Composer $0.155, 106 s, 36.8 (n=5); Codex × Composer $0.165, 137 s, 36.7 (n=3); Gemini × Composer $0.184, 134 s, 36.3 (n=5); Grok-high × Composer $0.292, 199 s, 36.3 (n=3); raw Opus $1.151, 278 s, 37 (n=2, the 2026-09-22/23 reference); raw Sonnet $0.330 (n=1, 0.973).

11.5 Does the X-ray map help an agent? Three experiments

Each experiment crossed three agents (Cursor Composer 2.5, Codex, Gemini via Antigravity) with a small and a wide change and three scenarios — the repository only, the X-ray intel only with the agent told not to explore, or both. One run per cell; graded statically, nothing compiled; Gemini reports no tool-call telemetry.

11.6 What the third edition did not run

11.7 Fourth-edition revision: the headline reference is Opus

How the headline is defined: the best measured antiloki result against raw Claude Opus on the same task and the same checks, not against the cheapest Claude model. The reasoning is that Opus is the Claude model a subscriber on the top plan reaches for. All figures below come from the merge record (merge_scores) and are the eight-module task of §11.4 (37 checks, ministore); the row ids are in the merge record.

Eight modules · cost per run against raw Opus and raw Sonnet · API-equivalent list prices
armnchecks$/runs/run× under Opus× under Sonnet
raw Claude Opus · 2026-09-22/23237.0 / 371.151278——
raw Claude Sonnet · 2026-09-23136 / 370.32971213.5—
Composer alone, no QA step336.3 / 370.06716117.24.9
Composer × Composer, with QA — the headline arm336.0 / 370.10179511.33.2
Sonnet plans, Composer writes, Sonnet judges536.8 / 370.15541067.42.1

The headline is the pipeline with its QA step: 11.3×, rounded down to 10×. Its quality is 0.973 of 1.0, inside the one tolerance of §11.1 and not equal to raw Opus; it passes 36.0 of the same 37 checks on average. Opus is not stable from day to day: the ladder run of 2026-09-22 recorded $0.61 for the same task, and against that run the pipeline is about 6.0×, so the honest range is 6×–11× and the headline is the best measured cell. The weekly-limit figure, +112 %, is the Sonnet-planned arm against raw Sonnet, $0.1554 against $0.3297 = 2.12×; with n=5 against n=1 and list-price dollars standing in for an allowance the vendors do not publish, it is a paired task result and not a share of a quota window.

The headline arm is the nearest measured arm to the shipped path, not the shipped path itself: for an ask of size L the shipped path is solo (no planning call) plus QA, and that exact path has not been run end to end. What the headline does not include: the separate screening call the shipped path makes before the plan ($0.006 on Gemini, $0.014 on Haiku), and any end-to-end run of the solo path in real use (§11.2). On the small-ask suite S2 the picture is different: raw Haiku $0.0437, Composer alone $0.0488, Composer × Composer with QA $0.086, so on a small ask there is no saving against Haiku. That is why the S family is not the evidence for the headline.

AReproducing a row

bench/merges/run.sh L claude-deepseek false     # one arm: pair · size · QA — clones ministore, drives /cockpit/goals/plan,
                                                  # audits with bench/ladder/audit-dir.sh, POSTs the audit to /settings/merges/record
bench/merges/matrix.sh                            # the built-in pairs × S · L, own serve on :41941, its own home
bench/merges/jobs.sh · jobs2.sh · jobs3.sh · jobs4.sh   # QA-on runs · blueprint (tests) · a wave of 3 (report) · raw singles + six pairs
GET  /api/v1/settings/merges                      # the table, weights included
POST /api/v1/settings/merges/record               # an Outcome; audit_only joins an existing row
GET  /api/v1/settings/audit                       # every call: model · ms · tokens · deterministic cost

RReferences

  1. [1]LMSYS, RouteLLM: an open-source framework for cost-effective LLM routing, July 2024.
  2. [2]Martian's tool automatically switches between LLMs to reduce costs, TechCrunch, Nov 2023.
  3. [3]Not Diamond — model routing for coding agents.
  4. [4]OpenRouter more than doubles valuation to $1.3B in a year, TechCrunch, May 2026; Sacra for volume.
  5. [5]aider, Separating code reasoning and editing, Sept 2024.
  6. [6]aider, R1 + Sonnet set SOTA on aider's polyglot benchmark, Jan 2025; R1 + V3 benchmark.
  7. [7]Cline: Plan & Act; Roo Code: using modes.
  8. [8]Wang et al., Mixture-of-Agents enhances large language model capabilities, 2024.
  9. [9]Zaharia et al., The shift from models to compound AI systems, BAIR, Feb 2024.
  10. [10]mergekit — what "model merging" means in ML, and what this paper does not do.