antiloki · the method

Compiled, not trained.

Measure every run. Rank the combinations. Bind the winner to a role. Let a router give each job to the model that earned it.

Updated 2026-10-01. The method on this page stands: compiled, not trained; deterministic where it can be, a model's opinion only where it must. Four details moved. The weights and tolerance are different now — one tolerance (a seat wins if it keeps at least 80 % of the best measured quality), then price, then time, rather than the 0.7 / 0.2 / 0.1 and three-point rule below. QA is obligatory and deterministic-first, not “switched on when you want it”. The board in section 04, and the Qwen and DeepSeek figures, are the first edition's withdrawn arms: the product is five subscription seats. And an ask of size L or smaller now runs on one seat alone with no planner. Current numbers: the technical report, fourth edition.

the claim, in one paragraph

The same model, checked, is a better model. One Qwen session passes 35 of 37 behavioural checks on an eight-module task and misses the same two every time. The identical model — planning for itself, verifying its own diff — passes all 37, three runs out of three. That is what a merge sells: not a better model, a checked one, at the cheap model's price.

01No weights change

A merge is not a fine-tune and not a distillation. Nothing is trained, and no weights exist afterwards that did not exist before. What a merge produces is a pipeline — a binding of one model to each job of a task, with its own measured score, price and speed.

A task has more jobs in it than "write the code". Something has to decide what to build, something has to write it, something has to check the diff against what was asked, something has to judge whether the result is acceptable, and something has to read the codebase well enough to answer the first question. Those are five different jobs with five different difficulty profiles, and the industry's habit is to give all five to the most expensive model in the house.

A merge gives each job to whichever model the benchmark measured best at that job — which, for the two jobs that dominate the token count, is usually not the expensive one.

02The adapter: what a fine-tune would have had to learn

Binding models is half of it. The other half is that a general model knows nothing about your codebase, and a fine-tune's whole value proposition is that it has been made to know something.

antiloki supplies that knowledge instead of training it in: the laws your repository lives by, the dependency graph, the per-file reviews and their axes, the blast radius of a change — all recomputed deterministically every commit — plus the skills, the specialist lenses and the tools the role works with. That package is what an agent starts from, and it is why a cheap model writing inside it produces work a cheap model does not otherwise produce.

Same models. Better orchestration. Measured results.

03The three steps

01 · measure

Every run is scored.

A behavioural audit of the result, a deterministic cost (model × tokens, cache reads at a tenth), and a real timer. One row per job × size × combination. Ours ships with the app; yours is added every time you work.

02 · rank

Quality first, then price.

0.7 quality + 0.2 cost + 0.1 time — and a candidate more than three points behind the best gets no credit for being cheap. A row earns its place by repeating; one lucky run ranks nothing.

03 · hire

You hire a role, not a stage.

Nobody hires a verifier. You hire a Backender, a Frontend designer, a Tester — and it arrives carrying the merge that won its kind of work, plus the skills, specialists and tools that role uses. Three of each: cheaper, faster, best.

04What the board looked like

the role you hireits merge, measured on its own workscoretimecost / run
BackenderCodex gpt-5.6-terra decides · OpenCode DeepSeek v4 Flash produces37/37305 s$0.012
FrontenderCodex gpt-5.6-terra decides · OpenCode Qwen 3.6 Plus produces100%620 s$0.043
ReviewerCodex gpt-5.6-sol checks every claim against the diff4/421 s$0.001
a one-file changeno merge — one OpenCode Qwen 3.6 Plus session, five for five5/534 s$0.011

Two of those rows are the method proving itself rather than the product selling itself. The cheap hands that win on backend came last on the screens — so the Frontender is bound to a different worker, because that is what the numbers said. And on a one-file change the benchmark says don't merge at all: the planner call is pure overhead, and the router obeys it.

A ranking that always recommended the product's most elaborate arrangement would not be a ranking.

05It keeps measuring after you install it

The table above is a starting point, not an answer. Every run you do is scored on your machine — which merge, which job, how long, what it cost, whether it passed first time — and filed into the same table, which grows with use. Your runs adjust the shipped scores, by at most a third, and a pairing that has only been watched working never displaces one that was benched: observation moves a score, it does not set one.

Nothing leaves the machine to make that happen.

06What is deterministic, and what is not

Worth being exact about, because "verified" is the most abused word in this industry.

Always on, and not a model's opinion: every turn is priced and timed; the fence around an agent is enforced where the work lands, for every seat, so a change outside it is shown, can be undone, and stops the merge or the delivery; on Claude panes a hook that sits outside the model also logs each tool call and refuses a write outside the fence before it happens, and the other seats have no such hook; and the codebase review's numbers are arithmetic over your ASTs and git history — size, complexity, coupling, density, churn, untested — so the same commit scores the same twice.

On every run: the project's own test, build and lint are available as gates the moment an agent claims it is finished, with a failing gate handing the agent the real log to try again; and QA checks each acceptance claim, deterministically where it can (a real page driven in Chrome, requests sent to a server) and with a judge where it cannot. The judge is a model, and it is labelled as one wherever its verdict appears.

← more from the antiloki blog  ·  every run behind the numbers →  ·  the long version →