The blog.
How model merging is measured, what the benchmarks actually found, and what is deterministic and what is not.
Compiled, not trained.
Measure every run, rank the combinations, bind the winner to a role, and let a router give each job to the model that earned it. No weights change — what a merge produces is a pipeline with its own score, price and speed. Includes what is deterministic in antiloki and what is a model's opinion.
Read → the first benchmark · 2026-09-22Every run, with the scripts.
One planner, one worker, with and without the QA gate, against raw Claude Opus on the same task — the wins, the losses, the arm where the benchmark says don't merge at all, and the two big runs that failed for reasons that were ours. Measured on Claude and DeepSeek; kept as history.
Read → the long version · fourth edition · 2026-10-01Heimdall, the merge and the adapter.
The whole argument at length: every rule, every formula and every number behind building a specialised coding model out of measured parts — now with the five subscription seats, the one tolerance, the QA judges' false-accept rates, and the cost against raw Claude: cell by cell, and on the eight-module task against Opus, the headline reference.
Read →