A desktop app for merging the subscriptions you have into one coding agent.
Windows today · macOS and Linux builds are older, see Download

Build +112% more on your weekly limit, 10× cheaper, with the same code quality.

Claude Grok, Gemini Codex — antiloki merges your existing model subscriptions into specialized coding agents. No new account and no extra subscription: antiloki drives the CLIs you are already signed in to, your code goes only to those vendors, and nothing is kept on a server of ours.

eight TypeScript modules in a real React store · tsc -b && vite build must pass · 37 behavioural checks · cost per run, same task, same checks · pipeline arms measured 2026-09-28/29, Claude references 2026-09-22/23 · Composer = Cursor Composer 2.5 (needs a Cursor seat)
antiloki · Composer alone, no planner, no QA step (an ask of size L) $0.067 36.3/37 · 61 s · n=3
antiloki · Composer plans, Composer writes, Composer judges (QA included) — the headline arm $0.102 36.0/37 · 95 s · n=3
antiloki · Sonnet plans, Composer writes, Sonnet judges $0.155 36.8/37 · 106 s · n=5
antiloki · Codex plans, Composer writes, Codex judges $0.165 36.7/37 · 137 s · n=3
antiloki · Gemini plans, Composer writes, Gemini judges $0.184 36.3/37 · 134 s · n=5
antiloki · Grok-4.6-high plans, Composer writes, Grok judges $0.292 36.3/37 · 199 s · n=3
raw Claude Sonnet, one session (the secondary reference, 2026-09-23) $0.330 36/37 · 121 s · n=1
raw Claude Opus, one session (the headline reference, 2026-09-22/23) $1.151 37/37 · 278 s · n=2
armcost per runagainst raw Opusagainst raw Sonnet
Composer alone, no QA step$0.067117.2× vs Opus4.9× vs Sonnet
Composer × Composer, with QA$0.101711.3× vs Opus3.2× vs Sonnet
Sonnet plans + judges, Composer writes$0.15547.4× vs Opus2.1× vs Sonnet
10× is our best measured result against raw Claude Opus on the eight-module task. Raw Opus, one session: $1.151 a run (n=2, 37 of 37 checks, 278 s). The antiloki pipeline, Composer planning, writing and judging with its QA step: $0.1017 a run (n=3, 36.0 of 37, 95 s). That is 11.3×, and the headline rounds it down to 10×. This headline arm is the nearest measured arm to the shipped path, not the shipped path itself: for an ask of size L the shipped path is solo (no planning call) plus QA, and that exact path has not been run end to end. The reference is Opus because it is the Claude model a Max subscriber reaches for; against raw Sonnet ($0.330, n=1) the same pipeline is 3.2× and the Sonnet-planned run 2.1×, both in the table above. The honest range is 6×–11×. Raw Opus does not cost the same every day: the ladder run of 2026-09-22 recorded $0.61 for this same task, and against that run the pipeline is about 6.0×. 11.3× is the best measured cell, not the typical one. The 17.2× row has no QA step; the shipped path is solo plus QA, and it was not measured end to end. The same code quality here means code that passes the same 37 checks: the pipeline averaged 36.0 of 37 against 37 for raw Opus, which is 0.973 of 1.0 — inside the product's 80 % tolerance, not identical. The checks are executed code written before any arm ran, not a model's opinion; quality is the mean over the audited runs (4 of 5 in the Sonnet and Gemini arms, so their means are 36.75 and 36.25, printed 36.8 and 36.3). +112 % on your weekly limit is a paired task result and not a measured percentage of anyone's weekly limit. On this task the arm where Claude still plans and judges while Composer writes cost $0.1554 against $0.3297 for raw Sonnet: 2.12×, so about 2.1× further on that task (n=5 against n=1, list-price dollars standing in for an allowance the vendors do not publish). What it does for your limit depends on your mix of work. Limits. One task in one repository, n of 1 to 5 per arm. Dollars are API list prices; Cursor, Codex and Gemini report tokens, not dollars, so theirs are estimated from the tokens. The shipped path adds a separate screening call before the plan ($0.006 on Gemini, $0.014 on Haiku) that is not in these figures, and the solo path (no planner, for an ask of size L or smaller) has not been watched end to end in a real run. The small-ask suite cuts the other way: on S2, raw Haiku cost $0.0437, Composer alone $0.0488 and the Composer pipeline with QA $0.086, so on a small ask there is no saving against Haiku — and in the report's small-ask cells Composer is priced about ten times lower than on S2 for asks of similar size, cause not found. That family is not the evidence for this headline. Across seven cells against the cheapest Claude arm inside the tolerance the geometric mean was 5.7× before QA and about 3.2× with QA added by estimate; the report keeps them as context. The eight: fuzzy product search, a stable sort, two localStorage-backed React hooks, money formatting with discount codes, a quantity-stepper component, and checkout validation with a Luhn card check — strict types, no any, JSDoc on every export, no new dependencies, no edits to existing files. The report, fourth edition → · The first edition, with the scripts →
one binary · no account · your current subscription is the engine
what it actually is

Merging is one part of it.

Your Claude, Codex, Cursor, Gemini and Grok subscriptions, on one repo, as one agent. antiloki routes each job of an ask to the seat measured best for it, automatically, under one rule: keep at least 80 % of the best measured quality, then the cheapest, then the fastest. The frontier seat is asked only for the parts that need it, and a seat that is out of quota — or nearly — is skipped before the job, not after it fails. It controls what may be touched, records every turn — priced, timed and attributed to a seat — tests every run before it counts, and measures the whole thing in seconds and dollars. The merge is what routing produces — the rest is why it can be trusted.

control

One agent, fenced where you say.

Decide which folders may be written. For every seat the fence is enforced where the work lands: a change beyond it is shown, can be undone in one click, and stops the merge or the delivery. On Claude panes a hook at the tool call, outside the model, also refuses the write before it happens; the other seats have no such hook.

verify

Every run ends in a test.

Every turn is priced and timed, and on Claude panes a hook outside the model logs each tool call. Then QA opens the real page in Chrome and uses the controls the change is about, or sends requests to the server, with checks written as plain expectations — the judge is shown a picture only when a check failed, none could decide, or the styles changed. Unmet claims go back to the worker; when the send-backs are spent, the run becomes a question to you. The judge is chosen on how often it let a broken run through.

route

Frontier money only where it earns it.

Most of an ask is not hard. An ask of size L or smaller goes to one seat, alone, with no planning call; a bigger one is split into jobs, and each job is bound to the seat the benchmark measured best on that work — usually not the expensive one. A seat at 95 % of its window, or out of quota, is skipped before the job starts. There are no per-job pickers to set. That is where the numbers above come from.

learn

It keeps measuring on your machine.

Every run is scored on your machine — which seat, which job, how long, what it cost, whether it passed first time — and filed in the benchmark that ships with the app. The numbers above are the starting point; your runs adjust the shipped scores, by at most a third, and a pairing that was only watched working never displaces one that was benched. The scores stay on the machine.

Adoption is finished and control never arrived: 90 % of professional developers use coding agents weekly and 68 % daily, but only 29 % trust what comes out — down from 40 %, with 3 % who highly trust it. And the failure mode is not theoretical: an agent quoted the rule that forbade it and deleted a production database and its backups, and the deletion call took nine seconds — 93 % of organisations have had an AI-caused incident, 19 % have governance for the next one. Rules in a prompt are not enforcement. JetBrains 2026 · Stack Overflow 2026 · PocketOS · Spacelift 2026

agents

One agent per role, on the seat measured best for the job.

You do not staff a roster. A setup wizard finds the AIs on your machine — Claude, Codex, Cursor, Gemini and Grok, each listed as ready, not signed in or not installed, with the command that fixes it — and asks for one main provider. Then you ask, and the agent takes the role the ask needs (Backender, Frontend designer, Tester, Reviewer and others), arriving with its scope, skills and specialists. The router chooses the seat for every job, from the seats you are signed in to and that still have room in their window; there is no per-job picker to set.

Each agent keeps a queue of asks you can reorder, cancel and re-prompt, with a checkpoint per response to undo to. A run can be paused, resumed or stopped, and the QA can be stopped on its own. The QA test is filmed — a live browser frame while it runs and the kept recording afterwards — and every turn closes with what it did, which seat did it, what it cost, and the QA verdict. A long conversation closes into episodes that consolidate into a summary.

X-ray

The codebase judged by arithmetic, before anything ships.

X-ray is the deterministic half. Every file gets a verdict — a score and why it is not 100, axis by axis — computed from your ASTs and git history, so the same commit scores the same twice and no model is asked for the number. Around it: file reviews with comments at the lines, the rules your codebase lives by checked at the line, a map of the codebase (every file a box; hubs, orphans, cycles), the coverage gaps, and the blast radius of a change before it lands.

It reads TypeScript and, since 2026-10-01, Flow-typed JavaScript: on a copy of React the Flow files it had to skip went from 1,160 to none, and 11 of 1,177 are still unreadable. It says what it did not index — excluded, an unsupported language, a parse that failed, changed since the index. The same facts reach the agent as its context pack. Measured on a wide TypeScript rename (three agents, one run per cell), the map alone was 31 % faster on 48 % fewer input tokens at the same pass rate. On a small change the map alone passed 0 of 3, because the index had left out the file that needed the change; that is why the pack now says what it could not read. On a second repository the map made Codex use more tokens (472k to 1.15M). The report has the cells.

around the work

What sits around the work.

QA

Watch the test run.

A browser frame with the live picture, a cursor gliding to each target, a subtitle for every step, and the kept film to replay afterwards; for a server, a terminal of requests, answers and checks. Before and after screenshots sit in the run story, linked at full size.

control

Pause, resume, stop — or stop only the QA.

Buttons in the run story. While a finished run is still being judged the button reads stop the QA: the work stands, is marked not verified, and nothing is sent back. Each agent also keeps a queue of asks (done, doing, next) you can reorder, cancel and re-prompt, with a checkpoint per response to undo to.

seats

Switch a whole seat off.

Settings has a switch per seat, and a seat you turn off is not called — not by a chat, a terminal, a hand-off, QA, a merge or the advice. Nothing is hard-coded to one vendor: a new chat starts on the first usable seat, and every role has a measured fallback when its first choice is off or out of quota.

limits

Where each seat stands.

Claude and Codex show their 5-hour and weekly windows. Cursor, Gemini and Grok publish no quota, so they show the calls and tokens sent in the last 5 hours and 7 days — a count, never dressed up as a percentage. The saving on screen counts only the current workspace's runs, and drops the old number the moment you switch; the windows and the counts are the account's and this machine's, whatever workspace made the call — and the counts cover only what went through antiloki.

setup

The first run finds your AIs.

The wizard lists the five seats as found on this machine — ready, not signed in, not installed — with the exact command to fix each one, lets you switch off the ones you do not have, and asks for one main provider. It offers a provider only when the seat is real, never merely because a binary exists. New workspaces open in a proper window with a folder browser.

ship

The pull-request button says what a click does.

Before it opens a PR it tells you what will happen; for a repo with no online remote it offers, with your confirmation, to create a private GitHub repo first. A GitHub token can be added in Settings: verified before it is saved, kept owner-only, never shown again, removable.

igryt

The plain parts are not asked of a model.

A yes, or a number, after a proposal is read by a deterministic front in English, Portuguese, Spanish, French, German and Italian. A gate rejects a run that changed nothing, wiped the tree or only deleted files before any model reads it. Gauges decide only when they can point at a fact.

Free for a week. Paid plans are coming.

Planned prices: $200 a year, or $500 once. Checkout is not open yet, so the 7-day trial is the only way in today.

Trial
$07 days
  • everything — agents and X-ray
  • no card, no account
  • a week is real work on a real repo
  • your projects stay when it ends
Download
Annual
$200/ year, planned
  • updates + support while it is active
  • the shipped routing table stays current — that is what renews
  • activated on a machine, deactivated to move it
Coming soon
Lifetime
$500once, planned
  • one payment — no yearly subscription
  • the update terms are set at checkout
  • activated on a machine, deactivated to move it
Coming soon

Prices are USD and may change before checkout opens. The app works offline for up to 14 days between license checks, once a key is activated; the trial makes no check.

Install. Open a repo. Ask.

needs git and at least one of Claude · Codex · Cursor · Gemini · Grok signed in; the browser QA needs Chrome or Chromium — antiloki reads the seat, not the binary, so an unauthenticated CLI is never offered, and the setup wizard shows which seats it found and the command to fix the rest. A merge wants two usable seats: one for the brain and a cheap one for the hands. With one it runs solo, still routed per job. OpenCode, and the Zen and open models behind it, are optional and off by default.

What is antiloki, in one sentence?

A desktop workspace that merges the model subscriptions you already have — Claude, Codex, Cursor, Gemini, Grok — into one coding agent: the seat measured best for each job of an ask, not one model for all of them. It makes your weekly limit last longer and each ask cost less, with folder-level write enforcement, a test at the end of every run and a full audit timeline: the control layer the agents don't ship with.

Does it replace Claude Code, Codex, Cursor, Gemini or Grok?

No — it runs them, twice over: as the agent in its box, and as the engine behind antiloki's own AI (the chat, the reviews, the ultra review), each through its own vendor CLI. Keep your subscriptions; there is no API key to paste.

What does the X-ray do that a linter doesn't?

It judges every file against your rules and its own measurements (tests, coupling, structure, change) and tells you why a file is not 100. Then the ultra review: a $0 survey of the territory, and your engine reading the worst units to report what it can verify — with the fix. The chat sees all of it, so "why is this file 81?" gets a real answer — and the intel handed to an agent says which files it could not read.

Do I need two subscriptions?

For a merge, two usable seats. A merge never puts two frontier models together; it puts cheap hands under a good brain. Google's free tier is capped at a small daily number of agent requests (about twenty when we checked) — enough to try, not a seat to rely on. antiloki only offers a provider you are actually signed in to: it reads the seat, not the installed binary, so a CLI sitting on your machine unauthenticated is never offered as something it can bind. With one seat it runs solo — still routed per job, still tested, just not merged.

Isn't this just worktrees and tmux?

Worktrees are the easy part. The fence enforced where the work lands, the verifier before merge, the QA test and the cost-per-task record are what you can't script in an afternoon — and what the trial is for.

Why is the headline measured against Opus, not a cheaper Claude?

Because Opus is the Claude model a Max subscriber reaches for, and the headline is our best measured result against it on one task: 11.3×, rounded down to 10×. Against raw Sonnet the same pipeline is 3.2× and a Sonnet-planned run 2.1×; against the cost Opus recorded on another day (the 2026-09-22 ladder run, $0.61) it is about 6×. The honest range is 6× to 11×, and the table under the headline shows every arm with its n.

Is +112 % a measured share of my weekly limit?

No. It is a paired task result: on the eight-module task, the run where Claude plans and judges while Composer writes cost $0.1554 against $0.3297 for raw Sonnet, 2.12×. That is dollars at list price standing in for an allowance the vendors do not publish, so how much further your own limit goes depends on your mix of work.

More calls must mean a bigger bill.

More calls, less money: the expensive model is asked only for the part that needs it, and the cheap one does the rest. That is the whole trade, and it is why the bill is on screen for every turn — including the runs where the router declined to split because splitting was not worth it, and the small asks it runs on one seat with no planning call. A failed run reports what it spent and claims no saving.

Can my company install it?

One binary, local, no account, no telemetry. The network calls are: your CLIs' own calls; a license check once a day, once you have activated a key (the trial makes none); an update check against a static file on GitHub Releases; the window's fonts, fetched from Google Fonts (fonts.googleapis.com and fonts.gstatic.com) when it opens; api.openai.com or openrouter.ai only if you set up the optional OpenCode route with your own key; and GitHub only if you add a token or open a pull request. Send that to security.