Which AI For What

Stop asking which AI is best. Ask, for each job, what is the cheapest worker that still clears the quality bar — and who is watching when it ships. This is the one page that answers that for every model on our bench. The roster, the parity table, the billing breakdown and the tier ladder are supporting pages; they all point here.

Lead visual · four tiers, cheapest first

Tier 0MechanicalPulling data, publishing, formatting, renaming, moving files.A plain script. No AI at all.
Tier 1BulkFirst drafts, harvests, transcript passes, repetitive edits across many pages.A cheap, fast model. Kimi K3 or Haiku.
Tier 2JudgmentThe honest score, the real strategy, anything written in a person’s voice.The top model. Claude Opus or Fable.
Tier 3Hands-onLogged-in publishing, anything behind a password, anything irreversible.A person — or an agent on their machine, watched.

Every job on the bench belongs to one of these four. Pick the tier first, then the model. Picking the model first is how a $914 day becomes an $8,700 one.

The ladder is the decision. The model names below are just today’s answer to it — and they change.

Goal: One page a stranger — or their agent — can open to decide which model runs a job, what it costs, and what it is not allowed to do.
Content: The four tiers, the six models with their real strengths and hard limits, the arithmetic of one factory day, and the one number that decides everything.
Targeting: Operators running more than one model, and the agents themselves. Not a vendor comparison and not a review.

The only question

What is the cheapest tier that still clears the quality bar?Not “which model is smartest.” Not “which one is newest.” The smartest model in the world is the wrong answer for renaming four hundred files, and the cheapest one is the wrong answer for telling a client an uncomfortable truth. The tier is the decision. The model is a detail that changes every few months.

This matters more than it sounds, because the difference between routing well and routing badly is not a few percent. On the numbers below it is roughly nine to one.

The four tiers, in detail

Tier The work What runs it Who is watching
0 · Mechanical Pulling data, publishing a file, formatting, renaming, moving things. A plain script. No model, no tokens, no cost. Nobody. It is deterministic.
1 · Bulk First drafts, mention harvests, transcript passes, the same edit across forty pages. A fast cheap model — Kimi K3 for the long ones, Haiku for the short ones. A different model reviews the output before anyone sees it.
2 · Judgment The honest score. The real strategy. Anything written in someone’s voice. Anything a client will quote back at you. The top model. This is the only work that earns top prices. A human reads it before it ships.
3 · Hands-on Logged-in publishing, anything behind a password, anything you cannot undo. A person, or an agent running on their machine where the diff is visible. The person, in real time.

Two rules ride on top of the ladder and never bend. The agent that checks the work is never the model that did the work — a second instance of the same model is not a second opinion, it is the same opinion with an echo. And agents draft; a human sends. Everything short of sending, publishing, spending or merging, an agent may do freely.

The six models, and what each is not

Read the third column first. Knowing what a worker must not touch is worth more than knowing what it is good at, because that is the column that stops the expensive mistakes.

Model Good at Never Roughly
Claude
Builder and judge
Long builds, real voice, the honest score, final QA. The conductor. Its own second opinion. Four Claudes agreeing is one opinion with three echoes. Opus 4.8 $5 in / $25 out · Fable 5 $10 / $50
Codex
Checker
Independent verification, diffs, research, “did we actually prove that.” The writer. A merge authority. A spend authority. Inside a ChatGPT plan
Kimi K3
Grinder
Long-horizon coding, million-token repo work, bulk harvests, overnight batches. A judge or a publisher. It runs under a QA gate and never sees client data. $3 in / $15 out · $0.30 cached
Grok Bot
Always-on desks
Named staff on a shared cloud computer — inbox, CRM, fleet, routines with the laptop closed. In the live cross-model room. No published adapter exists for that protocol. Inside SuperGrok or Cursor Pro+ and above
Gemini
Connected work
Anything where a Google connector is the point — Drive, Sheets, Calendar, Workspace data. A shared memory. Recall from earlier chats is a personal-account feature, not a work one. Google AI Pro from about $20
ChatGPT
Rented chat
A person thinking out loud. Counts as work only once it leaves a receipt in the repo. A production desk. Provider memory is a cache, never the company record. Plus from about $20

Prices move, so treat every figure here as dated. Everything on this page was checked on 25 August 2026. The tiers survive price changes; the model names and the numbers do not. When a price moves enough to change a routing decision, this page changes — not the ladder.

Model, desk, surface — three different things

An earlier version of our roster listed Cursor beside Claude and Grok as if they were the same kind of thing. They are not, and the confusion shows up the moment you add a worker. Claude is a brain. Cursor is a room — two rooms, actually: the watched desktop workshop, and a Cloud Agent on its own Linux VM. They share a GitHub repo. They do not share the Mac Keychain. Map: Cursor Cloud Agents. Kimi is a brain that can sit in that room, or run alone overnight.

Layer The question it answers Members
Model Which brain does the thinking Claude · Codex · Kimi K3 · Grok · Gemini · ChatGPT
Desk Where it sits, and whether a human can see the diff Cursor desktop (watched) · Cursor Cloud Agent (contractor VM) · Claude Code and Cowork · Kimi Code CLI · Grok Bot · a browser
Surface Where the output must land to count as real The repo · the client tool · the live room · the clock

Separate the three and the roster stops being a list to memorise. It becomes four questions with obvious answers: what is the job, which tier is it, which brain is cheapest at that tier, and whose eyes are on it when it ships.

Why Kimi K3 got the grinder seat

Kimi joined the bench in August 2026 as the bulk worker. The case for it is not that it is smarter. It is that grinding and judging are different jobs that are priced identically if you run them on the same model, and the benchmarks split along exactly that line.

Benchmark What it measures Kimi K3 Claude Fable 5
Terminal-Bench 2.1 long terminal sessions 88.3 84.6
SWE-Marathon multi-hour engineering 42.0 35.0
BrowseComp agentic web research 91.2 88.0
MCPMark-Verified tool orchestration 94.5
FrontierSWE hardest engineering 81.2 86.6
GDPval-AA v2 real knowledge work, Elo ~1,684 1,760
Humanity’s Last Exam reasoning under ambiguity 43.5 53.3

The top block is grinding. The bottom block is judgment. That is the whole argument for the seat, and it is also the argument for the gate on it: Moonshot’s own model card warns that K3 “may act excessively proactively on unclear instructions.” A worker that guesses when the brief is vague is fine inside a sandboxed harvest and a liability on a client’s live site. So Kimi builds and harvests. It does not judge, it does not decide voice, and it does not publish.

Cheap per token is not the same as cheap per task. K3 runs at roughly 35 to 39 tokens a second, which is about a third the speed of the fastest models on the bench, and independent testing put a long agentic task at around 56 minutes and $10.57. Price the deliverable, not the million tokens.

What one factory day actually costs

Same work, same volume — roughly 1.5 billion tokens in and 75 million out, with most of the input cached. Only the routing changes.

How it is routed Per day Per month
Everything on Fable 5 $8,700 ~$261,000
Everything on Opus 4.8 $4,350 ~$130,500
Everything on Kimi K3 $2,385 ~$71,600
Everything on Sonnet 5 $1,740 ~$52,200
Routed — 80% Haiku / 15% Sonnet / 5% Fable $1,392 ~$41,800
Routed, plus batch pricing on the overnight bulk $914 ~$27,400

Read that table the right way round. Moving everything to the cheap model saves about 73 percent. Routing properly saves about 90 percent. So the answer was never “switch to Kimi.” It was: stop running grind work on a judgment model. Kimi is a rung on the ladder, not a replacement for it.

The number that decides it

Dollars per passed deliverable. Never dollars per token.A cheap model that fails review twice is more expensive than an expensive model that passes once — in money, and in the hour somebody spent reading the failures. The only honest score for a routing decision is what it cost to get one thing that survived QA.

There is a second trick worth knowing, because it moves the line between tiers. A cheap model given a top model as an advisor — not as a rewriter, just as something it can ask — performed substantially better on hard research than the same cheap model alone, at a fraction of running the whole job on the expensive one. Where that pattern applies, Tier 1 stretches into work that used to need Tier 2.

Routing a job, in order

  1. Can a script do it? If yes, write the script. Tier 0 is free and it never hallucinates.
  2. Does the output need judgment, or just volume? Volume goes to Tier 1. Judgment goes to Tier 2. Most jobs people send to Tier 2 are actually Tier 1 work with one Tier 2 paragraph in them — split them.
  3. Does it touch a login, a payment, or something you cannot undo? That is Tier 3 and it needs a person, whatever else is true about it.
  4. Pick the cheapest model at that tier that has cleared the bar on this kind of work before.
  5. Name the reviewer before you start, and make it a different model than the one doing the work.
  6. Say where the receipt goes. A job with no receipt did not happen, however good the output was.

The effort dial is not a human job

Picking Extra High or Fast in Cursor is choosing the ceiling for that chat, not a routing plan. The agent (or a 15-second traffic-cop chat) must emit a kickoff card on turn one: job shape, ceiling, what is a script, what is bulk, what is judgment, who reviews. Then it runs the card. You do not sit in the picker for every client with three hundred videos.

On Grok 4.6, High and Extra High share the same rate card. Extra High is more hidden reasoning at the same dollars per token. Fast is twice the price and weaker judgment. Default the ceiling to High. Escalate one stuck judgment step to Extra High. Never open Extra High as the saved default for a video factory, a byline sweep, or a site crawl.

A receptionist desk (Grok Bot on the phone) may classify the job and hand back that card. It does not run the factory, and it is not the hourly mail cop. Cursor High plus scripts plus a cheap bulk lane is the factory. JVZoo on 26 August 2026 proved it: Extra High was the wrong kickoff; leftover Staff bylines were a REST script, not a smarter model. Phrase on this page: CEILING-NOT-FACTORY-2026-08-26.

What would change this page

Routing is a live decision, not a belief. These are the things that would move it, and they are the things the reprice job watches for:

  • A price change big enough to flip a tier — a model getting cheaper does not matter unless it crosses the line where the work would move.
  • A month of receipts showing dollars per passed deliverable is worse on the cheap lane than the routed one. That is the number that decides, and it beats any benchmark.
  • A model shipping something structural — real cross-session memory, a much larger context, a new safety gate — rather than a few points on a leaderboard.
  • A hosting change that moves where data physically sits, which can relax or tighten what a lane is allowed to see.

Questions people actually ask

Which one should I buy if I only buy one?

The top model, on the cheapest plan that covers your actual usage. One good judgment model plus scripts will carry a small operation a long way. Add a bulk lane when you can name the specific job that is costing you too much — not before, because a second subscription with no job attached is just a second bill.

Should I pick Extra High when a partner has hundreds of videos?

No. Inventory and captions are a script. First-pass how-tos are a cheap bulk lane only if the first file returns in about a minute; otherwise the conductor writes from the captions. ENHANCE-versus-NEW and the honest overlay stay on High. Extra High is the exception after High stalls, not the way you start a 300-video client.

Is the newest model automatically the best choice?

No, and this is the most expensive habit in the whole field. The newest model wins on a leaderboard. The right model wins on dollars per passed deliverable for one specific job. Those are different contests and only one of them is yours.

Can I just use one model for everything to keep it simple?

You can, and it is a defensible choice while you are small. Understand what it costs: on the day priced above, simplicity is roughly the difference between $914 and $8,700. At some volume that stops being simplicity and starts being a decision you are making by not making it.

Why does the checker have to be a different model?

Because a model is wrong in a consistent direction. Ask it to check itself and it agrees with itself, confidently, for the same reason it was wrong the first time. A different model has different blind spots, which is the entire point.

What about data going overseas?

It depends on the lane, and the rule does not bend for convenience. Public information and our own files can go anywhere. Client personal data, credentials and unpublished client work stay on lanes we have cleared for them. This is the same rule you would apply to any contractor and it is not about any one vendor.

Where do the supporting pages fit?

This page is the decision. The roster says who sits in which chair, the parity page says what each runtime can actually run, the billing page says how each vendor meters you, and the tier ladder is the short version of the four tiers above. Each of those goes deeper on one column of this page and links back here.

The one line to keep

The cheapest tier that clears the bar, checked by a model that did not do the work, shipped by a human, with a receipt someone else can open. Everything else on this page is detail that will be out of date before that sentence is.

Scroll to Top