When should I use each AI model?

Match the model to the shape of the work, not to the leaderboard. Reach for a frontier model like GPT-5.6 Sol or Claude Fable 5 when the path is unclear and a solution has to be discovered, a mid tier like GPT-5.6 Terra or Claude Opus 4.8 when the outcome is known but judgment is still needed, and a fast tier like GPT-5.6 Luna or Claude Sonnet 5 when the steps are defined and you just need them ticked off. The smartest model can cost five to eight times more per token, so reserve it for work that genuinely needs it.

A hand-drawn three-column table titled 'Model Choice by Work Type' with the BusyWork Dispatch logo in the corner. The columns are Models, Use When, and Desired Outcome. Row one: GPT-5.6 Sol and Claude Fable 5, use when the goal matters but the path is still unclear, desired outcome a strategy or solution needs to be discovered. Row two: GPT-5.6 Terra and Claude Opus 4.8, use when the outcome is known but judgment is still required, desired outcome the work is verifiable but not fully mechanical. Row three: GPT-5.6 Luna and Claude Sonnet 5, use when the steps are defined and the guardrails are in place, desired outcome a checklist can be tested and ticked off.

With Claude Fable 5 and OpenAI's new GPT-5.6 range (Sol, Terra and Luna) landing within weeks of each other, the question I get most from leaders is simple: which one do I actually use?

The instinct is to default to the smartest model for everything. That is the most expensive habit you can build.

The better approach is to match the model to the shape of the work, not to the top of the leaderboard. Here is how I think about it.

The one question that picks the model

Before you pick a model, answer one question about the task: how clear is the path from here to done?

That single question sorts almost every piece of work into one of three tiers.

Tier 1: the path is unclear, so you need to discover the solution

Use a frontier model when the goal matters but the path is still unclear, and the desired outcome is a strategy or solution that has to be discovered.

This is GPT-5.6 Sol and Claude Fable 5 territory.

These are the models built for the hardest problems: novel strategy, complex multi-step coding, security research, and long autonomous work where the model has to figure out the route itself. OpenAI positions Sol for exactly this, the hardest reasoning, coding and agentic tasks, and early reporting has Sol posting record scores on agentic coding benchmarks like TerminalBench 2.1.

Fable 5 tells the same story from the Anthropic side. It leads the major coding and agent benchmarks, and the headline example doing the rounds is Stripe, which reportedly ran a codebase-wide migration across 50 million lines of Ruby and had Fable 5 finish it in about a day, work a full engineering team would have spent months on by hand.

Reach for this tier when you genuinely do not know the answer yet and the quality of the thinking is what matters. That is where the extra capability earns its price.

Tier 2: the outcome is known, but it still needs judgment

Use a mid tier model when the outcome is known but judgment is still required, and the work is verifiable but not fully mechanical.

This is GPT-5.6 Terra and Claude Opus 4.8 territory.

Here you know what good looks like. The model is not inventing a strategy, it is applying judgment to get to a known outcome: customer support with real nuance, internal tools, document analysis, drafting that a human will check, most day-to-day coding. OpenAI pitches Terra as the balanced everyday model that matches the previous generation's quality at roughly half the cost. Opus 4.8 sits in the same band for people who want that last bit of accuracy on hard, multi-step reasoning and deep coding.

This is the workhorse tier. Most business work lives here, not at the frontier.

Tier 3: the steps are defined, so just tick them off

Use a fast tier model when the steps are defined and the guardrails are in place, and the outcome is a checklist that can be tested and ticked off.

This is GPT-5.6 Luna and Claude Sonnet 5 territory.

When the work is well specified and repeatable, you do not need a genius, you need speed and low cost at volume. Summarising, drafting, classifying, routine automation, high-volume customer-facing flows: this is what Luna and Sonnet 5 are built for. OpenAI describes Luna as the fast, cheap model for exactly this kind of everyday work, and Sonnet 5 is Anthropic's high-volume production model for customer-facing agents and content at scale.

If a task can be turned into a checklist and tested, run it on the cheapest model that passes the test.

Why you don't always use the smartest model

The smartest model is not just a bit more expensive. It is multiples more expensive.

At the pricing published around launch, the output tokens that do the heavy lifting land roughly like this per million tokens: Claude Fable 5 near $50 and GPT-5.6 Sol near $30 at the top, Opus 4.8 near $25 and the mid tier around $15 in the middle, and GPT-5.6 Luna near $6 at the bottom. So the frontier tier can cost five to eight times more per token than the fast tier for the same volume of output.

Now put that against the work. If a task is a defined checklist, a frontier model does not make it more correct, it just makes it cost eight times as much and often run slower. You are paying a premium for reasoning the task never needed.

Two more traps worth knowing:

  • More tokens, not just more dollars per token. Early testing of Sonnet 5 found it can use around 30 percent more tokens than previous models for the same task, so a lower per-token price does not always mean a lower total bill. Watch the total, not the sticker.
  • Smarter can mean over-eager. OpenAI's own system card notes GPT-5.6 is more likely than the previous version to act beyond what the user asked. On a locked-down, checklist task that is exactly the wrong trait, another reason the biggest model is not automatically the safest choice.

The point is not to be cheap. It is to spend your capability budget where it changes the outcome, and stop paying frontier prices for checklist work.

Bottom line: pick the model by how clear the path is, not by how clever the model is. Discover the solution on the frontier tier, apply judgment on the mid tier, and tick off defined work on the fast tier, so the smartest model is reserved for the work that actually needs it.

If you want help turning this into a simple rule your team can follow inside your AI loops, that is exactly the kind of thing I cover in my 60-minute AI Loops workshop. It is free and hands-on.

Keep reading — it's free

Pop in your email to keep reading and join AI Dispatch, our newsletter of practical advice for leaders scaling AI. Unsubscribe anytime.

Related