Skip to the content
Built by Colton

Updated Tue / 2026-10-06 / 4:55 AM CT

AI model selection · token cheat sheetVendor facts checked 2026-10-04

Which AI model, and how hard should it think?

Colton's field guide to Claude and ChatGPT: when to use what, when to bump up or down, when to stop and escalate, and how to stop paying for tokens you do not need. Every model name, price and limit below was checked against the vendor's own pages on 2026-10-04; the links are at the bottom.

The one-minute version

Start in the middleClaude Sonnet 5.5 or GPT-6.1 Sol handles most everyday work well and stretches your plan the furthest.
Bump up for stakesGo up a tier, or raise effort, when the output will be relied on, or the task has many moving parts.
Bump down for repeatsDrop a tier once a task has a written recipe and has already worked once on the bigger model.
New task, new chatLong threads re-send everything every turn. A fresh session is cheaper and sharper.

The Claude lineup

Four models on one ladder. Cost is the API output price per million tokens, indexed to Haiku 4.5 as 1x; on a subscription the same ratios roughly decide how fast you use up your limit. The effort gauge shows each model's default. Anthropic suggests starting with Opus 5.5 for most workloads.

HAIKU 4.5

Fast and cheap

The fastest, cheapest model, for simple, well-defined jobs that need no judgment.

Cost1x · $1 / $5
EffortNo setting
LMHXHMAX
200K tokens of contextRetires Oct 15, 2026
  • "I want a quick summary of an email thread."
  • "I need a batch of files renamed or reformatted the same way."
  • "I need fields pulled out of a spreadsheet or a log."
Good to know

Haiku has no effort setting and the smallest working memory of the four, so it is built for short, well-scoped jobs and helper work inside a bigger task. Of the four, it has the nearest retirement commitment, so expect a newer small model.

SONNET 5.5

Workhorse

The best mix of speed and intelligence. Your everyday default.

Cost2x · $2 / $10
EffortDefault high
LMHXHMAX
1M tokens of contextAll five effort levels
  • "I want my day-to-day coding across apps and sites handled."
  • "I need clean, client-ready copy and briefs written."
  • "I need my notes, Slack and email pulled into one answer."
Good to know

Start here. Raising Sonnet's effort is often cheaper than switching up to Opus, so try the dial before the tier.

OPUS 5.5

Heavy reasoning

Long-running coding and knowledge work. Anthropic's suggested starting point for most workloads.

Cost4x · $4 / $20
EffortDefault medium
LMHXHMAX
1M tokens of contextAll five effort levels
  • "I want a messy folder or repo structure solved end to end."
  • "I need conflicting information across systems reconciled."
  • "I need one-shot output I will not have time to double-check."
  • "I want a real, playable in-browser game built."
Good to know

Planning, multi-part analysis, anything that crosses several files, systems or documents. Opus earns its price at real effort; see the trap below before running it low.

FABLE 5.1

Escalation

The deepest reasoning and long, many-step work. Slowest. The break-glass option.

Cost10x · $10 / $50
EffortDefault high
LMHXHMAX
1M tokens of contextUsage credits on Pro
  • "I want a fresh run on a problem that stalled elsewhere."
  • "I need a breakthrough after two or three failed attempts."
  • "I want a genuinely polished interactive build: physics, animation, all of it."
  • "I need this rarely. It is not a daily driver."
Good to know

Reach for it when the other models at high effort still fall short. On Pro it runs on usage credits; on Max the plan table lists it at half of the weekly limits.

Effort: the dial inside each model

Picking a model is choosing the engine. Effort is how hard you let it think. In Claude Code, type /effort to change it and /model to switch models. Haiku 4.5 has no effort setting.

LOWFastest and cheapest, some loss of depth. Simple tasks and helper jobs.
MEDIUMBalanced. Opus 5.5's default. Day-to-day work that still needs judgment.
HIGHSpends what the task needs. Fable and Sonnet's default. Complex reasoning and hard coding.
EXTRA HIGHExtended thinking for long jobs of 30 minutes or more.
MAXNo limit on thinking. Costs the most for the smallest extra gain. Frontier problems only.
Decision trap to avoidDo not run Opus at low effort as a cheaper Sonnet. Opus earns its price at real effort; at low it skips the exploring and self-checking you are paying for. If a task cannot justify Opus at real effort, it was not an Opus task: use Sonnet.

When to bump up, when to bump down

BUMP UP: one tier higher, or more effort, when

  • The output will be relied on by someone who matters: a board memo, contract language, anything with your name on it.
  • The task touches several systems, files or documents at once.
  • You have never done this kind of task and there is no written recipe yet.
  • The input is long or contradictory: a 200-page RFP, sources that disagree.
  • The obvious fix did not work and you are now debugging the debugging.
  • The smaller model got the same step wrong twice.
  • You need a plan, a design or a decision, not just an answer.

BUMP DOWN: one tier lower, or less effort, when

  • The task repeats a template or checklist you already wrote.
  • You are summarizing, translating, rewording, classifying or tagging.
  • You want a quick fact, a definition or a one-line answer.
  • It is high volume: hundreds of rows, emails or files, each needing the same small treatment.
  • It is a helper job inside a bigger task, like searching files or reading pages.
  • The hard part is done. A strong model writes the plan, a fast model carries out the steps.
  • You are brainstorming out loud and speed matters more than polish.
The one rule that saves the most griefOnly move a task down after it has worked once on the bigger model and the steps are written down. Do not ask the smaller model whether it is good enough; a model cannot judge work above its own level. Judge it by the result.

Warning signs you picked the wrong model

TOO SMALL

  • You give the same correction twice.
  • It asks for something you already told it.
  • It builds something new instead of finding the thing that already exists.
  • Lots of confident activity, no finished piece after about 15 minutes.
  • Three failed attempts at the same step.
  • Answers turn generic and hedged just when the question got specific.

Fix: stop, ask for a short summary of where things stand, and start that summary in a new session one tier up.

TOO BIG

  • You wait a long time for answers to easy questions.
  • A yes-or-no question comes back as an essay.
  • You hit your usage limit midweek doing routine work.
  • It re-plans a job that only needed doing.

Fix: lower the effort first. If it is still slow, drop a tier for that kind of task.

The escalation flow

Start on Sonnet or Opusat the right effort
Stalls two or three timesgoing in circles, contradicting itself
Get a brieften lines: decisions and next steps
New sessionrun it fresh one tier up, on Fable

The ChatGPT lineup

OpenAI's current family is GPT-6, in three sizes. Older advice about "GPT-5" or "Thinking" modes is out of date. Cost is the API output price, indexed to Luna as 1x.

GPT-6 LUNA

Efficient

For focused, high-volume tasks.

Cost1x · $0.10 / $0.50
  • Small edits, well-scoped questions, simple extraction, frequent automations.

GPT-6.1 SOL

Everyday default

Everyday work that needs judgment. OpenAI's recommended default in Codex.

Cost20x · $2 / $10
  • Coding, research and workflows where completeness matters.

GPT-6 ASTRA

Most capable

For the hardest end-to-end work. Pro plans only.

Cost100x · $10 / $50
  • Ambitious projects that need broad context and complete results, when cost and wait time are not the constraint.
The effort names do not matchCodex lists low, medium and extra high, not Claude's five. Start Sol at medium for complex technical work, Luna at low for scoped edits, and Astra at low for concise writing; save extra high for high-stakes deliverables. GPT-5.5 retires from ChatGPT and Codex on October 14, 2026: use Sol on paid plans, Luna on Free and Go.

Plans and what they cost

PlanClaudeChatGPT
Free$0: Sonnet and Haiku, a baseline amount of use, no Claude Code.$0, and Go at $8 a month: GPT-6 Luna at standard speed, limited tools.
Pro / PlusPro, $20 a month ($17 billed yearly): at least 5x Free per five-hour session, adds Opus and Claude Code. Fable on usage credits.Plus, $20 a month: Sol and Luna, expanded Codex, projects and scheduled tasks.
Max / ProMax, from $100 a month: 5x or 20x Pro per five-hour session. Fable at half of weekly limits.Pro, $100, $200 or $500 a month: far more Codex, no five-hour cap, Astra at the top of the range.
TeamsTeam Standard $25 a seat ($20 yearly), Premium $125 ($100 yearly). Enterprise $20 a seat plus usage at API rates.Business $20 a user billed yearly ($25 monthly), two users minimum. Enterprise and Edu: contact sales.

How usage limits and overage billing work

CLAUDE

  • Every plan has a limit that resets on a rolling five-hour window; paid plans add weekly limits.
  • Chat on web, desktop and phone and Claude Code all draw from one shared pool. No fixed message count: long chats, bigger models and heavy features use it faster.
  • At the limit: wait for the reset, move up a plan, or turn on usage credits, billed separately at standard API rates.
  • Set the cap first: Settings, then Usage, sets a monthly spending limit. In Claude Code, /usage shows where you stand.

CHATGPT

  • Everyday text chats are listed as unlimited, within abuse guardrails. Codex is metered per five-hour period: on Plus roughly 15 to 150 Sol messages or 350 to 3,000 Luna messages.
  • On Plus and Pro you can buy more credits past the included limits. Business and Enterprise buy workspace credits.
  • Codex with an API key skips the plan limits and bills every token at API prices.
Where surprise bills come fromOverage is charged at the pay-as-you-go API rate, not the flat subscription rate. A monthly cap turns a surprise into a pause.

Fresh session or long thread?

Every new message sends the whole conversation back to the model. A long thread costs more per message, and older details start to blur. The top models hold about 1M tokens, roughly 555,000 words: a lot of room, and a lot of money to re-send on every turn.

  • Start fresh for any new task. In Claude Code, /clear starts over. In ChatGPT or the Claude app, open a new chat.
  • Compact when you must keep going. /compact summarizes the thread and frees space; /context shows what is filling it.
  • Hand off, do not drag on. Ask for a ten-line summary of decisions and next steps, then paste it into the new session.
  • Put standing material in a project. Files and instructions saved to a project load once instead of being pasted into every chat.
  • Keep standing instructions short. They load at the start of every session. Slim your CLAUDE.md shows how.

Cheap habits that add up

  • Text beats pictures. Paste the words instead of a screenshot. Browser automation and screenshots are the priciest way to get an answer.
  • Ask for the answer first. "Lead with the result, detail only if I ask" shortens every reply.
  • Batch similar requests into one message instead of ten.
  • Reuse, do not repeat. A cached prompt read on Anthropic's API costs 10 percent of the normal input price (5 percent on Opus 5.5, 2.5 percent on Fable 5.1). Batch jobs are half price.
  • Big model plans, small model does. The single biggest saving for repeat work.

Which app: Design, PowerPoint or Code

How this nests with the model pickerA different axis. Model choice is the engine; the app is the surface it runs on. Design and PowerPoint pick a model for you; Claude Code lets you pick model and effort yourself. Decide the app first by what you are producing, then, inside Claude Code, use the ladder above.

CLAUDE DESIGN

A visual canvas for decks, one-pagers and prototypes built from your brand system.

  • "I want an on-brand deck built from our templates and colors."
  • "I need a one-pager or microsite I can pass around, not code I will maintain."
  • "I want my meeting notes turned into a polished visual, no repo involved."

CLAUDE FOR POWERPOINT

Works inside a .pptx file you already have open.

  • "I already have a .pptx template and need it filled in."
  • "I need slides updated in a deck someone else built."
  • "I want this to stay a real PowerPoint file."

CLAUDE CODE

An agent that works in a real repo, on real code, with full model and effort control.

  • "I want this to actually deploy: the site, the app."
  • "I need it to read and write across files in an existing codebase."
  • "I need direct control over which model and effort run the job."
Handoff pattern worth knowingClaude Design exports straight to Claude Code. Prototype the look in Design, then hand it to Code when it is ready to become a real, maintained page. Do not build production pages in Design or sketch loose visual ideas in Code; each tool fights you outside its lane.

Sources, checked 2026-10-04

Vendors change names, prices and limits often. If something here disagrees with the vendor's page, the vendor's page wins.

AI MODEL SELECTION AND TOKEN CHEAT SHEET · builtbycolton.com/modelsGuide v3.0 · facts checked 2026-10-04