Which AI model, and how hard should it think?
Colton's field guide to Claude and ChatGPT: when to use what, when to bump up or down, when to stop and escalate, and how to stop paying for tokens you do not need. Every model name, price and limit below was checked against the vendor's own pages on 2026-10-04; the links are at the bottom.
The one-minute version
The Claude lineup
Four models on one ladder. Cost is the API output price per million tokens, indexed to Haiku 4.5 as 1x; on a subscription the same ratios roughly decide how fast you use up your limit. The effort gauge shows each model's default. Anthropic suggests starting with Opus 5.5 for most workloads.
HAIKU 4.5
Fast and cheapThe fastest, cheapest model, for simple, well-defined jobs that need no judgment.
- "I want a quick summary of an email thread."
- "I need a batch of files renamed or reformatted the same way."
- "I need fields pulled out of a spreadsheet or a log."
Good to know
Haiku has no effort setting and the smallest working memory of the four, so it is built for short, well-scoped jobs and helper work inside a bigger task. Of the four, it has the nearest retirement commitment, so expect a newer small model.
SONNET 5.5
WorkhorseThe best mix of speed and intelligence. Your everyday default.
- "I want my day-to-day coding across apps and sites handled."
- "I need clean, client-ready copy and briefs written."
- "I need my notes, Slack and email pulled into one answer."
Good to know
Start here. Raising Sonnet's effort is often cheaper than switching up to Opus, so try the dial before the tier.
OPUS 5.5
Heavy reasoningLong-running coding and knowledge work. Anthropic's suggested starting point for most workloads.
- "I want a messy folder or repo structure solved end to end."
- "I need conflicting information across systems reconciled."
- "I need one-shot output I will not have time to double-check."
- "I want a real, playable in-browser game built."
Good to know
Planning, multi-part analysis, anything that crosses several files, systems or documents. Opus earns its price at real effort; see the trap below before running it low.
FABLE 5.1
EscalationThe deepest reasoning and long, many-step work. Slowest. The break-glass option.
- "I want a fresh run on a problem that stalled elsewhere."
- "I need a breakthrough after two or three failed attempts."
- "I want a genuinely polished interactive build: physics, animation, all of it."
- "I need this rarely. It is not a daily driver."
Good to know
Reach for it when the other models at high effort still fall short. On Pro it runs on usage credits; on Max the plan table lists it at half of the weekly limits.
Effort: the dial inside each model
Picking a model is choosing the engine. Effort is how hard you let it think. In Claude Code, type /effort to change it and /model to switch models. Haiku 4.5 has no effort setting.
When to bump up, when to bump down
BUMP UP: one tier higher, or more effort, when
- The output will be relied on by someone who matters: a board memo, contract language, anything with your name on it.
- The task touches several systems, files or documents at once.
- You have never done this kind of task and there is no written recipe yet.
- The input is long or contradictory: a 200-page RFP, sources that disagree.
- The obvious fix did not work and you are now debugging the debugging.
- The smaller model got the same step wrong twice.
- You need a plan, a design or a decision, not just an answer.
BUMP DOWN: one tier lower, or less effort, when
- The task repeats a template or checklist you already wrote.
- You are summarizing, translating, rewording, classifying or tagging.
- You want a quick fact, a definition or a one-line answer.
- It is high volume: hundreds of rows, emails or files, each needing the same small treatment.
- It is a helper job inside a bigger task, like searching files or reading pages.
- The hard part is done. A strong model writes the plan, a fast model carries out the steps.
- You are brainstorming out loud and speed matters more than polish.
Warning signs you picked the wrong model
TOO SMALL
- You give the same correction twice.
- It asks for something you already told it.
- It builds something new instead of finding the thing that already exists.
- Lots of confident activity, no finished piece after about 15 minutes.
- Three failed attempts at the same step.
- Answers turn generic and hedged just when the question got specific.
Fix: stop, ask for a short summary of where things stand, and start that summary in a new session one tier up.
TOO BIG
- You wait a long time for answers to easy questions.
- A yes-or-no question comes back as an essay.
- You hit your usage limit midweek doing routine work.
- It re-plans a job that only needed doing.
Fix: lower the effort first. If it is still slow, drop a tier for that kind of task.
The escalation flow
The ChatGPT lineup
OpenAI's current family is GPT-6, in three sizes. Older advice about "GPT-5" or "Thinking" modes is out of date. Cost is the API output price, indexed to Luna as 1x.
GPT-6 LUNA
EfficientFor focused, high-volume tasks.
- Small edits, well-scoped questions, simple extraction, frequent automations.
GPT-6.1 SOL
Everyday defaultEveryday work that needs judgment. OpenAI's recommended default in Codex.
- Coding, research and workflows where completeness matters.
GPT-6 ASTRA
Most capableFor the hardest end-to-end work. Pro plans only.
- Ambitious projects that need broad context and complete results, when cost and wait time are not the constraint.
Plans and what they cost
| Plan | Claude | ChatGPT |
|---|---|---|
| Free | $0: Sonnet and Haiku, a baseline amount of use, no Claude Code. | $0, and Go at $8 a month: GPT-6 Luna at standard speed, limited tools. |
| Pro / Plus | Pro, $20 a month ($17 billed yearly): at least 5x Free per five-hour session, adds Opus and Claude Code. Fable on usage credits. | Plus, $20 a month: Sol and Luna, expanded Codex, projects and scheduled tasks. |
| Max / Pro | Max, from $100 a month: 5x or 20x Pro per five-hour session. Fable at half of weekly limits. | Pro, $100, $200 or $500 a month: far more Codex, no five-hour cap, Astra at the top of the range. |
| Teams | Team Standard $25 a seat ($20 yearly), Premium $125 ($100 yearly). Enterprise $20 a seat plus usage at API rates. | Business $20 a user billed yearly ($25 monthly), two users minimum. Enterprise and Edu: contact sales. |
How usage limits and overage billing work
CLAUDE
- Every plan has a limit that resets on a rolling five-hour window; paid plans add weekly limits.
- Chat on web, desktop and phone and Claude Code all draw from one shared pool. No fixed message count: long chats, bigger models and heavy features use it faster.
- At the limit: wait for the reset, move up a plan, or turn on usage credits, billed separately at standard API rates.
- Set the cap first: Settings, then Usage, sets a monthly spending limit. In Claude Code,
/usageshows where you stand.
CHATGPT
- Everyday text chats are listed as unlimited, within abuse guardrails. Codex is metered per five-hour period: on Plus roughly 15 to 150 Sol messages or 350 to 3,000 Luna messages.
- On Plus and Pro you can buy more credits past the included limits. Business and Enterprise buy workspace credits.
- Codex with an API key skips the plan limits and bills every token at API prices.
Fresh session or long thread?
Every new message sends the whole conversation back to the model. A long thread costs more per message, and older details start to blur. The top models hold about 1M tokens, roughly 555,000 words: a lot of room, and a lot of money to re-send on every turn.
- Start fresh for any new task. In Claude Code,
/clearstarts over. In ChatGPT or the Claude app, open a new chat. - Compact when you must keep going.
/compactsummarizes the thread and frees space;/contextshows what is filling it. - Hand off, do not drag on. Ask for a ten-line summary of decisions and next steps, then paste it into the new session.
- Put standing material in a project. Files and instructions saved to a project load once instead of being pasted into every chat.
- Keep standing instructions short. They load at the start of every session. Slim your CLAUDE.md shows how.
Cheap habits that add up
- Text beats pictures. Paste the words instead of a screenshot. Browser automation and screenshots are the priciest way to get an answer.
- Ask for the answer first. "Lead with the result, detail only if I ask" shortens every reply.
- Batch similar requests into one message instead of ten.
- Reuse, do not repeat. A cached prompt read on Anthropic's API costs 10 percent of the normal input price (5 percent on Opus 5.5, 2.5 percent on Fable 5.1). Batch jobs are half price.
- Big model plans, small model does. The single biggest saving for repeat work.
Which app: Design, PowerPoint or Code
CLAUDE DESIGN
A visual canvas for decks, one-pagers and prototypes built from your brand system.
- "I want an on-brand deck built from our templates and colors."
- "I need a one-pager or microsite I can pass around, not code I will maintain."
- "I want my meeting notes turned into a polished visual, no repo involved."
CLAUDE FOR POWERPOINT
Works inside a .pptx file you already have open.
- "I already have a .pptx template and need it filled in."
- "I need slides updated in a deck someone else built."
- "I want this to stay a real PowerPoint file."
CLAUDE CODE
An agent that works in a real repo, on real code, with full model and effort control.
- "I want this to actually deploy: the site, the app."
- "I need it to read and write across files in an existing codebase."
- "I need direct control over which model and effort run the job."
Sources, checked 2026-10-04
- Anthropic: Models overview (names, API prices, default effort)
- Anthropic: Claude plans and pricing (plans, five-hour and weekly limits, usage credits)
- Anthropic: Effort levels
- Anthropic Help: Usage credits for paid plans
- Claude Code: commands (/clear, /compact, /context, /effort, /usage)
- Anthropic: Model deprecations (retirement dates)
- OpenAI: ChatGPT pricing (plans and models)
- OpenAI: API models (GPT-6 Astra, Sol, Luna, prices)
- OpenAI: Codex plans and credits
- OpenAI: Codex models (plan availability, GPT-5.5 retirement)
- OpenAI: Choosing a model (reasoning levels, starting points)
Vendors change names, prices and limits often. If something here disagrees with the vendor's page, the vendor's page wins.