Which Claude Model Should You Use? The Guide I Wish I’d Read First

Haiku, Sonnet, Opus, Fable, the effort dial nobody found: And why almost everyone gets it wrong in both directions at once.Photo by Brecht Corbeel on UnsplashI had this question.I kept switching models in the picker without really knowing what changed. Sometimes I’d go with the expensive one…

Haiku, Sonnet, Opus, Fable, the effort dial nobody found: And why almost everyone gets it wrong in both directions at once.Photo by Brecht Corbeel on UnsplashI had this question.I kept switching models in the picker without really knowing what changed. Sometimes I’d go with the expensive one because it felt safer. Sometimes I’d stay on the fast one because I was in a hurry.I never knew if I was doing it right. So, I stopped and actually looked into it: documentation, tests, side-by-side comparisons. And found out I’d been doing it wrong. Both ways, depending on the day.This is the piece I wish I’d read first.There’s a pattern hereAnd it’s funnier than it looks. Almost everyone makes one of two mistakes. People assume they’re opposite mistakes. They’re not. It’s the same mistake, turned inside out.Mistake 1: using the most expensive model to summarize an emailThis is the most common one among people who’ve discovered there’s a model picker.The logic seems airtight: if there’s a better model, why not always use the better model? The output comes back good. Nobody complains.Except summarizing a four-paragraph email has no ambiguity in it. There is no “deeper reading” of that email. There’s one right answer, and every model in the lineup gets it.You paid for reasoning the task never asked for. You waited longer. And the part nobody sees: you burned a slice of your usage limit on something that could have cost a third as much.Then Thursday rolls around, you hit the limit, and you decide the problem is your plan.Mistake 2: brainstorming on the fast modelThis one’s the inverse, and it’s the expensive one.You open a conversation to decide between two job offers. Or to figure out why sales dropped last quarter. Or to think through how to handle a hard conversation with someone on your team.Back comes something organized. Well written. Nice bullets. And shallow.The problem is you have no way to notice. On a task with a right answer, when the model gets it wrong, you see it. On a judgment task, when the model hands you the obvious, that also looks like an answer. Nicely formatted, even.Nobody tells you the analysis skipped three alternatives. You get an analysis. And you decide based on it.What both mistakes shareIn both cases, the person picked a model based on how much they wanted to spend.The right move is to pick based on what kind of question they’re asking.The question that replaces the comparison tableThere’s one question that settles this, and it has nothing to do with price:Does this task have one right answer, or several defensible ones?One right answer → fast model. Extract, summarize, classify, reformat, translate. There is no “smarter” way to pull the dates out of a contract.Several defensible answers → heavy model. Decide, prioritize, weigh risk, choose between paths, find what’s missing. Here the depth of the reasoning is the product.Hold onto that question. It’ll still work long after every name below is out of date.The quick guide to the four modelsAnthropic organizes its models into four families. The family names are stable; the version numbers change every few months.Haiku: the fastest and lightestBuilt for everyday requests. Near-instant, minimal draw on your usage limit. In practice:Summarizing that seven-paragraph email from your kid’s schoolTurning a meeting transcript into a list of decisionsPulling the dates and dollar amounts out of a leaseSorting 200 survey responses into five categoriesTranslating somethingCleaning up the grammar in a document before you send itTurning a recipe into a shopping listThe test: if you already know exactly what format you want back, it’s Haiku.Sonnet: the workhorseThis is the default. Strong enough reasoning for most real work, fast enough that you can work alongside it in real time. In practice:Writing the awkward email to the client who’s late payingBuilding the budget spreadsheet for the kitchen renovationOutlining a class or a presentationCoding, debugging, refactoringWriting the product descriptions for your storeTurning a 30-page report into an executive summaryBuilding a project timeline out of a messy list of tasksThe test: if you’ve already decided what to do and you need it executed well, it’s Sonnet.When in doubt, start here.Opus: the specialist in hard problemsBuilt for problems that need sustained reasoning. It draws a lot more from your limit, so it isn’t for just anything. In practice:Comparing two job offers with different salary, benefits, and riskReading the whole contract and flagging what’s written against youFiguring out why sales dropped when there are five possible explanationsComparing three health plans whose coverage doesn’t line upTesting whether a business idea holds, and where it breaks firstStructuring the board presentation, deciding what’s in and what’s cutAnything you already tried on Sonnet that didn’t come out rightThe test: if you’re asking for an opinion rather than a deliverable, it’s Opus.Fable: the one that works aloneTop of the line. The difference isn’t only “smarter”: it’s built for long, autonomous work. You describe the outcome you want; it plans the steps and checks its own work, with fewer check-ins along the way. In practice:Research that spans dozens of sources and has to reconcile them at the endRewriting a whole system, not a functionThe three-hour job you want running while you do something elseThe test: if the request is “go do this and tell me when it’s done,” it’s Fable. If you’re going to watch it closely, it isn’t.The second dial almost nobody foundPicking the model decides who does the task.There’s a second control, called effort, that decides how much time that person gets. It sits right next to the model picker, and most people have never opened it.Five levels. Worth understanding all five, even if you only ever touch one or two.Low effort — Claude takes fewer detours. Less thinking before it writes, no preamble, straight to the point. This is the right setting when the task already has an obvious answer: reformatting a list, pulling data out of a table, translating a paragraph. You gain speed and save your limit. You lose a little capability, and on this kind of task you’ll never see it.Medium effort — The middle ground. Good for tasks with some nuance but no real difficulty, like a meeting summary that has to understand what was decided, not just what was said. It’s the smart pick when you’re running the same task dozens of times and want your limit to last the week.High effort — The default. If you’ve never touched this dial, this is where you are right now. And most of the time it’s where you should stay; it’s tuned to be the right answer in the majority of cases.Extra effort — For long, multi-step work that runs for a while. Here effort isn’t only “think harder”; it’s Claude running more searches, more checks, and more attempts before deciding it’s finished. It makes a real difference on deep research and autonomous work. On a short question, it makes no difference at all.Max effort — No ceiling on spend. It looks like the right button for “I want the best possible result,” and it rarely is. Anthropic’s own documentation warns that on most workloads it adds high cost for a small quality gain, and that on more structured tasks it can overthink and make things worse. Save it for genuinely hard problems.The most common mistake with effortAssuming that lowering effort makes the response shorter.It doesn’t. Effort controls how much Claude thinks, not how much it writes. Two different things, and mixing them up sends you to the wrong dial. If you want a short answer, ask for a short answer. That works better than any setting.Worth knowing, too: not every model has this control. Haiku doesn’t. It’s already the fast one by design.So how does this work in Cowork, Design, and Code?Picking a model and an effort level is one thing. Picking which door you walk through is another, and it changes the outcome more than switching models does.Most people only ever use the chat door. There are three others worth knowing.Claude Code: when the work is actually codeIf the task involves code files on your machine, working across a real project, not asking for a snippet, the door is Claude Code. It reads and edits the files directly, instead of you copying and pasting fragments into a conversation.If you’ve ever caught yourself pasting the fifth file in a row into a chat, that was the sign you were at the wrong door.Model and effort here: leave them at the defaults. Only drop to a lighter model when the work is mechanical: renaming things, applying the same change in twenty places. And only raise effort when it cuts corners: skipped a file, didn’t run the tests, abandoned a change halfway.Claude Cowork: when the work is long but isn’t codeThis is the door more people should know about and don’t. Cowork came out of a simple observation at Anthropic: non-technical teams were reaching for Claude Code to do work that wasn’t programming. So they built the same capability with a simplified experience, aimed at knowledge work.The difference from chat is the shape of the request. In chat, you ask a question and get an answer. In Cowork, you hand over an outcome and get a finished deliverable. In practice:Filling out the expense report from a folder full of receipt photosTurning 40 scattered note files into one organized reportReorganizing the folder that became a junk drawer, with logic that actually holdsCross-referencing several spreadsheets and returning a formatted analysisModel and effort here: this is where the heavy models earn their keep. The task is long, there are a lot of decisions along the way, and you won’t be watching. Spend it.Claude Design: when the thing has to be seenIf what you want is visual — a deck, a one-pager, a landing page, a screen prototype, the door is Claude Design. It works on a canvas: it generates a first version, and you refine by talking to it, commenting on it, or editing directly, until you export.Asking for a good-looking slide inside a text chat is like asking for a floor plan over the phone. It works, and you’ll redo it.Model and effort here: they usually don’t show up as choices at all. The product ships with the right configuration. One less thing to think about.The short ruleChat to think and write. Code for code. Cowork for long work that isn’t code. Design for anything that has to be seen.What almost nobody does (and what fixes the most)When the output comes back bad, the reflex is always the same: switch to a better model. Most of the time that’s the last place the problem was.The order that works runs the other way:1. Context. The model doesn’t know what you didn’t tell it. Did you attach the document? Give an example of the format you want? Say who’s going to read it? Most bad output dies right here.2. Constraints. “Write a post about leadership” is a request with no right answer, so it hands you the average of the internet. “Write 300 words for first-time managers, don’t mention Simon Sinek, open on a concrete scene” is a completely different request.3. Model and effort. Only after that. And only if the first two are already handled.Switching models before fixing your context is buying a better car to drive down the same dirt road.In the end, it’s two dials and one questionThe model decides who does the work. Effort decides how much time that person gets.And the question decides both: does this task have one right answer, or several defensible ones?A senior with five minutes loses to an intern with two hours on mechanical work. But on an ambiguous problem, no amount of time fixes the intern.That’s it. The rest is a table.This story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!Which Claude Model Should You Use? The Guide I Wish I’d Read First was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Source: Generative AI Pub — Published — Category: Image AI

🔗 Read full article on Generative AI Pub →