The keystone lesson. Three lessons of mechanism now become six named faults and six matching repairs — plus the one honest answer that isn't a repair at all.
You read the reply, something in you goes no, and the whole thing gets thrown out. Then you do one of three things: rewrite the prompt at random, press regenerate, or decide this isn't something Claude can do. All three are guesses, and two of them cost you a second bad answer to find out. This lesson replaces the guess with a diagnosis — six named failure classes, each with a repair that does nothing for the other five.
Everything in Lessons 1 to 3 was scaffolding for this. Each mechanism you learned turns into a diagnostic move here, and the table below has a column pointing back at where each one came from.
It's a description of your reaction. The reaction is the same whatever went wrong, which is exactly why it's useless for deciding what to do next. Here are two rejections of the same task that produce the identical sentence in your head and take opposite repairs.
A. "Write something for the team about the new returns policy."
→ Four hundred words. Warm, literate, correctly structured. Opens on our continued commitment to customer satisfaction, notes that the policy has been updated to better serve our customers, and closes by thanking everyone for their flexibility during the transition. Nowhere does it say what changed.
B. "Write a note to the 40 store managers explaining the three changes to the returns policy (attached), what they need to do before the 1st, and who to ask if a customer disputes one. Keep it practical."
→ Eleven hundred words. All three changes correct. The pre-1st actions correct and complete. The escalation contact correct. Delivered as a document with five headings, a summary table and a closing paragraph on change management — when what actually goes out is six lines in tomorrow's morning huddle post.
Same verdict both times. But A failed because nothing in the request narrowed what a good answer would say, so it produced the most typical thing an all-staff policy note can be. And B failed with the content entirely right — only the container is wrong.
Now look at what the wrong repair would do. Format instructions added to A specify the container for a note whose content was never narrowed. Context and constraints added to B pile requirements onto content that was already right. Neither edit touches the actual fault, and each costs you a round trip to discover that. Both are what people reach for, because "it gave me garbage" doesn't tell them apart.
These six cover nearly everything you will meet. Read the middle column first — that's the tell, the thing you can actually observe in the output in front of you.
| Failure class | Looks like | The repair | From |
|---|---|---|---|
| Missing information | Confidently answers a different situation than yours | Supply what it couldn't have had | L2 |
| Under-specified | Valid, competent, and not what you wanted | Add constraints, or an example | L3 |
| Fabrication | Specifics that are wrong — numbers, names, citations | Ground it, or scope it to a source | L1 |
| Wrong shape | Right content, wrong length, format or register | Specify the format, or show one | L3 |
| Lost the thread | Contradicts itself, drops instructions you gave earlier | Fresh chat, restated brief | L2 |
| Insufficiently decomposed | Does all of it, shallowly, at one pass | Break into steps; give it room to work | L1 |
There is a printable version of this table: the failure taxonomy card. It and the capstone's process are the two things from this course worth keeping after you've forgotten which lesson they came from.
Under-specified vs. wrong shape. This is the pair from the specimen above, and it's the one that genuinely trips people. The question that separates them: if you fixed nothing but the presentation, would you be happy? If yes, it's shape — the thinking is right and it's in the wrong container. If you'd still be looking at something bland or beside the point, it's under-specified, and reformatting it will not help.
Missing information vs. fabrication. Both arrive as confident wrongness in the same even register, which is why Lesson 1 spent so long on it. The tell is scale. Fabrication is a wrong specific sitting inside an otherwise sound answer — a figure, a date, a citation, a name. Missing information is wrong at the level of the whole situation: it has answered a sensible question competently, and it isn't your question, because it didn't have the thing that would have made yours different.
The most common thing people do with a bad answer is press regenerate. It's one click, it costs nothing, and it feels like a retry.
It isn't one. Output is sampled — the same prompt twice can give two different answers — so regenerating draws again from the same place. What varies is wording, ordering, which example it reached for. What doesn't vary is anything caused by the prompt, because the prompt didn't change.
Every one of the six classes above is a property of the request, not of the draw. So what a reroll can change is the wording; what it cannot change is a fault that lives in the request — and you spend a round trip finding out which you had. That's the reroll trap, and it's the single most common wasted action among people who use Claude casually.
The exception, stated precisely so you don't over-learn this: rerolling is a legitimate move when you have already decided the specification is right and you want a second draw for variety — a different opening line, another angle on a headline. That's sampling used deliberately. It is not a repair, and it is never the first move.
Here's a move almost nobody makes, and it's cheap: paste your own brief back alongside the output and ask what in it led there.
"Here is the brief I gave you, and the reply you produced. What in the brief left this under-determined? List what you had to assume, and what you'd have needed from me to not have to assume it."
What you are asking for is a list of gaps in your own specification — and that is a question you are badly placed to answer for yourself, because you cannot read your own brief without also reading everything you meant but never typed.
Now the part that matters more than the technique.
It is another generated answer — produced by the same machinery, to a question about itself, with the earlier exchange sitting in front of it. Claude does not have privileged, reliable introspective access to why it generated something, and its account of its own reasoning is not a recording of what happened. Ask for a diagnosis, receive a hypothesis, then test it.
Testing it is easy and takes seconds, which is why the epistemics don't cost you anything. Each item on that list is a claim about your brief — and your brief is right there. Read it. If it says you never specified the audience, check: did you? If you did, drop that item. What survives the check is a real gap you can now fill. What you must not do is take the list as an account of the machinery and start optimising against it.
Seventh: it's a settings problem, not a prompt problem. The model selected, how much effort it's allowed, whether thinking is on — these are three separate controls behind the model name, and most people never touch any of them. Some failures really do move when you change one, particularly the long, multi-step, genuinely hard ones. The dials are Lesson 7, and holding them back until then is deliberate.
Reaching for a bigger dial before you've checked your own specification is the wrong reflex, and it's a hard one to unlearn once it's installed. A more expensive setting applied to an under-determined brief produces a more elaborate version of the wrong thing — and now you're paying more and waiting longer for it. Check the specification first. Every time. The setting is what you try when the specification is already sound.
And one more that isn't your prompt either: the source you supplied is wrong. You attached the price list, the output quotes a price, the price is wrong — and when you open the attachment, the attachment says that price. There is no prompt that repairs that. It looks exactly like fabrication from the outside, and the two-second check that separates them is asking Claude which line it used. That check is the whole of Lesson 5.
"It can't do this."
This is the conclusion people reach after two bad answers, and it is nearly always wrong. It feels like a finding about the tool. It's usually a finding about a specification that never narrowed the target, or a fact the model was never given. The sentence is doing something else too, which is worth admitting: it ends the task and hands the fault to someone else. That's why it's attractive at exactly the moment you're frustrated.
The clinic below is the argument, not this paragraph. Five of its six cases are fixable — several with a single line — and you do the fixing yourself, which is the only form of this argument that survives contact with your next bad Monday.
One of the six is genuinely unfixable, and you aren't told which. It's there on purpose and it is not a trick. The fact that case needs is private, held by someone who hasn't shared it, and present in nothing you can attach. No wording reaches it. A diagnostic that can only ever say try harder is a motivational poster; the point of naming faults is calibrated persistence, not unlimited persistence.
Which gives you the question that separates the two, and it's the last thing this lesson asks you to remember:
If yes — a file on your desktop, a page on the web, a number in a system you have access to, a decision you've made but never typed — then this is a repair, and the taxonomy tells you which one.
If no — a colleague's unshared figure, a decision nobody has taken yet, a document that exists only in someone's head — then no prompt reaches it, and the correct move is to close the chat and go and get it. Recognising this case quickly is a skill, not a defeat.
The clinic is six supplied prompt-and-output pairs, one per failure class, shuffled so the class isn't guessable from where it sits in the list. Each one includes the material it was working from, so you can run your own repair against the same inputs.
Open the six clinic cases. Each has a hidden model repair. Don't reveal it until you've done both halves of the case yourself.
1 · Classify and prescribe (this half cannot fail — the outputs are already on the page). For each case, write two lines: which class, and what the repair is. Then write the repaired prompt out in full.
2 · Apply it live (this half is the one that matters). Run your repaired prompt against the material supplied with that case. Then reveal the model repair and compare — not for a match, but for what each of you specified that the other didn't. Record what actually happened, including "my first version was already fine". Classifying without applying produces people who pass the quiz below and still can't fix their own work.
3 · Your own worst one. Find the most frustrating Claude exchange you've had in the last month — the real one, the one you gave up on. Run the taxonomy on it: name the class, apply the repair, rerun. Then answer the question from the box above out loud. If the answer is that the thing it needed doesn't exist anywhere you can reach, write that down as the result. That is a completed diagnosis, not a failed exercise.
Six cases and your own one. Tap as you go — it saves in this browser only.
From memory, without scrolling up: name the six failure classes and the repair each one takes. Then name the question you ask when two honest repairs have both failed.
Missing information → supply what it couldn't have had. Under-specified → add constraints, or an example. Fabrication → ground it, or scope it to a source. Wrong shape → specify the format, or show one. Lost the thread → fresh chat, restated brief. Insufficiently decomposed → break it into steps and give it room to work.
The question: does the thing it needs exist somewhere I can reach? If yes, keep repairing. If no, stop — and go and get it. Bonus marks if you also wrote down that regenerating an unchanged prompt draws again from the same place, and that asking Claude to diagnose itself yields a hypothesis rather than a log.
Name the class before you touch the prompt.
Missing information · Under-specified · Fabrication
Wrong shape · Lost the thread · Insufficiently decomposed
Anthropic, Signs of introspection in large language models (29 October 2025). anthropic.com/research/introspection
Read it for the second habit on this page and nothing else. It is the rare piece of research writing that is more careful about its own limits than its headline — it finds some genuine introspective capability and then says plainly that the ability is unreliable and limited in scope, that models often fail to demonstrate it at all, and that self-reports need validating rather than believing. Skip the interpretability method sections unless you want them; the caveats near the end are the part that changes what you do on Monday. If you want the sharper companion piece, Anthropic's Reasoning models don't always say what they think (April 2025) makes the same point about reasoning that is shown to you rather than reported after the fact.