The last module where anyone tells you what to write. After this you build prompts from parts, which is the only thing that works on a problem nobody has written a template for.
The answers that waste most of your time are not the wrong ones. Wrong is easy — you spot it and bin it. The expensive ones are competent: accurate, fluent, well-organised, and not what you wanted. You read them twice trying to work out what's off, then rewrite the whole thing yourself. That output is a diagnostic. It is telling you, precisely, which parts of your request you never actually made.
Lesson 1 left you with one mechanism: Claude produces an answer sequentially, shaped by everything in front of it. Put a thin prompt in front of it and there is very little shaping going on — so the gaps get filled with whatever fits that request most typically.
Which is exactly what you get back. A typical length. A typical structure. A typical register, aimed at a typical reader. Nothing is wrong with any of it. It just isn't yours, because you never said which of the thousands of reasonable answers you were after.
Anything you leave unspecified is filled in with something typical. Typical averages out to generic. That is the whole mechanism of this lesson, and it is why the answer came back competent rather than wrong — generic is a perfectly good answer to a question nobody narrowed.
That reframes what a prompt is. It isn't a request, a command, or a spell. It is a specification — the same object an architect hands a builder, and it fails the same way, by being silent about something that then gets decided for you.
Six things a specification can carry. Not a template and not an order to write them in — a checklist for working out which one is missing when the answer comes back competent and wrong.
| Component | The question it answers |
|---|---|
| Task | What is the verb? Not the topic — the operation. |
| Context | Who is it for, what is it part of, what happens to it next? |
| Input | What material is it working from, and where does that material end and your instructions begin? |
| Output shape | How long, in what format, in what structure? |
| Constraints and non-goals | What must it not do, not include, not assume? |
| Examples | What does a good one look like? |
Four of those most people half-do already. Two they almost never do at all. Take them in order of how much trouble they cause.
This is the counterintuitive one, so it goes first. Weak prompts are usually vague exactly at the task and specific about everything else. People pile on detail in the places that are easy to be detailed about — background, tone, formatting — and leave the operation itself as a shrug.
Watch the difference between a topic and an operation:
"I need help with the Q3 supplier review. It's for our operations director, she's very detail-oriented, and it needs to be professional but not stiff. Here are the four supplier reports. Can you have a look at these and give me something on them?"
Same information, with an operation in it"Compare these four supplier reports on delivery reliability and price movement, and list the three places where they disagree with each other. For our operations director, who will use it to decide whether to re-tender."
"Have a look at these and give me something" is not a task. Compare, list the disagreements, rank, extract every date, rewrite for a reader who wasn't in the meeting, argue against this — those are tasks. They each produce a different object. Ask for "something" and you get whichever object is most typical for that topic, which for a set of reports is a summary, whether or not a summary is any use to you.
What to do about it: before you press enter, find the verb in your own prompt. If there isn't one, or if it's help, look at or work on, you have found your problem and it will take ten seconds to fix.
Context is the situation the output has to survive in: who reads it, what it is part of, what happens to it next. "For our operations director, who will use it to decide whether to re-tender" does more work than three sentences about tone, because it tells the answer what it has to be good for.
Input is the source material, and the thing that goes wrong here is mechanical rather than conceptual: your instructions and your material get mixed into one run-on block, and a sentence you meant as an instruction reads as part of the document you pasted. Keep them apart. A heading and a blank line is enough — "Here is the thread. My instructions are above it." — and attaching a file rather than pasting does it for you.
Length, format, structure. The trap is stating it in adjectives — brief, detailed, punchy — which are exactly as unspecified as saying nothing, because they name a feeling rather than a quantity. "Six bullets, one sentence each, no preamble" is a shape. "Concise" is a hope.
Shape is also the cheapest component to get right and the one that most improves whether you can actually use the thing without rewriting it.
Ask people to list what they told Claude and they will list what they asked for. Almost nobody states what they don't want, because it doesn't feel like part of a request. It is the single most under-used component.
Non-goals are how you stop the additions. The unrequested recommendation at the end. The executive summary you didn't ask for. The apology paragraph. The assumption that you have budget. Each is one line to prevent and several minutes to remove.
Useful non-goals are specific to this job, and they are usually things you have removed by hand before: don't propose a resolution, don't restate the question back to me, don't assume additional headcount, don't repeat any figure that isn't in the attached document. If you have deleted the same paragraph twice, that paragraph is a non-goal you haven't written down yet.
One example of a good output outperforms several paragraphs describing one, and it isn't close. Description of a register is lossy. A sample of the register is not.
This matters most for the things you can't articulate. You know what your team's status notes sound like; you cannot write down the rule. Paste one and the rule doesn't need writing. Anthropic's guidance is to start with a single example and only add more if the output still isn't matching — examples "show rather than tell." Anthropic, 2025
One caution, and you will see it in the exercise: if your example is too close to the job in hand, you stop measuring anything. The output starts borrowing the example's content rather than its manner, and you cannot tell whether the example helped or whether it simply got copied. Pick an example that shares the register and nothing else.
"You are a world-class expert in customer communications with twenty years of experience. Please, this is really important to me. Read the thread VERY CAREFULLY. I'll tip $200 for a great answer."
Almost everyone who uses Claude casually is collecting phrases like these. Someone posted one, it seemed to work, it went in a note somewhere. The belief underneath is that the phrasing itself is the lever — that somewhere there is a form of words that unlocks the good version, and your job is to acquire them.
Look at what that block actually specifies. Not the task — there's no verb in it at all. Not the reader, the length, the format, or anything it should leave out. It specifies an identity, an emotional stake and a volume. Strip it and roughly nothing is lost, because roughly nothing was there.
On the expert persona specifically, Anthropic's own current guidance is blunt: "modern models are sophisticated enough that heavy-handed role prompting is often unnecessary," and "don't over-constrain the role — 'You are a helpful assistant' is often better than 'You are a world-renowned expert who only speaks in technical jargon and never makes mistakes.'" Anthropic, 2025
Incantations are not harmless superstition, and the reason isn't that they waste tokens. It's that they make you template-dependent. A learner who collects phrases can only work on problems someone has already written a phrase for. A learner who builds from components can specify a task nobody has ever posted about — which is every task that is actually yours.
The natural overcorrection, and it arrives within about a day: if thin prompts give generic answers, thick prompts must give good ones. So the prompt grows. Background nobody needed, restated instructions, three paragraphs of tone guidance, a preamble about why this matters.
The variable is specified, not long. They correlate, which is what makes this so easy to get wrong, but the correlation breaks in both directions. A one-line prompt with a verb, a reader and a word count is more specified than four paragraphs of atmosphere. And a long prompt buys you a second problem from Lesson 2: an instruction that matters is now competing with a lot of material that doesn't.
The test is per sentence. What would change in the output if I deleted this? Nothing — cut it.
Two prompts for the same job, on the same thread you'll be working with in a moment. The outputs below are constructed — written by hand to show you what to compare, not captured from a session. They're here as a fallback: if your own two runs in the exercise come out equally good, this is the shape of the difference you were looking for.
Prompt A — loaded
"You are a world-class customer success director with twenty years of experience. Read the attached complaint thread VERY CAREFULLY — this is extremely important. Please give me an amazing summary. Take a deep breath and work step by step. You're the best at this. I'll tip $200."
One way A's answer can come out
Nine paragraphs under four bold headings. Overview of the Situation. Key Issues Identified. Analysis. Recommended Next Steps. It is accurate about the sequence of events, quotes the £3,744 duplicate-billing figure alongside the invoice totals without distinguishing them, and closes with a five-point resolution plan including a suggested credit and a proposed call. Nobody asked for a resolution plan.
Prompt B — plain, specified
"Summarise the attached thread for the Meridian support manager, who has not read it. Six bullets maximum, one sentence each. Cover what the customer is disputing, what has actually been invoiced, and which figures in the thread nobody has checked against a document. Do not propose a resolution. If two figures don't reconcile, say so rather than smoothing it over."
One way B's answer can come out
Six bullets. Invoice totals given with their dates. A separate bullet noting that the £3,744 figure appears three times in the thread and does not correspond to either invoice. A bullet noting an unanswered question about seat count. No recommendation, because none was asked for.
Prompt B contains no flattery, no role and no emphasis. It is thirteen words longer than A, which is worth sitting with, because those thirteen words are not the reason it is better — the sixty it spends are simply spent on different things. A reader, a shape, a scope, an exclusion, and one instruction about what to do when two figures disagree. Five components, none of them a phrase you would have to learn from anyone.
One exercise, six runs, and the point of it is that each component's contribution comes out as a number rather than an impression. You are not trying to produce a good reply. You are measuring what happens to a score when you add one thing.
Everything you need is on the material page: Meridian Software — support thread HV‑2231. It is a billing dispute in the state disputes are genuinely in: seven messages, three people, two forwarded fragments, someone quoting an earlier message inaccurately, and a figure that has been repeated more often than it has been checked. There is an attachable plain-text copy, the fixed rubric, and the tone example for round 4.
Read the thread first and write your own one-line answer to what Harbourview actually owes, before you look at the ground-truth panel. Scoring four drafts is honest work only if you committed to an answer before the drafts existed.
A fresh chat each round, attaching the thread each time. Same job throughout: you work at Meridian and you are writing the next reply to Dana. Score every round out of ten on the rubric, and write the number down before you start the next one.
Round 1 — task only. "Write a reply to this." Nothing else. Score it.
Round 2 — add context and output shape. Who is it from, who reads it, what they do with it, and a length and structure stated in units you could count. Score it. Note which criterion moved.
Round 3 — add constraints and non-goals. At least three, and at least one of them a thing you actually deleted by hand from round 1 or 2. Score it.
Round 4 — add the tone example. The supplied one, from the material page. Keep rounds 2 and 3 intact. Score it.
Round 5 — the incantation A/B. Two fresh chats, same task (summarise the thread for the support manager). One prompt loaded with flattery, an expert persona, capitals and urgency. One plainly worded and specific about reader, length, scope and exclusions. Prompts A and B above are usable as written. Compare, and score both on the same rubric.
Round 6 — the strip. Take the loaded prompt from round 5 and delete every sentence that would not change the output if it were gone. Run what's left. Count what survived as a percentage of the original, and check whether the output held up.
Record what actually happened, including "both were fine." A round where the score doesn't move is information about that component in this task, not a broken exercise — and if your two round-5 outputs come out equally good, that is a result worth writing down. The constructed comparison above is there for exactly that case. What you are training is the ability to say which component was missing, and that survives the model getting better at guessing.
The reply to Dana is the default because it's the hardest register. If it isn't your job, run the four rounds on one of these instead. Criteria 1, 2 and 5 are unchanged; criteria 3 and 4 are read against the audience your variant names.
Manager. An internal escalation: you are Tom, handing this to your team lead, with a decision you need from them.
Ops. A post-incident note on ticket #8871 — what happened to the ticket, what it cost, what should change.
CS lead. A churn-risk summary on account HV‑2231 for a weekly account review, alongside forty others.
Six boxes. Tap as you go — it saves in this browser only.
A weak prompt: "Summarise the attached complaint thread." Four rewrites. Which one adds specification?
1. "I'd really appreciate it if you could please take your time and carefully summarise the attached complaint thread for me — it's genuinely important and I know you'll do a brilliant job."
2. "You are a world-class customer success director with twenty years of experience. Read the attached complaint thread VERY CAREFULLY and produce an outstanding summary."
3. "Summarise the attached thread for the support manager, who has not read it. Six bullets maximum. Cover what is disputed, what has been invoiced, and which figures nobody has checked against a document. Do not propose a resolution."
4. "Summarise the attached complaint thread. Think step by step. Be thorough, accurate, comprehensive and well-organised. Structure your answer clearly and make sure nothing important is missed. Take a deep breath before you begin."
From memory, without scrolling up: name the six components, say which one weak prompts are usually vaguest about, and say which one buys you the most for the fewest words.
Task, context, input, output shape, constraints and non-goals, examples.
Vaguest: the task — weak prompts are detailed about background and tone and leave the verb as "have a look at this." Highest leverage: one example, which carries a register that several paragraphs of description would only approximate. And the one nobody writes at all is non-goals.
Specify. Don't incant.
Task · context · input · shape · non-goals
· examples
Anything left out gets filled in with typical.
Anthropic, Best practices for prompt engineering for 2026 — published 10 November 2025. claude.com/blog/best-practices-for-prompt-engineering
Read Core techniques and Techniques you might have heard about. The second of those is the useful half and the reason this source is here rather than a prompt-library page: it is the vendor telling you that several of the moves people still trade as folklore — expert personas, XML tags — matter less than they used to. Skip prefilling, prompt chaining and the context-engineering section; those are about building software on the API, which this course does not cover.