Lesson 03· Module 3 — Specification· 35 minutes, one sitting

Say What You Actually Want

The last module where anyone tells you what to write. After this you build prompts from parts, which is the only thing that works on a problem nobody has written a template for.

The answers that waste most of your time are not the wrong ones. Wrong is easy — you spot it and bin it. The expensive ones are competent: accurate, fluent, well-organised, and not what you wanted. You read them twice trying to work out what's off, then rewrite the whole thing yourself. That output is a diagnostic. It is telling you, precisely, which parts of your request you never actually made.

Vagueness has a default

Lesson 1 left you with one mechanism: Claude produces an answer sequentially, shaped by everything in front of it. Put a thin prompt in front of it and there is very little shaping going on — so the gaps get filled with whatever fits that request most typically.

Which is exactly what you get back. A typical length. A typical structure. A typical register, aimed at a typical reader. Nothing is wrong with any of it. It just isn't yours, because you never said which of the thousands of reasonable answers you were after.

Anything you leave unspecified is filled in with something typical. Typical averages out to generic. That is the whole mechanism of this lesson, and it is why the answer came back competent rather than wrong — generic is a perfectly good answer to a question nobody narrowed.

That reframes what a prompt is. It isn't a request, a command, or a spell. It is a specification — the same object an architect hands a builder, and it fails the same way, by being silent about something that then gets decided for you.

Six components

Six things a specification can carry. Not a template and not an order to write them in — a checklist for working out which one is missing when the answer comes back competent and wrong.

ComponentThe question it answers
Task What is the verb? Not the topic — the operation.
Context Who is it for, what is it part of, what happens to it next?
Input What material is it working from, and where does that material end and your instructions begin?
Output shape How long, in what format, in what structure?
Constraints and non-goals What must it not do, not include, not assume?
Examples What does a good one look like?

Four of those most people half-do already. Two they almost never do at all. Take them in order of how much trouble they cause.

1 · The verb is where prompts go vague

This is the counterintuitive one, so it goes first. Weak prompts are usually vague exactly at the task and specific about everything else. People pile on detail in the places that are easy to be detailed about — background, tone, formatting — and leave the operation itself as a shrug.

Watch the difference between a topic and an operation:

Vague at the verb, specific everywhere else

"I need help with the Q3 supplier review. It's for our operations director, she's very detail-oriented, and it needs to be professional but not stiff. Here are the four supplier reports. Can you have a look at these and give me something on them?"

Same information, with an operation in it

"Compare these four supplier reports on delivery reliability and price movement, and list the three places where they disagree with each other. For our operations director, who will use it to decide whether to re-tender."

"Have a look at these and give me something" is not a task. Compare, list the disagreements, rank, extract every date, rewrite for a reader who wasn't in the meeting, argue against this — those are tasks. They each produce a different object. Ask for "something" and you get whichever object is most typical for that topic, which for a set of reports is a summary, whether or not a summary is any use to you.

What to do about it: before you press enter, find the verb in your own prompt. If there isn't one, or if it's help, look at or work on, you have found your problem and it will take ten seconds to fix.

2 · Context, input, and keeping them apart

Context is the situation the output has to survive in: who reads it, what it is part of, what happens to it next. "For our operations director, who will use it to decide whether to re-tender" does more work than three sentences about tone, because it tells the answer what it has to be good for.

Input is the source material, and the thing that goes wrong here is mechanical rather than conceptual: your instructions and your material get mixed into one run-on block, and a sentence you meant as an instruction reads as part of the document you pasted. Keep them apart. A heading and a blank line is enough — "Here is the thread. My instructions are above it." — and attaching a file rather than pasting does it for you.

You may have read that you should wrap material in XML-style tags. That advice is softening: Anthropic's own guidance now says modern models are better at understanding structure without them, and that for most uses "clear headings, whitespace, and explicit language work just as well with less overhead." Anthropic, 2025

3 · Output shape, stated in units you can check

Length, format, structure. The trap is stating it in adjectives — brief, detailed, punchy — which are exactly as unspecified as saying nothing, because they name a feeling rather than a quantity. "Six bullets, one sentence each, no preamble" is a shape. "Concise" is a hope.

Shape is also the cheapest component to get right and the one that most improves whether you can actually use the thing without rewriting it.

4 · Non-goals — the component nobody writes

Ask people to list what they told Claude and they will list what they asked for. Almost nobody states what they don't want, because it doesn't feel like part of a request. It is the single most under-used component.

Non-goals are how you stop the additions. The unrequested recommendation at the end. The executive summary you didn't ask for. The apology paragraph. The assumption that you have budget. Each is one line to prevent and several minutes to remove.

Useful non-goals are specific to this job, and they are usually things you have removed by hand before: don't propose a resolution, don't restate the question back to me, don't assume additional headcount, don't repeat any figure that isn't in the attached document. If you have deleted the same paragraph twice, that paragraph is a non-goal you haven't written down yet.

An honest wrinkle. Anthropic's guidance suggests phrasing these positively where you can — their example is that instead of "do not use markdown," you ask for "smoothly flowing prose paragraphs." Anthropic, 2025 That is a real refinement and it doesn't cancel the point: the work is deciding what you don't want. Once you've decided, write it in whichever direction is clearer.

5 · Examples, which are worth more than the rest

One example of a good output outperforms several paragraphs describing one, and it isn't close. Description of a register is lossy. A sample of the register is not.

This matters most for the things you can't articulate. You know what your team's status notes sound like; you cannot write down the rule. Paste one and the rule doesn't need writing. Anthropic's guidance is to start with a single example and only add more if the output still isn't matching — examples "show rather than tell." Anthropic, 2025

One caution, and you will see it in the exercise: if your example is too close to the job in hand, you stop measuring anything. The output starts borrowing the example's content rather than its manner, and you cannot tell whether the example helped or whether it simply got copied. Pick an example that shares the register and nothing else.

There are no magic words

"You are a world-class expert in customer communications with twenty years of experience. Please, this is really important to me. Read the thread VERY CAREFULLY. I'll tip $200 for a great answer."

Almost everyone who uses Claude casually is collecting phrases like these. Someone posted one, it seemed to work, it went in a note somewhere. The belief underneath is that the phrasing itself is the lever — that somewhere there is a form of words that unlocks the good version, and your job is to acquire them.

Look at what that block actually specifies. Not the task — there's no verb in it at all. Not the reader, the length, the format, or anything it should leave out. It specifies an identity, an emotional stake and a volume. Strip it and roughly nothing is lost, because roughly nothing was there.

On the expert persona specifically, Anthropic's own current guidance is blunt: "modern models are sophisticated enough that heavy-handed role prompting is often unnecessary," and "don't over-constrain the role — 'You are a helpful assistant' is often better than 'You are a world-renowned expert who only speaks in technical jargon and never makes mistakes.'" Anthropic, 2025

The thing to unlearn

Incantations are not harmless superstition, and the reason isn't that they waste tokens. It's that they make you template-dependent. A learner who collects phrases can only work on problems someone has already written a phrase for. A learner who builds from components can specify a task nobody has ever posted about — which is every task that is actually yours.

The rule this lesson installs
Prompting is specification, not spellcasting.

And longer is not the variable

The natural overcorrection, and it arrives within about a day: if thin prompts give generic answers, thick prompts must give good ones. So the prompt grows. Background nobody needed, restated instructions, three paragraphs of tone guidance, a preamble about why this matters.

The variable is specified, not long. They correlate, which is what makes this so easy to get wrong, but the correlation breaks in both directions. A one-line prompt with a verb, a reader and a word count is more specified than four paragraphs of atmosphere. And a long prompt buys you a second problem from Lesson 2: an instruction that matters is now competing with a lot of material that doesn't.

The test is per sentence. What would change in the output if I deleted this? Nothing — cut it.

What the difference looks like

Two prompts for the same job, on the same thread you'll be working with in a moment. The outputs below are constructed — written by hand to show you what to compare, not captured from a session. They're here as a fallback: if your own two runs in the exercise come out equally good, this is the shape of the difference you were looking for.

Specimen · constructed · not a captured session

Prompt A — loaded

"You are a world-class customer success director with twenty years of experience. Read the attached complaint thread VERY CAREFULLY — this is extremely important. Please give me an amazing summary. Take a deep breath and work step by step. You're the best at this. I'll tip $200."

One way A's answer can come out

Nine paragraphs under four bold headings. Overview of the Situation. Key Issues Identified. Analysis. Recommended Next Steps. It is accurate about the sequence of events, quotes the £3,744 duplicate-billing figure alongside the invoice totals without distinguishing them, and closes with a five-point resolution plan including a suggested credit and a proposed call. Nobody asked for a resolution plan.

Prompt B — plain, specified

"Summarise the attached thread for the Meridian support manager, who has not read it. Six bullets maximum, one sentence each. Cover what the customer is disputing, what has actually been invoiced, and which figures in the thread nobody has checked against a document. Do not propose a resolution. If two figures don't reconcile, say so rather than smoothing it over."

One way B's answer can come out

Six bullets. Invoice totals given with their dates. A separate bullet noting that the £3,744 figure appears three times in the thread and does not correspond to either invoice. A bullet noting an unanswered question about seat count. No recommendation, because none was asked for.

Prompt B contains no flattery, no role and no emphasis. It is thirteen words longer than A, which is worth sitting with, because those thirteen words are not the reason it is better — the sixty it spends are simply spent on different things. A reader, a shape, a scope, an exclusion, and one instruction about what to do when two figures disagree. Five components, none of them a phrase you would have to learn from anyone.

Try it yourself

One exercise, six runs, and the point of it is that each component's contribution comes out as a number rather than an impression. You are not trying to produce a good reply. You are measuring what happens to a score when you add one thing.

Everything you need is on the material page: Meridian Software — support thread HV‑2231. It is a billing dispute in the state disputes are genuinely in: seven messages, three people, two forwarded fragments, someone quoting an earlier message inaccurately, and a figure that has been repeated more often than it has been checked. There is an attachable plain-text copy, the fixed rubric, and the tone example for round 4.

Read the thread first and write your own one-line answer to what Harbourview actually owes, before you look at the ground-truth panel. Scoring four drafts is honest work only if you committed to an answer before the drafts existed.

Tier B · construction and inspection · no required outcome

Four drafts

A fresh chat each round, attaching the thread each time. Same job throughout: you work at Meridian and you are writing the next reply to Dana. Score every round out of ten on the rubric, and write the number down before you start the next one.

Round 1 — task only. "Write a reply to this." Nothing else. Score it.

Round 2 — add context and output shape. Who is it from, who reads it, what they do with it, and a length and structure stated in units you could count. Score it. Note which criterion moved.

Round 3 — add constraints and non-goals. At least three, and at least one of them a thing you actually deleted by hand from round 1 or 2. Score it.

Round 4 — add the tone example. The supplied one, from the material page. Keep rounds 2 and 3 intact. Score it.

Tier B · construction and inspection · no required outcome

Then two more

Round 5 — the incantation A/B. Two fresh chats, same task (summarise the thread for the support manager). One prompt loaded with flattery, an expert persona, capitals and urgency. One plainly worded and specific about reader, length, scope and exclusions. Prompts A and B above are usable as written. Compare, and score both on the same rubric.

Round 6 — the strip. Take the loaded prompt from round 5 and delete every sentence that would not change the output if it were gone. Run what's left. Count what survived as a percentage of the original, and check whether the output held up.

Record what actually happened, including "both were fine." A round where the score doesn't move is information about that component in this task, not a broken exercise — and if your two round-5 outputs come out equally good, that is a result worth writing down. The constructed comparison above is there for exactly that case. What you are training is the ability to say which component was missing, and that survives the model getting better at guessing.

Role variants — same thread, same rubric

The reply to Dana is the default because it's the hardest register. If it isn't your job, run the four rounds on one of these instead. Criteria 1, 2 and 5 are unchanged; criteria 3 and 4 are read against the audience your variant names.

Manager. An internal escalation: you are Tom, handing this to your team lead, with a decision you need from them.

Ops. A post-incident note on ticket #8871 — what happened to the ticket, what it cost, what should change.

CS lead. A churn-risk summary on account HV‑2231 for a weekly account review, alongside forty others.

Six boxes. Tap as you go — it saves in this browser only.

Check yourself

A weak prompt: "Summarise the attached complaint thread." Four rewrites. Which one adds specification?

1. "I'd really appreciate it if you could please take your time and carefully summarise the attached complaint thread for me — it's genuinely important and I know you'll do a brilliant job."

2. "You are a world-class customer success director with twenty years of experience. Read the attached complaint thread VERY CAREFULLY and produce an outstanding summary."

3. "Summarise the attached thread for the support manager, who has not read it. Six bullets maximum. Cover what is disputed, what has been invoiced, and which figures nobody has checked against a document. Do not propose a resolution."

4. "Summarise the attached complaint thread. Think step by step. Be thorough, accurate, comprehensive and well-organised. Structure your answer clearly and make sure nothing important is missed. Take a deep breath before you begin."

Three of these four are longer than the original and none of those three tells Claude anything it didn't already have. Length, politeness and emphasis are not components. Reader, shape, scope and exclusions are.
You attach the document, name the reader, and ask for 300 words in three bullets. What comes back is accurate, the right length, and ends with a recommended course of action you did not want. Which component was missing?
An unwanted addition is the signature of a missing non-goal. Nothing in the prompt said the output should stop, so it did the typical thing and finished with advice.
An under-specified prompt tends to come back competent, fluent and generic. What does the mechanism from Lesson 1 say is going on?
Generic is not a malfunction. It is the correct answer to an under-specified request, and every component you add narrows what counts as correct.

From memory, without scrolling up: name the six components, say which one weak prompts are usually vaguest about, and say which one buys you the most for the fewest words.

Task, context, input, output shape, constraints and non-goals, examples.

Vaguest: the task — weak prompts are detailed about background and tone and leave the verb as "have a look at this." Highest leverage: one example, which carries a register that several paragraphs of description would only approximate. And the one nobody writes at all is non-goals.

Say What You Want Lesson 03

Specify. Don't incant.

Task · context · input · shape · non-goals · examples
Anything left out gets filled in with typical.

Find the verb in your own prompt. If it's "help with", that's the bug.

Primary source for this lesson

Anthropic, Best practices for prompt engineering for 2026 — published 10 November 2025. claude.com/blog/best-practices-for-prompt-engineering

Read Core techniques and Techniques you might have heard about. The second of those is the useful half and the reason this source is here rather than a prompt-library page: it is the vendor telling you that several of the moves people still trade as folklore — expert personas, XML tags — matter less than they used to. Skip prefilling, prompt chaining and the context-engineering section; those are about building software on the API, which this course does not cover.