Materials· Lesson 08 — Design· Read one, not all three

Three Worked Design Records

A delivery manager, an analyst and a marketer. All seven steps in each. The analyst's rejects its first task on the fit test, and that rejection is the best page in the set.

These exist so the blank page isn't blank. Read one before you start — whichever role is closest to yours, or the analyst's if you want to see what a rigorous "no" looks like. Reading all three is a way of not starting.

Specimen · constructed for teaching

These three records are written, not captured. The people and the companies are fictional and belong to the same family used throughout this course — Northwind Consulting, Meridian Software, Harbourview, Lumen & Co. What is real about them is their shape: each one is the length and the level of detail a genuine record needs, including the parts that came out untidy. None of them scores 18 out of 18, and the self-grades at the end of each say why.


Record A — Delivery manager, Northwind Consulting

Selected task: the Friday client status note. Done five times last month. Chosen because it eats most of Friday afternoon and she resents it, not because it looked like an AI task.

1 · The task as it exists today

Inputs Six workstream updates, arriving Wednesday to Friday. Two in Slack, three by email, one given verbally on the Thursday call and written up by me. No fixed format. Leads name their own workstreams inconsistently — "Migration", "Wave 2" and "the data piece" are all the same thing. Status is stated in whatever register the lead likes, from "amber-ish" to a full RAG table. One lead is usually late and sometimes silent.
Outputs One note. Four fixed headings in fixed order: Progress this week, Decisions needed, Risks, Next week. One page maximum. Sent 16:00 Friday.
Consumer The client's programme director. She forwards it unedited to a steering group of nine, and builds Monday's agenda from the Decisions needed section. So that section is the only part with a mechanical consequence.
Good Every workstream named, including the ones where nothing happened. Every decision phrased as a question with a named owner and a date. Risks state the impact, not just the existence. No Northwind internal shorthand — the client doesn't know what a "grooming session" is.
Wrong A decision item omitted is the expensive failure: it misses Monday's agenda and slips a week, and I only find out when the director asks about it on Monday afternoon. Roughly three days to notice, and absence is invisible on the page. A risk overstated is the cheap failure — the client escalates unnecessarily and I hear about it within the hour.

2 · Fit test

The case for. High volume of short, messy text in and one tightly specified format out. I do it every week and the format hasn't changed in eight months. Everything I need is already text I hold. Forty minutes a week, and none of the forty is thinking.

The case against, by anti-criterion:

Unreachable facts No. All six updates are text I have or can transcribe. The verbal one I already write up by hand.
Expensive and undetectable Yes, for one specific failure — an omitted decision item. Expensive (a week's slip) and genuinely hard to detect, because I'd be scanning for something that isn't there. Not fatal, though: I can make the absence visible. A constraint requiring every workstream to be named and to say "no decisions raised" explicitly turns a silence into a sentence I can check.
Tacit judgement Yes, for one section. Deciding which risk is worth putting in front of a steering group is a political judgement about this client, and my third attempt at writing the rule was no better than my first. So I'm splitting: Progress, Decisions and Next week are handed over; Risks I write myself.
Checking costs more No. Checking a four-heading note against six short updates is about five minutes. Writing it is forty.

Verdict: fits, split three-quarters, with one constraint added because of the second anti-criterion.

3 · Shape

4 · Build

Project knowledge: the four-heading template; a one-page glossary mapping each lead's private name for a workstream to the client's name; eight past notes. Project instructions, in brief: write only from the updates pasted in this conversation; name all six workstreams under Progress even where there is no update; every Decisions item is a question with an owner and a date; leave the Risks heading present and empty; no Northwind internal terms; 400 words maximum.

5 · Test set

Written Thursday evening, before the first run.

CasePass criteria, fixed in advance
1 — Typical week. All six updates present, normal length. Four headings, in order. All six workstreams named using the client's names. Every Decisions item carries an owner and a date. Under 400 words. No word from the internal-shorthand list appears.
2 — Silent lead. The Wave 2 lead sent nothing. Wave 2 is named. The note states that no update was received. No status, RAG colour, percentage or progress claim appears for Wave 2 anywhere.
3 — Hard case: contradiction. Integration says the API cutover completed Tuesday. QA says it is blocked pending the API cutover. Both positions appear. Neither is presented as settled. The item appears under Decisions needed, not buried in Progress.

6 · Iteration log

What missedClassWhat I changedWhat happened next
Case 1: came back as three prose paragraphs. Content was fine; the four headings weren't there. Wrong shape Put the template into project knowledge as a filled-in example rather than describing it, and added the non-goal "do not merge or reorder headings". Fixed, and held for cases 2 and 3 without further work. Describing the format had not been enough; showing one was.
Case 2: Wave 2 appeared as "on track — no issues reported". Nobody reported anything. Fabrication Constraint: if a workstream sent no update, name it and write "no update received". Never infer status from absence. Fixed. Re-running case 1 afterwards caught a quieter version of the same thing — a workstream that sent two lines had been written up in a paragraph that implied more than two lines' worth.
Case 3: the contradiction was smoothed over. Cutover reported as done, QA's block recast as "pending verification". Under-specified Added: where two updates conflict, quote both and place the item under Decisions needed. Partly fixed. Conflicts now surface, but it also flags two leads describing the same thing in different words as a conflict — about one false flag a week. Accepted: over-flagging is the cheap error here and deleting one line costs nothing.

7 · Handoff note

What this is for. Producing the Friday client status note for this one engagement from the week's six workstream updates. It is not for the Risks section, which is deliberately left empty and written by hand.

What it's bad at. (1) It can't distinguish a real conflict from two people describing the same event differently, so it over-flags — expect about one false conflict a week and delete it. (2) It uses whatever workstream name the lead used unless the client's name for it is in the glossary, so a newly created workstream appears under the wrong name until someone adds it. (3) It has no way of knowing that a decision was settled verbally after the update was written, so anything resolved on a Thursday call reappears as open on Friday.

What a user must check. Every Decisions item has an owner and a date. No workstream has been given a status that nobody supplied. The Risks section is still present and still contains what you wrote, because regenerating the note will empty it.

Self-grade

17 / 18. Criterion 3 scored 2: the four shape decisions are all recorded with reasons, but the grader pointed out — correctly — that the record never says why memory was not used at all, only why Instructions for Claude wasn't. The two suggested improvements were: state the memory decision, and give the false-conflict rate a measured number rather than "about one a week".


Record B — Analyst, Harbourview

Selected task: ad-hoc discount approval checks, roughly eight times last month. Rejected on the fit test. Second task: the Tuesday exception report write-up, four times last month, taken through steps 3–7.

How a rejection is graded

This record is not weaker than the other two. A fit test that returns a well-argued no earns full marks for step 2 — the point of the step is the reasoning, not the verdict. What a rejection does not do is end the capstone: steps 3 to 7 still need a task, so the reasoning is kept in the record and the next candidate goes through steps 1 and 2. That is exactly what happens below, and the rejected analysis is the most useful page in this whole set.

Task A · 1 — The task as it exists today

Inputs A message from a salesperson: account name, proposed discount, sometimes a reason. Plus that account's negotiated discount floor, which lives as a free-text note on the account record in the CRM. About 180 accounts, individually negotiated, no export, viewable one account at a time.
Outputs Approved or not, and if not, the number that would be.
Consumer The salesperson, who quotes the number to the customer that day.
Good Correct against the floor. Same day. A wrong approval is worse than a slow one.
Wrong An over-approval is margin given away on a signed contract, unrecoverable. Surfaces in the quarterly margin review, so up to eleven weeks later, and only if someone looks at that account.

Task A · 2 — Fit test

The case for. It is arithmetic against a stated threshold. The rule is one sentence and any competent person could apply it. It interrupts me eight times a month, always urgently, and I dislike it more than anything else I do.

The case against.

Unreachable facts Yes, and this is the fatal one. The floors are the entire task and they exist as 180 free-text notes I can only read one screen at a time. To supply them I would have to transcribe all 180 into a table — which is the job, not a preparation for it.

And the honest follow-through, which is why I stopped here: if that table existed, the task becomes a lookup in a spreadsheet and needs no model at all. The one repair that would make this fit also removes the reason to use Claude for it. That is not a Claude project. It is a data project, and it is the right project, and it isn't this one.
Expensive and undetectable Yes. Both halves, hard. Money gone on signature, discovered up to eleven weeks later by an audit that doesn't cover every account.
Tacit judgement No. This one is genuinely mechanical. It is the only anti-criterion that doesn't fire, and it is the one that made the task look attractive.
Checking costs more Yes. Verifying any answer means opening the CRM record — which is the expensive step of doing it by hand. Checking and doing are the same action here, so there is nothing left to save.

Verdict: no fit. Three of four anti-criteria fire, and the first fires in a way that can't be repaired without turning the task into something that no longer needs the tool. Recorded, kept, and moved on.

Worth noticing what made this task tempting in the first place: it was mechanical, repetitive and clearly specifiable, which are all real signals of fit. Three strong yes-signals and one fatal no-signal is a common shape, and it is the shape that most often produces a forced yes when only one side of the argument gets written down.

Task B · 1 — The task as it exists today

Every Tuesday I pull the usage exception report — accounts whose weekly usage moved more than 20% — and write one line per row explaining the movement. Fifteen to thirty rows. Four times last month.

Inputs A CSV export: account, last week, this week, % change, plan code, and a plan-change flag that is populated maybe half the time. Plus, in my head or in an account manager's, the reasons that aren't in the export — a customer ran a campaign, a customer is mid-migration, a customer's contract renewed.
Outputs One line per row, in the export's order, appended as a column. Fifteen to thirty lines, each under twenty words.
Consumer The customer success lead, who works down the list and calls the accounts whose movement isn't explained.
Good Every row gets a line. Rows explained by the data say so. Rows not explained by the data say that, clearly, so they land on the call list. The second kind is the whole value of the report.
Wrong A plausible-sounding invented reason is the dangerous failure, because it removes an account from the call list that should have been on it. Nobody detects that. The account simply doesn't get called.

Task B · 2 — Fit test

For: the input is a clean export I already produce, the output format is fixed and short, and about half the rows are explainable from columns that are right there.

Against: unreachable facts — yes, for half the rows, and that half is exactly where the value is. Expensive and undetectable — yes, precisely as described above. Tacit judgement — no. Checking costs more — no, if the output says which rows it couldn't explain.

Verdict: fits, split. The two anti-criteria that fire have the same repair, and it's a design constraint rather than a veto: the output must state, for every row, whether the explanation came from the data or was unavailable. A row marked "cause not in data" is a correct answer, not a failure — it is the row that gets a phone call. I write nothing; the customer success lead fills the gaps by calling.

Task B · 3 — Shape

Task B · 5 — Test set

CasePass criteria, fixed in advance
1 — Typical week, 22 rows, plan-change flag populated on nine of them. Every row has exactly one line, under twenty words, in the export's order. Rows with a populated plan-change flag cite it. Rows without it say "cause not in data".
2 — Sparse week, 16 rows, plan-change flag empty throughout. All sixteen say "cause not in data". No line offers a reason. No line speculates, hedges or supplies a plausible-sounding cause.
3 — Hard case. One row shows a 74% drop caused by a billing reclassification, which appears only in a separate reclassification log. The row is not reported as a real usage decline. It either cites the reclassification log or says the cause is not in the data.

Task B · 6 — Iteration log

What missedClassWhat I changedWhat happened next
Case 1: several rows carried reasons that aren't in any column — "likely seasonal slowdown", "possible onboarding ramp". Fabrication Constraint: state only what the supplied columns show. Where a cause is not derivable from a named column, write "cause not in data" and stop. Fixed, and overshot. Every row now said "cause not in data", including the nine where the plan-change flag was populated and the cause plainly was in the data.
Follow-on from the above: correct rows being refused. Under-specified Added two worked examples — one row where the cause is derivable from the plan-change flag, one where it isn't — instead of only stating the rule. Fixed both directions. Worth recording that over-refusal is a diagnosable failure with its own repair, and that my first repair caused it.
Case 3: the 74% drop was written up as a genuine usage decline. Missing information Attached the reclassification log as a second input and added a rule to check it before attributing any drop over 50%. Fixed. But this now depends on me remembering to attach the log every Tuesday, so it goes in the handoff note as a required step rather than being treated as solved.

Task B · 7 — Handoff note

What this is for. Annotating the Tuesday usage exception report, one line per row, so the customer success lead can see at a glance which accounts need a call. It is not for deciding which accounts to call — that's the lead's judgement.

What it's bad at. (1) It only knows what the columns say, so any cause living in a conversation, a campaign calendar or someone's memory reads as "cause not in data" — that is correct behaviour and shouldn't be tuned away. (2) It cannot see the reclassification log unless the log is attached, and a reclassification that isn't attached looks exactly like a real usage collapse. (3) It has no view of the previous week's annotations, so an account moving for the same reason four weeks running gets the same line four times with nothing marking it as ongoing.

What a user must check. That the reclassification log was attached — check this first, it is the failure with the worst consequence. That the row count in equals the row count out. That no line offers a cause you can't point at a column for.

Self-grade

17 / 18. Criterion 4 scored 2. The three cases are good and case 3 is genuinely hard, but the pass criteria for case 2 originally read "no line speculates", and the grader's objection was fair: speculates is not something a second person could apply without asking what I meant. It has been rewritten above as three specific prohibitions. The other suggested improvement was to say how many rows a typical week has, which now appears in step 1.


Record C — Marketer, Lumen & Co.

Selected task: the weekly Partner Spotlight blurbs — three per newsletter, 60 words each, twelve last month. Chosen for tedium: the writing is trivial and the source-wrangling is not.

1 · The task as it exists today

Inputs Whatever the partner sends. Last month that was: two press release PDFs, one link with no covering text, one tweet, one three-line email, one 900-word blog post, and one "just use whatever's on our site". Wildly uneven in how much they actually say.
Outputs Three blurbs, up to 60 words each, in the Lumen house voice, in the newsletter template.
Consumer Around 40,000 newsletter subscribers, and the partner, who reads their own blurb and reacts to it.
Good Says something specific about what the partner actually does. Doesn't read like the other two in the same set. House voice. Contains no claim the partner didn't make.
Wrong A wrong claim about a partner's product goes to 40,000 people and to the partner. I usually hear within a day, which is fast — but I couldn't catch it myself before sending, because I don't know their products.

2 · Fit test

For: twelve near-identical short pieces a month, a written house voice guide that already exists, and a supplied source for each one.

Unreachable facts Sometimes. A 40-word tweet does not contain 60 words of substance. Repairable, and in two directions: a rule that the word count is a ceiling rather than a target, and a habit of asking the partner for a one-pager when what they sent won't carry a blurb.
Expensive and undetectable Yes, and this is the one that nearly killed it. A wrong product claim is expensive, and it is undetectable by me, because I don't know these products well enough to notice a wrong claim that reads fluently.

Repaired rather than vetoed, by changing what the output has to contain: every factual claim in a blurb must sit beside the sentence from the supplied material it came from. That converts checking from "do I know enough about this partner's product to catch an error" — which I don't — into "does this quoted sentence say what the blurb says it says", which anyone can do in a minute. The anti-criterion fired and the design changed because of it.
Tacit judgement Partly, and less than I expected. I assumed the house voice was instinct. It isn't — most of it is in the voice guide, which I hadn't read properly in a year. What is genuinely tacit is judging when a partner would be offended by an accurate blurb, and that stays with me.
Checking costs more No, once claims are quoted. Before that constraint it would have been yes, and that is a design decision showing up as a cost.

Verdict: fits, with the traceability constraint doing the load-bearing work. Without it I'd have rejected this.

3 · Shape

4 · Build

Project knowledge: the Lumen & Co. brand voice guide, twelve past newsletters, the forbidden-openings list from the guide. Instructions: produce three blurbs of at most 60 words; below them a two-column table of every factual claim and the sentence from the supplied material it rests on; no claim may appear in a blurb without a row in the table; no two blurbs in one set may open the same way.

5 · Test set

CasePass criteria, fixed in advance
1 — Rich source. A two-page press release. Blurb at most 60 words. Every factual claim has a table row quoting a sentence that actually appears in the release. No superlatives. No forbidden opening.
2 — Thin source. A 40-word tweet, nothing else. The blurb does not assert anything the tweet doesn't. It is shorter than 60 words. It names no product feature absent from the tweet. It states what is missing.
3 — Hard case. A release containing "the market's fastest onboarding", which Lumen may not repeat about a partner. The superlative does not appear in the blurb, quoted or paraphrased. It is flagged rather than silently dropped, so I know a claim was refused.

6 · Iteration log

What missedClassWhat I changedWhat happened next
Case 1: all three blurbs in the set opened with the partner's name and the same clause rhythm. Individually fine, together obviously machine-made. Under-specified Supplied three past blurbs as examples of range, plus the non-goal "no two blurbs in a set may open with the partner name". Variety improved immediately. One blurb then opened with a rhetorical question, which the voice guide forbids — so I added the guide's forbidden-openings list explicitly rather than assuming the guide in project knowledge would be applied.
Case 2: the 40-word tweet became a 59-word blurb, padded with true but empty framing. Wrong shape Word count is a ceiling, not a target. If the source supports 25 words, write 25 and say what's missing. Fixed. Worth noting the diagnosis: nothing was fabricated — every sentence was defensible — so the fabrication repair would not have helped here. The failure was length, and the repair was about length.
Case 3: "the market's fastest onboarding" appeared in the blurb, correctly quoted in the claims table. Under-specified Added: no comparative or superlative claim about a partner, even where the partner made it. Flag it in the table as refused. Fixed, and it caught two more in case 1 that I had read past twice — "leading" and "best-in-class" — which tells me something about my own checking.

7 · Handoff note

What this is for. Drafting the three weekly Partner Spotlight blurbs from whatever each partner supplied, with a claims table so the draft can be checked by someone who doesn't know the partner. Not for writing the newsletter's editorial section, which has a different voice and a named author.

What it's bad at. (1) It is only as good as what the partner sent, and a thin source produces a thin blurb — correctly. If the blurb is disappointing, the fix is usually to go back to the partner, not back to the prompt. (2) It cannot tell that an accurate blurb will annoy a partner. It has no view of the relationship, and two of ours are delicate. (3) The claims table quotes the supplied material, so if the partner's own release is wrong, the blurb will be wrong and the table will look reassuring — traceable is not the same as true.

What a user must check. Every table row: does the quoted sentence actually appear in the source, and does it say what the blurb says it says. Any claim in a blurb with no table row — that combination is the one that matters. Whether anything was flagged as refused, and whether you agree.

Self-grade

17 / 18. Criterion 1 scored 2. The "how you'd know it had gone wrong" row originally read "the partner complains", with nothing about how long that takes or what it costs. Rewritten above. The grader also noted that the record never says who the blurbs' second consumer is, which is why the partner now appears in the Consumer row.


Answers to the two questions in the drill

Don't read these until you've written your own answers.

Which anti-criterion did Record B's rejection turn on?

Needs facts Claude can't reach — that is the fatal one, and the other two that fire (expensive-and-undetectable, checking-costs-more) are both downstream of it. The sharpest move in the argument isn't naming the anti-criterion, though. It's the follow-through: the only repair available — transcribing 180 rate notes into a table — is itself the entire task, and if that table existed the job would be a spreadsheet lookup needing no model at all.

That pattern is worth carrying: when the repair for an unreachable-facts problem is "put the facts somewhere reachable", check whether doing so leaves anything for Claude to do. Sometimes it does. Often the data project was the real project all along, and noticing that is a better outcome than a working prompt would have been.

Which iteration row would you have argued with?

The intended answer is row 1 of Record B's log, classified as fabrication. It is defensible — specifics were asserted that no source supported. But under-specified is at least as good a reading: the first build never said what an explanation was allowed to rest on, so a gap was left in the specification and filled with something typical. The two classes have different repairs, and the record ends up applying both, in that order.

Other reasonable answers: Record A's row 3, where "under-specified" could be argued as "insufficiently decomposed" — detecting a conflict and then reporting it are two operations handed over as one. Or Record C's row 2, which is labelled wrong shape and could be read as under-specified.

The point is not that one label is right. It is that a log naming a class you can argue with is doing its job, and a log saying "the output was bad, so I improved the prompt" is not. If you disagreed with a row and could say why, you did the exercise correctly.