One per shape the capstone commonly takes. Each is built to the structure of the real thing, including the parts that are annoying, so a capstone built on one of these is a complete capstone.
Task selection in the capstone is real — you pick a tedious thing you actually do. What you feed it is your decision. If your material can't leave your systems, or you'd simply rather not, take the analogue below whose shape matches your task and build against that. Nothing is withheld from this path. Every step works, every grading criterion applies, and the design record you produce describes the structure of your real task either way.
Do step 1 — the task description — about your own task. Its inputs, its outputs, its consumer, what good looks like, how you'd know it went wrong. That is structure, not contents, and it needs nothing confidential in it.
Then pick the analogue whose shape matches, and run steps 3 to 7 on that. Your build, your test set and your iteration log are all about the analogue. Your handoff note is about the analogue too — and it will still tell you most of what you need to know about the real one, because the limitations you find are properties of the shape rather than of the data.
All six analogues are written, not captured. They use the same fictional companies as the rest of the course. They are deliberately untidy: inconsistent naming, contradictions, missing entries, dates without months, values stated three different ways. A tidy analogue would let every build succeed and would teach you nothing about your own inputs, which are not tidy either.
Yours is this shape if: several people send you something in whatever format suits them, on a schedule, and you merge it into one document with a fixed structure that somebody else acts on.
The output template: four headings in this order — Progress this week, Decisions needed, Risks, Next week. One page. Sent to the client's programme director, who builds Monday's agenda from Decisions needed.
What good looks like. All six workstreams named, including Change & comms, which sent nothing. Every decision phrased as a question with an owner and a date. No internal shorthand. The client calls Dev's workstream "Wave 2" and Marcy's "Data migration" — the note uses the client's names.
The hard case is already in there. Rob says the cutover completed Tuesday; QA says it has been blocked pending that cutover since Monday. One of those is stale and the material doesn't say which. A build that picks a winner has failed; a build that surfaces both under Decisions needed has passed.
Where the anti-criteria bite on this shape. Usually the second one: an omitted item is expensive and invisible, because you'd be scanning for something that isn't there. The standard repair is to force the absence to become a sentence — name every workstream, and write "no update received" rather than saying nothing.
Yours is this shape if: a table arrives on a schedule and you write a line of commentary per row, some of which the table supports and some of which it doesn't.
What good looks like. One line per row, under twenty words, in the export's order. Rows the data explains cite the column that explains them. Rows the data doesn't explain say so plainly, because those are the rows that get a phone call — "cause not in data" is a correct answer and the most valuable one in the file.
Hard cases to put in your test set. Bellweather's 74% drop, which is a billing reclassification and not a usage collapse, and which is only knowable from the second file. The Halloway row, which has no current-week figure at all. And Ashgrove, which appears twice under two spellings with wildly different volumes — two accounts, or one?
Where the anti-criteria bite. The first: half of what makes a row explicable lives outside the export, in someone's head. The repair is not to guess better — it's to make the output say which rows it couldn't explain. Watch for the over-correction, where everything comes back as "cause not in data" including the rows where the plan-change flag plainly explains it.
Yours is this shape if: you turn source material of wildly varying quality into short pieces that all have to sound the same, and being wrong about the source is worse than being dull.
What good looks like. Three blurbs that don't read like each other, each saying something specific, each carrying a claims table: every factual assertion in the blurb sitting beside the sentence from the source it came from.
The hard cases are all three, for different reasons. Partner 1 contains "the market's fastest onboarding", which the voice guide forbids repeating — a good build refuses it visibly rather than dropping it silently. Partner 2 supports perhaps 20 words, so a 60-word blurb has been padded or invented. Partner 3 supplied nothing at all, and the only correct output is a statement that there is nothing to write from.
Where the anti-criteria bite. The second, hard: a wrong claim about somebody else's product is expensive and you are the wrong person to detect it. The repair that saves this shape is making every claim traceable, which turns checking from "do I know this product" into "does this quoted sentence say that".
Yours is this shape if: free text arrives and you pull the same handful of fields out of it, over and over, into something with columns.
What good looks like. Every field filled or explicitly marked unknown. Version stated as it appears, not normalised into a guess. Severity justified by something in the text rather than by tone — #8841 shouts and is genuinely P1; #8843 is calm and is a data-integrity discrepancy, which may matter more than it sounds.
Hard cases. #8841 contains two separate issues, one urgent and one cosmetic, and a version that isn't a version. #8847 is a follow-up to #8841 and adds a new symptom — is that one ticket or two? #8852 isn't a support ticket at all.
Where the anti-criteria bite. Usually the third: severity is a judgement and you may find you can't write the rule. Try writing it three times before deciding. Often the extraction of stated facts hands over cleanly and the severity call doesn't — which is a split, not a rejection.
Yours is this shape if: a queue arrives and each item needs a category, a priority and a destination, and the cost of the wrong destination is that it sits unread.
What good looks like. Every item categorised and routed, with the phrase from the item that justifies the category. Duplicates linked rather than counted twice. Items that can't be categorised from what's written land in a "needs a human" bucket rather than getting a plausible guess.
Hard cases. R-203 and R-208 are described as annoyances but are bugs. R-201 and R-204 are the same request from two customers. R-205 is a question wearing a feature request's clothes. R-207 says nothing at all. R-206 is a compliance item that a naive read routes to product, which is the one misroute here with a commercial consequence.
Where the anti-criteria bite. The third and the second together. Priority depends on "how blocked they are", which nobody has defined — so either you define it now, in writing, or you hand over the categorisation and keep the priority. And a misroute is cheap unless it's R-206, which is expensive and silent: nothing tells you a security item is sitting in the product backlog.
Yours is this shape if: several submissions arrive and you score each against criteria that already exist on paper, and the reason for the score matters as much as the score.
What good looks like. A score per criterion per vendor, each with the sentence it rests on quoted. Where the response doesn't address the criterion, the score reflects that rather than being inferred from the surrounding confidence.
Hard cases. Vendor A's backups are "replicated for resilience" — replicated where? The response doesn't say, which is a 1 and not a 2. Vendor B's C2 answer is about uptime, which is a different question entirely, and its C1 answer is reassurance rather than information. Vendor C is strong and still not a 3 on C2, because it says outright that it doesn't publish historical performance — a build that rounds a good answer up to full marks has failed the case.
Where the anti-criteria bite. The fourth, quietly: if checking a score means re-reading the whole response anyway, you've saved nothing. The repair is the quoted sentence, which turns verification into a scan. And the second, if a score feeds a real decision: a wrong score on a criterion nobody re-reads is expensive and silent.
That is a useful finding rather than a gap. Two things it usually means.
Your task is more than one shape. Most real tasks are. A weekly report that extracts fields, scores them and then writes prose is shapes 4, 6 and 1 stacked. Build one layer, test it, and only then decide whether the next layer belongs in the same handover — that's the Lesson 4 decomposition move, and stacking shapes without testing each is the most common way a promising capstone stops being diagnosable.
Or the thing that makes your task tedious isn't the text handling at all. If the tedium is opening eleven tabs, waiting for an export, or chasing people for their input, then no arrangement of these six shapes touches it. That's a fit test returning no, for the first anti-criterion, and it is worth writing up as exactly that.