Design document· 8 modules · ~5 hours· Rev. 3.1

From Casual to Capable

The full curriculum design — every module’s objective, the misconception it has to defuse, its exercise, and how a learner checks themselves with nobody grading.

This is the design, not the course. It is written for whoever builds the modules, so it argues with itself in places and says plainly where it is uncertain. Learners get the lessons.

What changed in rev. 3

A design review marked three items P0 and several P1. All are applied. The three structural changes worth knowing before reading: Module 1 no longer claims Claude's reasoning is visible in its output (it isn't necessarily, and extended thinking makes the old claim plainly wrong); Module 2 no longer claims the context window is the boundary of what Claude can reach (memory, project memory, chat search and automatic compaction all cross it); and no exercise anywhere now depends on Claude failing on command. That last one is new design premise 4 and it governs the entire synthetic material spec.

Module 5 is also reframed from verification to evidence, which is a better concept and survives models getting more accurate.

Rev. 3.1 adds two production rules, both aimed at the same failure — a course that ages badly. The tier test turns premise 4 into a check a scriptwriter can run, and every exercise is now labelled. Volatile specifics are out, locations stay in removes page thresholds, effort-tier names and model lineups throughout, while keeping the handful of interface locations objective 4 depends on.

Four design premises

These are the load-bearing choices. Everything downstream follows from them, so they're the right thing to argue with first.

1. Diagnosis is the point; the mental model is scaffolding for it. The second objective — learners repair a bad prompt instead of giving up — is the one that actually changes behaviour. But you cannot debug a system you have no model of. So Modules 1–3 teach mechanism, and each one is cashed out as diagnostic leverage in Module 4. No mechanism is taught that doesn't later become a diagnostic move. If a concept can't earn its place that way, it's trivia and it gets cut.

2. Features are taught as answers to problems the mental model predicts. The obvious way to teach Projects, memory and Artifacts is a feature tour. Don't. A feature tour trains learners to ask "what can I do with a Project?" — exactly the tool-first thinking the third objective is trying to break. Instead, Module 2 establishes how information reaches a response, and Module 6 opens with: here are the mechanisms that put it there for you.

3. The course must be able to say "no." A course that only ever demonstrates AI succeeding produces learners who can't tell a fixable prompt from an unfixable task. So the diagnostic clinic in Module 4 contains one genuinely unfixable item, and the capstone accepts "I analysed this and AI is the wrong tool here" as a full-credit answer.

4. No learning objective may depend on Claude making a particular mistake live. New in rev. 3, and it constrains every exercise in the course. The original outline had six demonstrations whose teaching payoff required a stochastic system to misbehave on cue — miscount letters, fabricate a citation, forget a constraint, ignore an instruction not to fabricate. Any one of them may work today. None will work reliably across model releases, and this course is meant to outlive the model it was written against.

The rule in practice: the guaranteed teaching event is always a supplied artifact; the live run is always second and its outcome is always interesting. Learner examines a captured, dated real output that exhibits the failure and works out what's wrong with it — that part cannot break. Then they run it themselves. If today's Claude reproduces the failure, they've confirmed it. If it does better, the debrief has an answer ready: this is a system that improves, which is exactly why you check rather than assume. Either way the objective is met.

This is the premise with the largest production consequence, and it is the reason the review's ordering should be adjusted: it belongs at P0, not P1. The synthetic material set is the longest-lead deliverable in the whole build, and this rule changes what has to be in it — every affected exercise needs a captured "before" artifact that doesn't exist yet. Build the materials first and apply the rule second, and you build them twice.

The tier test — apply to every exercise, during scripting

Premise 4 stated as a check a scriptwriter can actually run. Every exercise in the course is labelled below, and anything that can't be labelled A or B doesn't ship.

TierWhat it isWhy it holds
A — cannot breakAnalysis of a supplied artifactThe teaching event already happened and is captured. No live system involved.
B — robust liveLearner constructs something and inspects it, with no required outcomeTests construction and inspection, not prediction. Works whatever the model does.
C — forbiddenObjective depends on a specific model behaviourShips a silent time bomb. Rewrite as A, or cut.

The companion copy rule, and it applies to every word of learner-facing text: never assert what Claude will do — say what to look for and why it matters. "Claude will miscount the letters" is a Tier C sentence even inside a Tier A exercise. "Check whether the count is right, and note what you find" is not.

What the tier test is not: an argument for dropping live work. A quiz-only course delivers objective 1 and abandons 2, 3 and 4 — you can test whether someone can classify a broken prompt, but not whether they can repair one, and confidence with Projects is not obtainable by reading about Projects. If the course ever does go assessment-only, the objectives should be rewritten down to match rather than quietly under-delivered.

Sequencing logic

The conceptual spine, in the designer's vocabulary rather than the learner-facing titles. Renaming Module 5 from verification to evidence is not cosmetic: it shifts the module from "catch the robot lying" to "know what any claim rests on and how much being wrong costs," which is a professional literacy skill that stays useful however accurate models get.

One reorder is defensible: moving M6–M7 earlier for a quick win. I'd still resist it. Learners who meet Projects before they understand how information reaches a response use them as folders, and that misuse is hard to undo. If drop-off data says otherwise, the cheaper fix is a five-minute unassessed tour preface, not a resequence.

Module 1 · Generation

Autocomplete That Read the Library

30 min

describe what Claude is doing between pressing enter and seeing text, and predict from that description which kinds of task need extra care.

Core concept

Claude produces its answer sequentially, one chunk of text at a time, each shaped by everything in front of it. Three consequences do most of the explanatory work for the rest of the course:

Plus the small one that explains a daily annoyance: text is handled in chunks, not letters, which is why character-level tasks are shaky. Name them once and move on — these chunks are what everyone else calls tokens, and they're the unit usage limits are counted in, so the word is worth recognising even though the course won't use it again. And output is sampled, not fixed — the same prompt twice can give two different answers, which sets up the reroll trap in Module 4.

Biggest misconception

"It's looking things up." The database metaphor is the most damaging wrong model a casual user carries, and the parent of every fabrication surprise later in the course. Nearly all of them hold it implicitly.

The move, under premise 4: lead with a supplied captured Claude response — dated, archived, web search off — asking for sources on an obscure topic, complete with perfectly formatted citations, plausible author names and real-sounding journals that do not exist. The learner's task is to work out what's wrong and how they'd check. That teaching event is guaranteed because it's a static artifact.

Then they run it live. If today's Claude fabricates too, they've confirmed it first-hand. If it declines, or hedges, or searches instead, the debrief says: good — this is a system that improves, which is exactly why the skill is checking rather than assuming. The original design staged this live and would have broken silently the first time the model got better at it.

Hands-on · "Find the edges" · Tier A → B

Each item is a supplied captured artifact first, then an optional live re-run. Predictions recorded before either.

  1. Letter-counting in a long word — supplied transcript showing the miscount, then try it live.
  2. Six-digit multiplication, supplied as a captured pair: bare answer, and the same task worked through step by step. Learner compares. Then live.
  3. Same question asked three times — supplied triple showing the variation, then live.
  4. The fabricated-citation artifact above, then live with search off and on.
  5. A recent-events question, supplied with search off and on. Same mechanic as #4 on purpose, and the repetition is the point: #4 fails because it never knew, #5 because it stopped learning. One repair — grounding — for two causes.

Every live re-run carries the same instruction: record what actually happened, including "it worked fine." That framing is what makes the exercise robust instead of fragile.

Self-check · auto-graded, 7 items

Predict-the-failure: a task is described, learner selects needs care / straightforward and then the reason from four mechanism-based options. Graded on the reason — a learner right for the wrong reason will fail Module 4. One item must separate never knew it from knew nothing after that date, since both produce identical-looking confident wrongness.

Module 2 · Available information

What It Can Actually See Right Now

30 min

say what information is available to Claude for the response it's about to give, name where each piece came from, and take deliberate action when a conversation stops being reliable.

Core concept

Claude can only reason over information made available to the current response. Some of that comes directly from the conversation in front of it. Some is brought in by other machinery: memory, project knowledge, chat search, uploaded files, connectors, web search.

This replaces rev. 2's claim that "anything outside the context window does not exist," which was already false at the time of writing. Claude can search previous chats, memory injects persistent entries, projects retrieve from project knowledge, and long conversations can be automatically compacted — earlier messages summarised so the conversation continues, rather than simply falling off the end. Teaching the absolutist version gives learners a model that Module 6 then contradicts, and being contradicted by a later module is how a course loses trust. The reframe is also a better foundation for M6, which becomes "here is the machinery that makes information available" rather than a rescue from a limitation.

The two things the learner should be able to do:

And the honest version of the old claim: long conversations can become less reliable or less focused. Not monotonically, and automatic context management makes any stronger claim brittle. Teach the recognisable symptom — dropped constraints, drift, self-contradiction — and the cheap repair, which is a fresh chat with a restated brief.

Biggest misconception

Two, pulling in opposite directions.

"But I already told it that." Learners conflate having said something with it being reliably available now, and read the failure as carelessness rather than as a question of what actually reached this response.

"More is better, so I'll paste everything." The reasonable inference from the first half of the lesson, and it produces worse output.

The move: a supplied transcript — a long, realistic, drifting conversation where a constraint stated early is quietly violated late. The learner finds where it went and diagnoses why. Guaranteed to work, because it already happened. Then the live half is the contrast, not the failure: same question in a fresh chat with just the brief, compared against the tail end of the supplied transcript. Producing a good answer live is easy to reproduce; producing a bad one on cue is not.

Hands-on · "The drifting brief" · Tier A → B

Supplied: Northwind Consulting — Q3 Board Memo, 12 fictional pages, with a constraint stated on page 2 ("no recommendation may assume additional headcount"), plus a captured 20-turn transcript in which that constraint is eventually violated.

  1. Read the transcript's final recommendation. Find the violation. Trace back to where the constraint stopped being honoured, and write one sentence on why.
  2. Live: fresh chat, memo plus constraint plus the final question, nothing else. Compare.
  3. Live: the narrow-question test — all 12 pages versus the two relevant pages. Compare precision. This one is robust because it tests focus, not failure.
  4. Account for availability: for your live answer, list what was available and where each piece came from.
Self-check · auto-graded, 7 items

"Where would this come from?" — a scenario is described (a file uploaded three chats ago, a project instruction, something mentioned yesterday, a document still on the learner's desktop) and the learner identifies whether it could reach the current response and by what route. Directly pre-loads M6's routing quiz, and the "still on your desktop" item is the one that separates understanding from guessing.

Module 3 · Specification

Say What You Actually Want

35 min

build a prompt from components rather than copying a template, and explain what each component is doing to the output.

Core concept

A prompt is a specification, and vagueness in the spec gets filled in with whatever's most typical. Six components:

The unifying explanation, from M1: an under-specified prompt leaves Claude to pick something typical, and typical averages out to generic. Every component above narrows the target.

Biggest misconception

"There are magic words." Casual users collect incantations — "You are a world-class expert…", elaborate politeness, ALL CAPS, offering tips — and believe the phrasing itself is the lever. This is what makes learners template-dependent and unable to build a prompt for a problem nobody has written a template for. It's the direct obstacle to the third course objective.

The move: a controlled A/B in the exercise. Same task, one prompt loaded with flattery and emphasis, one plainly worded but specific about audience, length and exclusions. Then the reframe: prompting is specification, not spellcasting. Reinforce by having the learner strip a supplied incantation-heavy prompt to its actual content — usually about a third survives and the output holds up.

Premise-4 note: this A/B is safe to run live. It doesn't need Claude to fail — it needs specificity to help, which is stable across model versions and the whole point of the module. Supply a captured comparison anyway as the fallback for a learner whose two outputs come out equally good.

Hands-on · "Four drafts" · Tier B

Supplied: a messy fictional customer complaint thread (Meridian Software, billing dispute, three participants, one of them wrong about the facts). Learner writes a reply four times:

  1. Task only. ("Write a reply to this.")
  2. Task + context + output shape.
  3. Add constraints and non-goals.
  4. Add one example of the desired tone (supplied).

Scored each round against a fixed supplied rubric, so each component's marginal contribution is a number rather than a vibe. Then the incantation A/B.

Role variants: the same thread as an internal escalation (manager), a post-incident note (ops), or a churn-risk summary (CS lead). Shared rubric across all three.

Self-check · auto-graded, 8 items

Component labelling: given a strong prompt, tag each part. Rewrite selection: given a weak prompt and four rewrites, pick the one that adds specification — distractors are longer, more polite and more emphatic but no more specified. That item type is the misconception, tested directly.

Module 4 · Diagnosis

It Gave Me Garbage

40 min

look at a bad output, name which of six failure classes it belongs to, and apply the matching repair — instead of rerolling, rewriting at random, or giving up.

Core concept

The keystone module. Bad output is not one thing and the repairs are not interchangeable.

FailureLooks likeRepairFrom
Missing informationConfidently answers a different situation than yoursSupply what it couldn't haveM2
Under-specifiedValid, competent, not what you wantedAdd constraints or an exampleM3
FabricationSpecifics that are wrong — numbers, names, citationsGround it, or scope it to a sourceM1
Wrong shapeRight content, wrong length/format/registerSpecify format, or show oneM3
Lost the threadContradicts itself, drops earlier instructionsFresh chat, restated briefM2
Insufficiently decomposedDoes all of it, shallowlyBreak into steps; give it room to workM1

Renamed from "task too big" on review, and the reviewer was right. "Too big" names a fixed capability boundary while the repair column immediately says it isn't one — the label was arguing against its own row. "Insufficiently decomposed" points at the fix.

Three habits taught alongside the taxonomy:

Biggest misconception

"It can't do this." Learners attribute a capability ceiling when the cause is nearly always their own specification or missing information. They stop one move too early, and that single habit separates a casual user from a capable one.

The move: the clinic's answer key is the argument. Five of six items are fixable, several with a one-line change, and the learner does the fixing.

The sixth is genuinely unfixable — a fact that is private and in no supplied source — and its key says so plainly. This should survive review. A diagnostic that can never return "stop, wrong tool" is not a diagnostic; it's a motivational poster. Learners need calibrated persistence, not unlimited persistence.

Note the unfixable item is now a private-fact item, not character counting. Under premise 4, "Claude can't count characters" is exactly the kind of claim that may stop being true; "Claude cannot know your colleague's unshared salary figure" will not.

Hands-on · "The clinic" · Tier A → B

Six supplied prompt-and-output pairs — captured, dated, shuffled, one per failure class. For each: classify, state the repair, then apply it live and compare against the supplied model repair. Applying is mandatory; classification alone produces learners who pass the quiz and can't fix their own work.

Wholly supplied inputs, so premise 4 is satisfied by construction — this exercise was already robust and needed no redesign.

Closing drill: the learner brings their own worst recent Claude interaction and runs the taxonomy on it. First time they point the method at their own material.

Self-check · auto-graded, 8 items

Vignette classification: prompt plus bad output, learner picks failure class and repair as separate scored fields. Distractors are the adjacent classes — under-specified and wrong-shape are the pair people genuinely confuse. At least one item's correct answer is "not a prompting problem."

Ship the taxonomy as a permanent reference card, alongside the capstone's seven-step process. On review these are the two most reusable objects in the course, and both should outlive the modules that introduce them.

Module 5 · Evidence

What Is This Based On?

30 min

for any claim in an output, say what it rests on, how to inspect that basis, and how much it matters if it's wrong.

Core concept

Reframed on review from verification to evidence, and the reframe is the improvement. The old module was organised around catching the model lying, which dates badly as models improve. This one is organised around a question that stays useful regardless: what is this claim based on?

Two independent questions, which rev. 2 collapsed into one and shouldn't have:

A · Can I trace it?B · Is the source any good?
A source you suppliedAuthority — who produced it, and do they know?
A source Claude retrieved (search, connector, project knowledge)Relevance — is it about your actual question?
Unsupported model knowledgeRecency — is it current enough to act on?
Completeness — is it the whole picture or one slice?

Column A is a traceability ordering, not an accuracy ordering, and rev. 2 wrongly implied otherwise. A document you uploaded can be outdated, wrong, incomplete, adversarial, or simply the wrong document — and a retrieved primary source from a statistical agency may be far more trustworthy than a PDF someone emailed you. The two questions are orthogonal and both have to be asked.

Which yields the sentence the module exists to install:

Grounded does not mean true. Grounded means you have something concrete to inspect.

Then triage, because effort has to be proportional: risk is roughly cost of being wrong × how traceable the claim is. The reliably high-attention zones, taught as a checkable list: specific numbers, citations and URLs, direct quotes, names and dates, anything legal/medical/financial, recent events, and anything about the learner's own organisation.

And the behavioural rule, stated as behaviour rather than as an empirical absolute: do not use confident tone as evidence of accuracy. Claude writes an unsupported claim in the same register as a well-supported one. Most people are unconsciously reading prose style as a confidence signal, and the fix is a habit, not a belief about what tone does or doesn't correlate with.

Biggest misconception

A matched pair, and the module must hit both or it pushes learners from one into the other.

"I'll just tell it not to make things up." Instruction doesn't reach the mechanism. Taught from a supplied captured example — an ungrounded question with "only state facts you are certain of" attached, and an unsupported claim in the answer anyway, complete with a confidence disclaimer. Supplied rather than live, per premise 4: this demo previously required the model to ignore an instruction on cue.

"It makes things up, so I have to check everything." The overcorrection. Produces learners who verify a tone adjustment as carefully as a revenue figure, find the tool nets out slower than doing it themselves, and quit. Quieter than the first failure and probably more common among cautious professionals.

The move: triage as the organising idea, with a worked procedure. The learner should leave with a specific answer to "what do I check on a busy Tuesday" — the numbers, the names, the citations, and not much else when the rest traces to a document you supplied and trust.

Hands-on · "The summary that went too far" · Tier A → B

Supplied: a fictional Harbourview Market Assessment (six pages) and a summary of it containing three factual errors and two unsupported inferences — claims that sound sourced but appear nowhere in the document. Fully supplied, so it was already premise-4 safe.

  1. Find all five unaided. Most learners find the errors and miss the inferences; the debrief should predict that gap before revealing it.
  2. Re-run the summary live with a grounding prompt — quote the supporting line for each claim, flag anything inferred rather than read. Compare what becomes inspectable.
  3. Column B drill: the assessment cites two sources — one a trade-body press release, one a statistical agency release. Same traceability, very different authority. Which claims would you act on?
  4. Apply the triage list to a supplied output; mark only what genuinely warrants checking.
Self-check · auto-graded, 8 items

Two-dimensional sorting: for six sentences, learner marks both traceable to what and act on it / check it first. Correct answers include several "act on it" — a learner who flags everything scores badly, which is the calibration being tested. Two items pit a traceable claim from a weak source against an untraceable claim that happens to be plausible, which is where the two dimensions come apart.

Module 6 · Persistence

Making It Stick

35 min

choose the right scope for a piece of persistent context, and set up a Project that makes a recurring task meaningfully faster.

Core concept

Calls back to M2: information has to be made available to a response, and typing it every time doesn't scale. These features do it for you. Teach them as scopes, not as boxes — which is a rev. 3 correction, because the tidy three-bucket taxonomy is no longer accurate:

Why "scopes" and not "boxes." Memory is not a single account-wide store: each project has its own separate memory space and its own project summary, kept apart from other projects and from non-project chats. So "Project" and "Memory" are not parallel alternatives — memory has a project scope. Chat search adds another retrieval route with its own boundaries (all chats outside projects; within a project, that project only). The honest teaching is a single question — how widely should this apply? — answered along a scope gradient, rather than three tidy containers a learner will later find leaking into each other.

The retained decision rule, reworded for scope: true about me everywhere → Instructions. True about this body of work → that Project. Happy for it to accumulate on its own → Memory, at whichever scope you're working in.

Biggest misconception

"I told Claude that already — why doesn't it know?" Learners don't distinguish said in a conversation from stored at a scope. Everything gets filed as "I told it," so the behaviour looks arbitrary — sometimes it remembers, sometimes not, no discernible rule. This is the confusion that makes people quietly distrust the tool.

The move: make the invisible visible. Have them open Settings → Memory and read their own entries. Nothing else lands as hard as seeing discrete, editable, listed statements where they'd imagined a vague cloud of recollection.

Runner-up — don't skip it: quiet unease about memory. A learner who is uncomfortable and can't name why simply won't use the feature. Address it factually and early — pause, reset, export, incognito, per-project separation, entries you can read — rather than reassuringly. The controls are the reassurance.

Hands-on · "Build the container" · Tier B

Supplied: a fictional Lumen & Co. brand voice guide and three past newsletters, for a recurring "draft this month's newsletter section" task.

  1. Create the project, upload both knowledge files, write project instructions including at least one output-format constraint (carried from M3).
  2. Run the recurring task inside the project.
  3. Run the identical task in a bare chat. Compare — the delta is the argument for Projects, and it doesn't depend on anything failing.
  4. Write one line of Instructions for Claude that would be true of every project you ever run, and notice how few things qualify.
  5. Open Settings → Memory. Inspect one entry and say why it belongs at that scope.

Step 5 previously required read, edit and delete. Cut on review, and correctly: requiring a learner to destroy their own real memory entry is needlessly invasive for an educational exercise, and inspection alone proves the concept. Editing and deletion move to an optional extension using a deliberately disposable entry the learner creates first.

Role variants: newsletter section, weekly status roll-up, or candidate screening notes.

Self-check · auto-graded, 8 items

Scope routing: eight facts ("I always want metric units"; "our fiscal year starts in April"; "this client hates bullet points") placed at the right scope. Deliberate near-misses, since the everywhere-versus-this-project line is where the real judgement lives. One item should have a defensible answer at two scopes, with feedback that explains the trade rather than declaring a winner — that's where scope-thinking beats box-thinking.

Module 7 · Surfaces

Real Material In, Real Work Out

40 min · see timing note

get source material into Claude, judge how completely it was actually read, produce something in an Artifact that survives outside the conversation, and set how hard Claude works on a task.

Core concept

And the dials underneath — how hard Claude works

Clicking the model name opens three separate controls, and casual users almost never touch any of them. This pays off Module 4's seventh repair and closes the loop opened in Module 1.

Deliberately deferred from Module 4. A learner who reaches for max effort before checking whether their prompt was under-specified has swapped a diagnostic habit for a dial — the exact behaviour Module 4 exists to prevent.

Biggest misconception

"I uploaded it, so it's read the whole thing." Learners treat a successful upload as a guarantee of complete, accurate reading, then trust a page-40 claim as much as a page-1 claim. Genuinely dangerous, because it manufactures confidence at exactly the moment M5 taught them traceable sources are better.

The move: stop relying on a live degradation demo, and stop reaching for specifics too. Teach the principle:

"Uploaded" is not a single file-reading mode. How completely a file gets read depends on its format and its size, and you cannot tell which you got by looking at the reply.

A PDF, a spreadsheet and a scanned page are not handled the same way, and a long file and a short one are not handled the same way. The learner does not need the mechanism, and giving them a number would be worse than giving them nothing — see the note below. What they need is the conclusion: if it matters, demand the supporting line.

Why no thresholds here. An earlier revision taught specific page counts from Anthropic's platform documentation. Cut, for two reasons. First, those describe API behaviour and were never verified against the consumer path this course teaches — an unverified number is exactly the kind of claim Module 5 tells learners not to accept. Second, and more durably: a named threshold invites learners to treat it as a reliable operating parameter, so it fails twice over when it changes — the number is wrong, and the habit it built was wrong. "Format and size both matter, and you can't tell by looking" survives every future change, and generalises to formats and limits nobody has documented yet.

The live half becomes a check rather than a gotcha: ask a question whose answer sits deep in a dense table, demand the exact supporting line, verify it against the source yourself. "Quote the line" is the durable habit either way.

Hands-on · "From mess to deliverable" · Tier B

Supplied: a deliberately untidy fictional spreadsheet (inconsistent date formats, merged cells, a stray total row) plus a short covering memo.

  1. Upload both; run the quote-the-line check before trusting anything.
  2. Produce a one-page summary as an Artifact.
  3. Iterate twice in place with targeted change requests — not regenerations.
  4. Take the hardest analytical question the spreadsheet supports; answer it on defaults, then with thinking on. Compare, and note honestly whether it was worth the wait — on an easy question it won't be.
  5. Publish, and open the link in a private window to see exactly what a recipient sees.
  6. Optional extension: branch by editing an earlier message, then return to the original.

Timing, flagged on review and it's a fair hit. Forty minutes for six live tasks involving an upload, an Artifact, two iterations, a settings comparison, publishing and a branch was optimistic — it ignored upload time, response latency and reading. The honest options were: split into two modules (breaks the 6–8 module brief), extend past the 40-minute ceiling (breaks the sitting constraint), or cut. I cut: branching is now optional, which is defensible because it's the one item here that isn't load-bearing for the capstone. 40 minutes is still the tightest module in the course and should be the first thing playtested.

Connectors remain a parallel path, never a gate. A meaningful share of learners won't connect a personal Gmail for a course exercise, and that's a reasonable position. The file-upload route satisfies the same objective and both count. Progression must never depend on a learner granting a permission they'd rather not grant.

Self-check · auto-graded, 8 items + artifact checklist

Quiz on how each route makes information available. The file items must test the conclusion — given a scenario, does this claim need the supporting line demanded? — and never a threshold or a processing rule, which is where a quiz bank rots fastest and hardest, because a stale item marks a correct answer wrong. Two dial items, scenario-based rather than definitional: given a task, is this a prompt problem or a settings problem? "Prompt" is correct more often, which protects the Module 4 reflex.

Plus a binary self-verified checklist: published link opens in a private window ✓, two in-place iterations ✓. Binary and self-evident — the learner can't half-convince themselves they did it.

Module 8 · Design

Capstone — Design Your Own Use Case

45–60 min

start from a real work problem, decide honestly whether Claude fits it, and if so build, test and document a working setup — with no template supplied at any point.

P0 fix — the data contradiction

The review caught a genuine blocker and it was the sharpest catch in the pass. Rev. 2 held both "synthetic only — no exercise may require the learner's real work data" and a capstone that said pick a real recurring task, load its knowledge, test it on real cases, and publish the result to a public link. Those cannot coexist. Publishing on Free/Pro/Max makes an Artifact reachable by anyone with the link.

Two changes. First: Publish is removed from the capstone entirely. The review is right that it contributes nothing to the capstone's objectives — the learner already demonstrated publishing in Module 7 — while creating a real disclosure risk at exactly the moment they're working with material closest to their actual job. The design record is submitted, not published.

Second, and here I've landed somewhere different from the review. The review proposed learners build against a supplied synthetic analogue instead of their own material. I think that over-corrects and costs the capstone its best property — motivation and transfer come precisely from building something the learner actually wants on Monday. Read the constraint literally: no exercise may require real work data. It does not say real data is forbidden; it says the course can never make it the price of admission.

So: task selection stays real, data is at the learner's discretion, and a synthetic analogue is supplied for every capstone shape so anyone who can't or won't use their own material has a complete path with nothing withheld. The graded artifact is the design record, which describes a task's structure — inputs, outputs, what good looks like, how you'd know it failed — and needs no real content to be excellent. That satisfies the constraint as written, keeps the capstone worth doing, and removes the disclosure risk that actually mattered.

The open prompt

Pick a task you personally did more than three times in the last month and did not enjoy. Do not pick a task because it sounds like a good AI task. Pick it because it is tedious, and then find out whether AI helps. Describe its structure in full. Use your own material only if you're free to — otherwise take the supplied synthetic analogue, which is built to the same shape and costs you nothing.

The second sentence is the whole capstone. Left off, roughly everyone submits summarisation or email drafting and the course has assessed template-matching under another name. Constraining the selection rather than the solution is what forces problem-first reasoning.

The seven-step design process

This process — not any individual module — is the durable artifact of the course. Ship it as a reference card, alongside Module 4's failure taxonomy.

  1. Describe the task as it exists today. Inputs, outputs, who consumes it, what "good" looks like, how you'd know it had gone wrong. Structure, not contents.
  2. Run the fit test — argue both sides. Anti-criteria supplied: tasks needing facts Claude can't reach; tasks where a wrong answer is expensive and hard to detect; tasks that are mostly judgement you can't articulate; tasks where checking costs more than doing. "This doesn't fit" is a full-credit answer if the reasoning holds — the learner then picks a second task and continues.
  3. Decide the shape. One-off chat or a Project? What should the output rest on (M5)? At what scope should context persist (M6)? Does the output want to be an Artifact (M7)?
  4. Build it. Instructions, knowledge, prompt.
  5. Write the test set before looking at any output. Three cases including one deliberately hard or edge case, pass criteria fixed in advance. Writing criteria after seeing output is how people talk themselves into mediocre results, and naming that trap is part of the instruction.
  6. Run, diagnose, iterate — at least twice, documented. Every failure classified with the Module 4 taxonomy and repaired accordingly. A capstone with no documented iteration is incomplete regardless of how good the output looks, because the iteration is the skill being assessed.
  7. Write the handoff note. What this is for, what it's bad at, what a user must check. Forcing articulation of limits is the cleanest proof that Modules 4 and 5 landed.
Deliverable

The working setup, plus a one-page design record covering all seven steps. Produced as an Artifact — not published. Submitted or self-graded privately.

Biggest risk at this stage

Not a misconception — a completion risk. Highest effort, lowest structure, arriving when novelty has worn off. This is where an async cohort haemorrhages learners.

Mitigations: four separately-saveable checkpoints so it survives several sittings; three fully-worked example design records from different roles — including one that legitimately concludes "no fit" — so the blank page is never truly blank; and step 1 moved to the end of Module 7 as a five-minute pre-task, so learners arrive already holding a candidate task instead of facing selection and design at once.

Assessment and completion, with no one grading

1. Auto-graded knowledge checks — per module

Six to eight items, 80% to pass, unlimited retries from a shuffled pool so retrying isn't memorising. The critical rule: wrong answers return a mechanism-level explanation, not "incorrect." With no instructor, the quiz is the only thing that can catch a misconception at the moment it's exposed, so it must teach rather than merely measure. Budget four distractor explanations per item — the most underestimated cost in the build.

2. Behavioural gates — per hands-on

Binary, self-evident, hard to fool yourself on. "Your project instructions contain an output-format constraint — yes or no." These carry more weight than any confidence rating, because the learner checks a fact about an object rather than rating a feeling. Never ask "do you feel confident with Projects?" — the answer is uncorrelated with anything.

3. Claude as the capstone grader

The learner pastes their design record into a supplied grading prompt that scores against published criteria and returns specific gaps. Defensible rather than a cop-out because the criteria are public and auditable, and because by Module 5 the learner has been taught not to take Claude's judgement on trust. The instruction should say so outright — challenge this grade; if you think it's wrong, argue with it and see if it holds. The assessment becomes one more repetition of the skill. Note this also works on a structure-only design record, so it never requires real content.

4. Spaced, interleaved retrieval — day 7 and day 30

Two short cumulative checks, drawing across all modules and deliberately interleaved — a diagnosis item, then a scope item, then an evidence item — rather than blocked by module. Blocked practice inflates in-the-moment performance and does much less for retention. The only element targeting durable recall rather than course-completion, and the cheapest thing here to build.

Completion criteria

Day-7 and day-30 checks are reported but deliberately don't gate the certificate. Gating them creates an incentive to defer completion, and the retention value doesn't depend on compulsion.

Before the build starts

  1. The synthetic material set is the critical path, and premise 4 just enlarged it. Every affected exercise now needs a captured, dated "before" artifact that doesn't exist yet: a fabricated-citation response, a 20-turn drifting transcript, a step-by-step arithmetic pair, an ignored-instruction example. These carry the pedagogy; too clean and every exercise succeeds and nothing is learned. Scope as a named deliverable with its own review pass — and build it after premise 4 is settled, not before.
  2. Volatile specifics are out; locations stay in. The test applied throughout: does the learner need this to take an action, or only to feel informed? Page thresholds, effort-tier names and model lineups are informed-only, and they're cut — a named number fails twice when it changes, since the number is wrong and the habit it built was wrong. What stays is where things live (Settings → Memory, the project knowledge panel, the Publish button, the model-name menu), because objective 4 is unreachable without it. That's roughly half a dozen interface locations across the whole course — small enough to be a pre-cohort checklist rather than a versioned specifics file with an owner. Keep screenshots in one replaceable asset set with capture dates, and prefer the written path in body copy so the text survives a UI change even when the image doesn't.
  3. Module 5 still needs the most review, though less than before. The evidence reframe means it leans less on an authored account of hallucination mechanics and more on a question anyone can apply, which is more robust. Still: get it read by someone who understands the mechanism and someone who represents the audience.
  4. Market it as 5–6 hours, not 5. Designed instruction is ~285–300 minutes, but that excludes quiz retries, setup friction, response latency, project creation, publishing and reading feedback. Under-promising learner time in an async course is a completion risk, not a marketing win.
Still open. (1) Module 7 at 40 minutes remains the tightest sitting in the course even after cutting branching — playtest it first, and if it fails, the real choice is a ninth module against the 6–8 brief. (2) The capstone data policy above is my judgement, not the reviewer's — I kept real task selection with optional real data where they'd have moved wholly to synthetic. If the deployment context is regulated or the learner population is enterprise, take their stricter version instead; the synthetic analogues exist either way. (3) I'd still put premise 4 at P0 rather than P1, because the synthetic material spec depends on it. (4) One course-wide fictional company family, and learners assumed to have Claude open in a second window.