The full curriculum design — every module’s objective, the misconception it has to defuse, its exercise, and how a learner checks themselves with nobody grading.
This is the design, not the course. It is written for whoever builds the modules, so it argues with itself in places and says plainly where it is uncertain. Learners get the lessons.
A design review marked three items P0 and several P1. All are applied. The three structural changes worth knowing before reading: Module 1 no longer claims Claude's reasoning is visible in its output (it isn't necessarily, and extended thinking makes the old claim plainly wrong); Module 2 no longer claims the context window is the boundary of what Claude can reach (memory, project memory, chat search and automatic compaction all cross it); and no exercise anywhere now depends on Claude failing on command. That last one is new design premise 4 and it governs the entire synthetic material spec.
Module 5 is also reframed from verification to evidence, which is a better concept and survives models getting more accurate.
Rev. 3.1 adds two production rules, both aimed at the same failure — a course that ages badly. The tier test turns premise 4 into a check a scriptwriter can run, and every exercise is now labelled. Volatile specifics are out, locations stay in removes page thresholds, effort-tier names and model lineups throughout, while keeping the handful of interface locations objective 4 depends on.
These are the load-bearing choices. Everything downstream follows from them, so they're the right thing to argue with first.
1. Diagnosis is the point; the mental model is scaffolding for it. The second objective — learners repair a bad prompt instead of giving up — is the one that actually changes behaviour. But you cannot debug a system you have no model of. So Modules 1–3 teach mechanism, and each one is cashed out as diagnostic leverage in Module 4. No mechanism is taught that doesn't later become a diagnostic move. If a concept can't earn its place that way, it's trivia and it gets cut.
2. Features are taught as answers to problems the mental model predicts. The obvious way to teach Projects, memory and Artifacts is a feature tour. Don't. A feature tour trains learners to ask "what can I do with a Project?" — exactly the tool-first thinking the third objective is trying to break. Instead, Module 2 establishes how information reaches a response, and Module 6 opens with: here are the mechanisms that put it there for you.
3. The course must be able to say "no." A course that only ever demonstrates AI succeeding produces learners who can't tell a fixable prompt from an unfixable task. So the diagnostic clinic in Module 4 contains one genuinely unfixable item, and the capstone accepts "I analysed this and AI is the wrong tool here" as a full-credit answer.
4. No learning objective may depend on Claude making a particular mistake live. New in rev. 3, and it constrains every exercise in the course. The original outline had six demonstrations whose teaching payoff required a stochastic system to misbehave on cue — miscount letters, fabricate a citation, forget a constraint, ignore an instruction not to fabricate. Any one of them may work today. None will work reliably across model releases, and this course is meant to outlive the model it was written against.
The rule in practice: the guaranteed teaching event is always a supplied artifact; the live run is always second and its outcome is always interesting. Learner examines a captured, dated real output that exhibits the failure and works out what's wrong with it — that part cannot break. Then they run it themselves. If today's Claude reproduces the failure, they've confirmed it. If it does better, the debrief has an answer ready: this is a system that improves, which is exactly why you check rather than assume. Either way the objective is met.
This is the premise with the largest production consequence, and it is the reason the review's ordering should be adjusted: it belongs at P0, not P1. The synthetic material set is the longest-lead deliverable in the whole build, and this rule changes what has to be in it — every affected exercise needs a captured "before" artifact that doesn't exist yet. Build the materials first and apply the rule second, and you build them twice.
Premise 4 stated as a check a scriptwriter can actually run. Every exercise in the course is labelled below, and anything that can't be labelled A or B doesn't ship.
| Tier | What it is | Why it holds |
|---|---|---|
| A — cannot break | Analysis of a supplied artifact | The teaching event already happened and is captured. No live system involved. |
| B — robust live | Learner constructs something and inspects it, with no required outcome | Tests construction and inspection, not prediction. Works whatever the model does. |
| C — forbidden | Objective depends on a specific model behaviour | Ships a silent time bomb. Rewrite as A, or cut. |
The companion copy rule, and it applies to every word of learner-facing text: never assert what Claude will do — say what to look for and why it matters. "Claude will miscount the letters" is a Tier C sentence even inside a Tier A exercise. "Check whether the count is right, and note what you find" is not.
What the tier test is not: an argument for dropping live work. A quiz-only course delivers objective 1 and abandons 2, 3 and 4 — you can test whether someone can classify a broken prompt, but not whether they can repair one, and confidence with Projects is not obtainable by reading about Projects. If the course ever does go assessment-only, the objectives should be rewritten down to match rather than quietly under-delivered.
The conceptual spine, in the designer's vocabulary rather than the learner-facing titles. Renaming Module 5 from verification to evidence is not cosmetic: it shifts the module from "catch the robot lying" to "know what any claim rests on and how much being wrong costs," which is a professional literacy skill that stays useful however accurate models get.
One reorder is defensible: moving M6–M7 earlier for a quick win. I'd still resist it. Learners who meet Projects before they understand how information reaches a response use them as folders, and that misuse is hard to undo. If drop-off data says otherwise, the cheaper fix is a five-minute unassessed tour preface, not a resequence.
describe what Claude is doing between pressing enter and seeing text, and predict from that description which kinds of task need extra care.
Claude produces its answer sequentially, one chunk of text at a time, each shaped by everything in front of it. Three consequences do most of the explanatory work for the rest of the course:
Plus the small one that explains a daily annoyance: text is handled in chunks, not letters, which is why character-level tasks are shaky. Name them once and move on — these chunks are what everyone else calls tokens, and they're the unit usage limits are counted in, so the word is worth recognising even though the course won't use it again. And output is sampled, not fixed — the same prompt twice can give two different answers, which sets up the reroll trap in Module 4.
"It's looking things up." The database metaphor is the most damaging wrong model a casual user carries, and the parent of every fabrication surprise later in the course. Nearly all of them hold it implicitly.
The move, under premise 4: lead with a supplied captured Claude response — dated, archived, web search off — asking for sources on an obscure topic, complete with perfectly formatted citations, plausible author names and real-sounding journals that do not exist. The learner's task is to work out what's wrong and how they'd check. That teaching event is guaranteed because it's a static artifact.
Then they run it live. If today's Claude fabricates too, they've confirmed it first-hand. If it declines, or hedges, or searches instead, the debrief says: good — this is a system that improves, which is exactly why the skill is checking rather than assuming. The original design staged this live and would have broken silently the first time the model got better at it.
Each item is a supplied captured artifact first, then an optional live re-run. Predictions recorded before either.
Every live re-run carries the same instruction: record what actually happened, including "it worked fine." That framing is what makes the exercise robust instead of fragile.
Predict-the-failure: a task is described, learner selects needs care / straightforward and then the reason from four mechanism-based options. Graded on the reason — a learner right for the wrong reason will fail Module 4. One item must separate never knew it from knew nothing after that date, since both produce identical-looking confident wrongness.
say what information is available to Claude for the response it's about to give, name where each piece came from, and take deliberate action when a conversation stops being reliable.
Claude can only reason over information made available to the current response. Some of that comes directly from the conversation in front of it. Some is brought in by other machinery: memory, project knowledge, chat search, uploaded files, connectors, web search.
This replaces rev. 2's claim that "anything outside the context window does not exist," which was already false at the time of writing. Claude can search previous chats, memory injects persistent entries, projects retrieve from project knowledge, and long conversations can be automatically compacted — earlier messages summarised so the conversation continues, rather than simply falling off the end. Teaching the absolutist version gives learners a model that Module 6 then contradicts, and being contradicted by a later module is how a course loses trust. The reframe is also a better foundation for M6, which becomes "here is the machinery that makes information available" rather than a rescue from a limitation.
The two things the learner should be able to do:
And the honest version of the old claim: long conversations can become less reliable or less focused. Not monotonically, and automatic context management makes any stronger claim brittle. Teach the recognisable symptom — dropped constraints, drift, self-contradiction — and the cheap repair, which is a fresh chat with a restated brief.
Two, pulling in opposite directions.
"But I already told it that." Learners conflate having said something with it being reliably available now, and read the failure as carelessness rather than as a question of what actually reached this response.
"More is better, so I'll paste everything." The reasonable inference from the first half of the lesson, and it produces worse output.
The move: a supplied transcript — a long, realistic, drifting conversation where a constraint stated early is quietly violated late. The learner finds where it went and diagnoses why. Guaranteed to work, because it already happened. Then the live half is the contrast, not the failure: same question in a fresh chat with just the brief, compared against the tail end of the supplied transcript. Producing a good answer live is easy to reproduce; producing a bad one on cue is not.
Supplied: Northwind Consulting — Q3 Board Memo, 12 fictional pages, with a constraint stated on page 2 ("no recommendation may assume additional headcount"), plus a captured 20-turn transcript in which that constraint is eventually violated.
"Where would this come from?" — a scenario is described (a file uploaded three chats ago, a project instruction, something mentioned yesterday, a document still on the learner's desktop) and the learner identifies whether it could reach the current response and by what route. Directly pre-loads M6's routing quiz, and the "still on your desktop" item is the one that separates understanding from guessing.
build a prompt from components rather than copying a template, and explain what each component is doing to the output.
A prompt is a specification, and vagueness in the spec gets filled in with whatever's most typical. Six components:
The unifying explanation, from M1: an under-specified prompt leaves Claude to pick something typical, and typical averages out to generic. Every component above narrows the target.
"There are magic words." Casual users collect incantations — "You are a world-class expert…", elaborate politeness, ALL CAPS, offering tips — and believe the phrasing itself is the lever. This is what makes learners template-dependent and unable to build a prompt for a problem nobody has written a template for. It's the direct obstacle to the third course objective.
The move: a controlled A/B in the exercise. Same task, one prompt loaded with flattery and emphasis, one plainly worded but specific about audience, length and exclusions. Then the reframe: prompting is specification, not spellcasting. Reinforce by having the learner strip a supplied incantation-heavy prompt to its actual content — usually about a third survives and the output holds up.
Premise-4 note: this A/B is safe to run live. It doesn't need Claude to fail — it needs specificity to help, which is stable across model versions and the whole point of the module. Supply a captured comparison anyway as the fallback for a learner whose two outputs come out equally good.
Supplied: a messy fictional customer complaint thread (Meridian Software, billing dispute, three participants, one of them wrong about the facts). Learner writes a reply four times:
Scored each round against a fixed supplied rubric, so each component's marginal contribution is a number rather than a vibe. Then the incantation A/B.
Role variants: the same thread as an internal escalation (manager), a post-incident note (ops), or a churn-risk summary (CS lead). Shared rubric across all three.
Component labelling: given a strong prompt, tag each part. Rewrite selection: given a weak prompt and four rewrites, pick the one that adds specification — distractors are longer, more polite and more emphatic but no more specified. That item type is the misconception, tested directly.
look at a bad output, name which of six failure classes it belongs to, and apply the matching repair — instead of rerolling, rewriting at random, or giving up.
The keystone module. Bad output is not one thing and the repairs are not interchangeable.
| Failure | Looks like | Repair | From |
|---|---|---|---|
| Missing information | Confidently answers a different situation than yours | Supply what it couldn't have | M2 |
| Under-specified | Valid, competent, not what you wanted | Add constraints or an example | M3 |
| Fabrication | Specifics that are wrong — numbers, names, citations | Ground it, or scope it to a source | M1 |
| Wrong shape | Right content, wrong length/format/register | Specify format, or show one | M3 |
| Lost the thread | Contradicts itself, drops earlier instructions | Fresh chat, restated brief | M2 |
| Insufficiently decomposed | Does all of it, shallowly | Break into steps; give it room to work | M1 |
Renamed from "task too big" on review, and the reviewer was right. "Too big" names a fixed capability boundary while the repair column immediately says it isn't one — the label was arguing against its own row. "Insufficiently decomposed" points at the fix.
Three habits taught alongside the taxonomy:
"It can't do this." Learners attribute a capability ceiling when the cause is nearly always their own specification or missing information. They stop one move too early, and that single habit separates a casual user from a capable one.
The move: the clinic's answer key is the argument. Five of six items are fixable, several with a one-line change, and the learner does the fixing.
The sixth is genuinely unfixable — a fact that is private and in no supplied source — and its key says so plainly. This should survive review. A diagnostic that can never return "stop, wrong tool" is not a diagnostic; it's a motivational poster. Learners need calibrated persistence, not unlimited persistence.
Note the unfixable item is now a private-fact item, not character counting. Under premise 4, "Claude can't count characters" is exactly the kind of claim that may stop being true; "Claude cannot know your colleague's unshared salary figure" will not.
Six supplied prompt-and-output pairs — captured, dated, shuffled, one per failure class. For each: classify, state the repair, then apply it live and compare against the supplied model repair. Applying is mandatory; classification alone produces learners who pass the quiz and can't fix their own work.
Wholly supplied inputs, so premise 4 is satisfied by construction — this exercise was already robust and needed no redesign.
Closing drill: the learner brings their own worst recent Claude interaction and runs the taxonomy on it. First time they point the method at their own material.
Vignette classification: prompt plus bad output, learner picks failure class and repair as separate scored fields. Distractors are the adjacent classes — under-specified and wrong-shape are the pair people genuinely confuse. At least one item's correct answer is "not a prompting problem."
Ship the taxonomy as a permanent reference card, alongside the capstone's seven-step process. On review these are the two most reusable objects in the course, and both should outlive the modules that introduce them.
for any claim in an output, say what it rests on, how to inspect that basis, and how much it matters if it's wrong.
Reframed on review from verification to evidence, and the reframe is the improvement. The old module was organised around catching the model lying, which dates badly as models improve. This one is organised around a question that stays useful regardless: what is this claim based on?
Two independent questions, which rev. 2 collapsed into one and shouldn't have:
| A · Can I trace it? | B · Is the source any good? |
|---|---|
| A source you supplied | Authority — who produced it, and do they know? |
| A source Claude retrieved (search, connector, project knowledge) | Relevance — is it about your actual question? |
| Unsupported model knowledge | Recency — is it current enough to act on? |
| Completeness — is it the whole picture or one slice? |
Column A is a traceability ordering, not an accuracy ordering, and rev. 2 wrongly implied otherwise. A document you uploaded can be outdated, wrong, incomplete, adversarial, or simply the wrong document — and a retrieved primary source from a statistical agency may be far more trustworthy than a PDF someone emailed you. The two questions are orthogonal and both have to be asked.
Which yields the sentence the module exists to install:
Grounded does not mean true. Grounded means you have something concrete to inspect.
Then triage, because effort has to be proportional: risk is roughly cost of being wrong × how traceable the claim is. The reliably high-attention zones, taught as a checkable list: specific numbers, citations and URLs, direct quotes, names and dates, anything legal/medical/financial, recent events, and anything about the learner's own organisation.
And the behavioural rule, stated as behaviour rather than as an empirical absolute: do not use confident tone as evidence of accuracy. Claude writes an unsupported claim in the same register as a well-supported one. Most people are unconsciously reading prose style as a confidence signal, and the fix is a habit, not a belief about what tone does or doesn't correlate with.
A matched pair, and the module must hit both or it pushes learners from one into the other.
"I'll just tell it not to make things up." Instruction doesn't reach the mechanism. Taught from a supplied captured example — an ungrounded question with "only state facts you are certain of" attached, and an unsupported claim in the answer anyway, complete with a confidence disclaimer. Supplied rather than live, per premise 4: this demo previously required the model to ignore an instruction on cue.
"It makes things up, so I have to check everything." The overcorrection. Produces learners who verify a tone adjustment as carefully as a revenue figure, find the tool nets out slower than doing it themselves, and quit. Quieter than the first failure and probably more common among cautious professionals.
The move: triage as the organising idea, with a worked procedure. The learner should leave with a specific answer to "what do I check on a busy Tuesday" — the numbers, the names, the citations, and not much else when the rest traces to a document you supplied and trust.
Supplied: a fictional Harbourview Market Assessment (six pages) and a summary of it containing three factual errors and two unsupported inferences — claims that sound sourced but appear nowhere in the document. Fully supplied, so it was already premise-4 safe.
Two-dimensional sorting: for six sentences, learner marks both traceable to what and act on it / check it first. Correct answers include several "act on it" — a learner who flags everything scores badly, which is the calibration being tested. Two items pit a traceable claim from a weak source against an untraceable claim that happens to be plausible, which is where the two dimensions come apart.
choose the right scope for a piece of persistent context, and set up a Project that makes a recurring task meaningfully faster.
Calls back to M2: information has to be made available to a response, and typing it every time doesn't scale. These features do it for you. Teach them as scopes, not as boxes — which is a rev. 3 correction, because the tidy three-bucket taxonomy is no longer accurate:
Why "scopes" and not "boxes." Memory is not a single account-wide store: each project has its own separate memory space and its own project summary, kept apart from other projects and from non-project chats. So "Project" and "Memory" are not parallel alternatives — memory has a project scope. Chat search adds another retrieval route with its own boundaries (all chats outside projects; within a project, that project only). The honest teaching is a single question — how widely should this apply? — answered along a scope gradient, rather than three tidy containers a learner will later find leaking into each other.
The retained decision rule, reworded for scope: true about me everywhere → Instructions. True about this body of work → that Project. Happy for it to accumulate on its own → Memory, at whichever scope you're working in.
"I told Claude that already — why doesn't it know?" Learners don't distinguish said in a conversation from stored at a scope. Everything gets filed as "I told it," so the behaviour looks arbitrary — sometimes it remembers, sometimes not, no discernible rule. This is the confusion that makes people quietly distrust the tool.
The move: make the invisible visible. Have them open Settings → Memory and read their own entries. Nothing else lands as hard as seeing discrete, editable, listed statements where they'd imagined a vague cloud of recollection.
Runner-up — don't skip it: quiet unease about memory. A learner who is uncomfortable and can't name why simply won't use the feature. Address it factually and early — pause, reset, export, incognito, per-project separation, entries you can read — rather than reassuringly. The controls are the reassurance.
Supplied: a fictional Lumen & Co. brand voice guide and three past newsletters, for a recurring "draft this month's newsletter section" task.
Step 5 previously required read, edit and delete. Cut on review, and correctly: requiring a learner to destroy their own real memory entry is needlessly invasive for an educational exercise, and inspection alone proves the concept. Editing and deletion move to an optional extension using a deliberately disposable entry the learner creates first.
Role variants: newsletter section, weekly status roll-up, or candidate screening notes.
Scope routing: eight facts ("I always want metric units"; "our fiscal year starts in April"; "this client hates bullet points") placed at the right scope. Deliberate near-misses, since the everywhere-versus-this-project line is where the real judgement lives. One item should have a defensible answer at two scopes, with feedback that explains the trade rather than declaring a winner — that's where scope-thinking beats box-thinking.
get source material into Claude, judge how completely it was actually read, produce something in an Artifact that survives outside the conversation, and set how hard Claude works on a task.
Clicking the model name opens three separate controls, and casual users almost never touch any of them. This pays off Module 4's seventh repair and closes the loop opened in Module 1.
Deliberately deferred from Module 4. A learner who reaches for max effort before checking whether their prompt was under-specified has swapped a diagnostic habit for a dial — the exact behaviour Module 4 exists to prevent.
"I uploaded it, so it's read the whole thing." Learners treat a successful upload as a guarantee of complete, accurate reading, then trust a page-40 claim as much as a page-1 claim. Genuinely dangerous, because it manufactures confidence at exactly the moment M5 taught them traceable sources are better.
The move: stop relying on a live degradation demo, and stop reaching for specifics too. Teach the principle:
"Uploaded" is not a single file-reading mode. How completely a file gets read depends on its format and its size, and you cannot tell which you got by looking at the reply.
A PDF, a spreadsheet and a scanned page are not handled the same way, and a long file and a short one are not handled the same way. The learner does not need the mechanism, and giving them a number would be worse than giving them nothing — see the note below. What they need is the conclusion: if it matters, demand the supporting line.
Why no thresholds here. An earlier revision taught specific page counts from Anthropic's platform documentation. Cut, for two reasons. First, those describe API behaviour and were never verified against the consumer path this course teaches — an unverified number is exactly the kind of claim Module 5 tells learners not to accept. Second, and more durably: a named threshold invites learners to treat it as a reliable operating parameter, so it fails twice over when it changes — the number is wrong, and the habit it built was wrong. "Format and size both matter, and you can't tell by looking" survives every future change, and generalises to formats and limits nobody has documented yet.
The live half becomes a check rather than a gotcha: ask a question whose answer sits deep in a dense table, demand the exact supporting line, verify it against the source yourself. "Quote the line" is the durable habit either way.
Supplied: a deliberately untidy fictional spreadsheet (inconsistent date formats, merged cells, a stray total row) plus a short covering memo.
Timing, flagged on review and it's a fair hit. Forty minutes for six live tasks involving an upload, an Artifact, two iterations, a settings comparison, publishing and a branch was optimistic — it ignored upload time, response latency and reading. The honest options were: split into two modules (breaks the 6–8 module brief), extend past the 40-minute ceiling (breaks the sitting constraint), or cut. I cut: branching is now optional, which is defensible because it's the one item here that isn't load-bearing for the capstone. 40 minutes is still the tightest module in the course and should be the first thing playtested.
Connectors remain a parallel path, never a gate. A meaningful share of learners won't connect a personal Gmail for a course exercise, and that's a reasonable position. The file-upload route satisfies the same objective and both count. Progression must never depend on a learner granting a permission they'd rather not grant.
Quiz on how each route makes information available. The file items must test the conclusion — given a scenario, does this claim need the supporting line demanded? — and never a threshold or a processing rule, which is where a quiz bank rots fastest and hardest, because a stale item marks a correct answer wrong. Two dial items, scenario-based rather than definitional: given a task, is this a prompt problem or a settings problem? "Prompt" is correct more often, which protects the Module 4 reflex.
Plus a binary self-verified checklist: published link opens in a private window ✓, two in-place iterations ✓. Binary and self-evident — the learner can't half-convince themselves they did it.
start from a real work problem, decide honestly whether Claude fits it, and if so build, test and document a working setup — with no template supplied at any point.
The review caught a genuine blocker and it was the sharpest catch in the pass. Rev. 2 held both "synthetic only — no exercise may require the learner's real work data" and a capstone that said pick a real recurring task, load its knowledge, test it on real cases, and publish the result to a public link. Those cannot coexist. Publishing on Free/Pro/Max makes an Artifact reachable by anyone with the link.
Two changes. First: Publish is removed from the capstone entirely. The review is right that it contributes nothing to the capstone's objectives — the learner already demonstrated publishing in Module 7 — while creating a real disclosure risk at exactly the moment they're working with material closest to their actual job. The design record is submitted, not published.
Second, and here I've landed somewhere different from the review. The review proposed learners build against a supplied synthetic analogue instead of their own material. I think that over-corrects and costs the capstone its best property — motivation and transfer come precisely from building something the learner actually wants on Monday. Read the constraint literally: no exercise may require real work data. It does not say real data is forbidden; it says the course can never make it the price of admission.
So: task selection stays real, data is at the learner's discretion, and a synthetic analogue is supplied for every capstone shape so anyone who can't or won't use their own material has a complete path with nothing withheld. The graded artifact is the design record, which describes a task's structure — inputs, outputs, what good looks like, how you'd know it failed — and needs no real content to be excellent. That satisfies the constraint as written, keeps the capstone worth doing, and removes the disclosure risk that actually mattered.
Pick a task you personally did more than three times in the last month and did not enjoy. Do not pick a task because it sounds like a good AI task. Pick it because it is tedious, and then find out whether AI helps. Describe its structure in full. Use your own material only if you're free to — otherwise take the supplied synthetic analogue, which is built to the same shape and costs you nothing.
The second sentence is the whole capstone. Left off, roughly everyone submits summarisation or email drafting and the course has assessed template-matching under another name. Constraining the selection rather than the solution is what forces problem-first reasoning.
This process — not any individual module — is the durable artifact of the course. Ship it as a reference card, alongside Module 4's failure taxonomy.
The working setup, plus a one-page design record covering all seven steps. Produced as an Artifact — not published. Submitted or self-graded privately.
Not a misconception — a completion risk. Highest effort, lowest structure, arriving when novelty has worn off. This is where an async cohort haemorrhages learners.
Mitigations: four separately-saveable checkpoints so it survives several sittings; three fully-worked example design records from different roles — including one that legitimately concludes "no fit" — so the blank page is never truly blank; and step 1 moved to the end of Module 7 as a five-minute pre-task, so learners arrive already holding a candidate task instead of facing selection and design at once.
Six to eight items, 80% to pass, unlimited retries from a shuffled pool so retrying isn't memorising. The critical rule: wrong answers return a mechanism-level explanation, not "incorrect." With no instructor, the quiz is the only thing that can catch a misconception at the moment it's exposed, so it must teach rather than merely measure. Budget four distractor explanations per item — the most underestimated cost in the build.
Binary, self-evident, hard to fool yourself on. "Your project instructions contain an output-format constraint — yes or no." These carry more weight than any confidence rating, because the learner checks a fact about an object rather than rating a feeling. Never ask "do you feel confident with Projects?" — the answer is uncorrelated with anything.
The learner pastes their design record into a supplied grading prompt that scores against published criteria and returns specific gaps. Defensible rather than a cop-out because the criteria are public and auditable, and because by Module 5 the learner has been taught not to take Claude's judgement on trust. The instruction should say so outright — challenge this grade; if you think it's wrong, argue with it and see if it holds. The assessment becomes one more repetition of the skill. Note this also works on a structure-only design record, so it never requires real content.
Two short cumulative checks, drawing across all modules and deliberately interleaved — a diagnosis item, then a scope item, then an evidence item — rather than blocked by module. Blocked practice inflates in-the-moment performance and does much less for retention. The only element targeting durable recall rather than course-completion, and the cheapest thing here to build.
Day-7 and day-30 checks are reported but deliberately don't gate the certificate. Gating them creates an incentive to defer completion, and the retention value doesn't depend on compulsion.
Settings → Memory, the project knowledge panel,
the Publish button, the model-name menu), because objective 4 is unreachable without it.
That's roughly half a dozen interface locations across the whole course — small enough to be a
pre-cohort checklist rather than a versioned specifics file with an owner. Keep screenshots in
one replaceable asset set with capture dates, and prefer the written path in body copy so the
text survives a UI change even when the image doesn't.