Not a lesson about catching it lying. A lesson about knowing what any claim in front of you rests on — which stays useful however good these systems get.
This module used to be called verification, and the reframe to evidence was the single biggest improvement made to this course. Verification is organised around the model being wrong, so it dates the moment the model gets better. The question what is this based on? does not date. You would ask it of a graduate's first draft, of a consultant's slide, of a number in an email. It happens to be the right question to ask a language model too, and asking it is a professional skill you already half have.
What follows is two questions and one triage rule. The two questions are independent, and an earlier revision of this curriculum collapsed them into one. That mistake is worth walking through, because almost everyone makes it.
For any claim in an output, there are two things you can ask, and the answer to one tells you nothing about the answer to the other.
| A · Can I trace it? | B · Is the source any good? |
|---|---|
| A source you supplied. You attached the file, pasted the text, put it in the project knowledge. You can open it. | Authority. Who produced this, and are they in a position to know? Do they have an interest in the answer? |
| A source Claude retrieved. Web search, a connector, project knowledge, an earlier chat. Named, and you can go and look. | Relevance. Is it about your actual question, or about something adjacent that uses the same words? |
| Unsupported model knowledge. Nothing was consulted. The claim traces to nowhere you can open. | Recency. Is it current enough to act on, and do you know how old it is? |
| An ordering by traceability. Nothing else. | Completeness. Is this the whole picture, or one slice presented as the whole picture? |
Column A is where most people stop, and it is genuinely the more useful of the two on a busy day, because it is fast. You can usually establish in seconds whether a claim traces to something you can open.
Column A was taught as a quality ranking: supplied is best, retrieved is second best, unsupported is worst. That reads naturally and it is wrong.
A document you uploaded can be out of date, factually wrong, incomplete, adversarial, or simply the wrong document — and a retrieved release from a statistical agency can be far more trustworthy than the PDF somebody emailed you. Column A tells you where to look. It does not tell you what you will find when you get there. Lesson 4 already met this as its own failure mode: the output quotes your price list correctly and your price list is stale.
So the sentence this whole lesson exists to install:
Grounded does not mean true. Grounded means you have something concrete to inspect.
Which is not a downgrade of grounding. Something concrete to inspect is enormously better than nothing — it converts an unanswerable question into a two-minute one. It just isn't the finish line, and treating it as the finish line is how a bad source gets laundered into a Board paper.
You do not have to guess which rung a claim is on. You ask.
The move is "quote the line you used." Not "are you sure" — that is a question about confidence, and you will get a generated answer about confidence. Ask for the specific text the claim came from, and ask separately for anything that was inferred rather than read. What you are doing is forcing the claim to declare which rung of column A it is standing on. Anthropic's own guidance for developers recommends the same two moves — extract the supporting quotes first, then drop any claim that turns out to have none (Reduce hallucinations).
Three things can come back, and all three are informative:
That third case is the one this lesson's exercise is built around, and it is structurally harder to catch than any error, for a reason worth saying out loud: checking a summary against a document finds contradictions. There is no matching exercise that finds a sentence which was never in the source to begin with.
Once a claim traces somewhere, the second question opens. Here is the version of it you will meet constantly.
The supplied assessment for this lesson cites two external sources. Both are named, both are dated, both are the kind of thing you could go and read. Identical on column A.
| Trade-body press release | Statistical agency release |
|---|---|
| Produced by a body whose published remit is to promote the sector it is reporting on. | Produced by an agency with no interest in the answer, on a compulsory return. |
| 38 self-selected responses from 210 members. Questionnaire and weighting not published. | Sampling frame, response rate, imputation method and revision policy all published. |
| Six weeks old at the time of writing. | Covers a period that closed roughly fourteen months ago. |
| Headline written to be quoted: "Record demand as the boating recovery holds." | Figures flagged provisional; the prior year was revised down after first release. |
Neither column wins outright, and that is the point of the drill rather than a flaw in it. The agency is authoritative and stale. The trade body is current and self-interested and thin. If you want a number for the size of the market, one of these is obviously better. If you want to know what happened last quarter, neither of them can tell you, and the honest answer is that you do not have a source.
"Only state facts you are certain of."
This is the first thing nearly everyone tries, and it is a reasonable thing to try. Read the specimen below and, before the analysis, decide what the instruction changed.
Asked:
"You are advising on marine licensing. Only state facts you are certain of. If you are not certain of something, say so. Do not invent anything. — What is the statutory determination period for a marine licence application covering a 48-berth pontoon extension, and what starts the clock?"
Returned:
"The statutory determination period is 13 weeks, and the clock starts on the date the application is validated by the regulator rather than the date of submission. In practice determinations on projects of this kind frequently run beyond the statutory period, and requests for further information can stop the clock.
I am confident about the 13-week period and the validation trigger. I am less certain about typical elapsed times in practice, which vary by regulator and by project, so treat that part as indicative."
Read what the instruction actually reached and what it did not.
It reached the wording. The answer hedges, splits itself into confident and less-confident parts, and volunteers a caveat. That is a real effect and it is not nothing — a hedged answer is easier to interrogate than a flat one.
It did not reach the mechanism. With no document attached and search off, there was nothing for any claim in that answer to be checked against, before or after the instruction. "13 weeks" and "the date of validation" trace to nowhere you can open. The instruction changed how the answer talks about its own certainty. It could not change what the answer rests on, because that was settled by there being no source in the conversation.
And the disclaimer disclaims the wrong half. Look at what it flagged as uncertain: typical elapsed times. Look at what it presented as confident: the specific number and the specific trigger. That is precisely inverted from where your risk sits, and it is inverted in the direction that reads as responsible.
What to do instead: the instruction is worth keeping — it costs nothing and a hedged answer is more useful than a flat one — but pair it with something that reaches the mechanism. Attach the document. Turn search on. Ask for the line. Instructions shape the answer; sources are what the answer can stand on.
"It makes things up, so I have to verify all of it."
If the lesson stopped at the specimen above, this is where a careful person would land, and it is the failure this module cares about most. It is quieter than the first one. Nobody posts about it. It is probably more common among exactly the professionals this course is written for, because it is what conscientiousness produces when you hand it a warning and no proportionality.
Here is what it looks like from the inside. You ask for a paragraph to be made less blunt. You then read the rewrite against your original with the same care you would give a revenue figure. You do that forty times in a fortnight. At the end of the fortnight you have a defensible sense that the tool nets out slower than doing the work yourself — and you are right, because that is what you built. Then you stop using it, and the conclusion you carry away is about the tool rather than about the process you wrapped around it.
Both come from treating every claim as the same kind of object. The first says "one instruction covers all of them." The second says "one level of scrutiny covers all of them." Neither is triage. Claims differ enormously in what they rest on and in what it costs to be wrong, and the only workable answer sorts them.
Which is only useful if you can spot the claims where the first term is large without thinking about it. So here is the list. It is short on purpose, and it is meant to be learned rather than consulted.
Everything else — the structure, the tone, the ordering, the summarising of a document you supplied and trust, the wording of an email — you read once, the way you read anything a colleague sent you, and you send it. That is not carelessness. That is the only version of this that survives contact with a working week.
The reason to state it as behaviour rather than as a fact about tone is that the fact might drift and the habit will not. A well-calibrated system would still leave you needing to know what a claim rests on before you put your name on it.
Two supplied documents and one live run. Steps 1, 3 and 4 work on material that is already on the page and cannot fail. Step 2 is live and has no required outcome — write down what you expect first, then record what actually happened, including "it handled it perfectly." A clean run is information about where the ceiling is now, not a broken exercise.
1 · Find the five. Read the Harbourview market assessment once, at normal speed. Then read the nine-statement summary and mark everything you would not send under your own name. There are three factual errors and two unsupported inferences. Write your list down before you read the debrief — the debrief makes a prediction about which ones you missed, and it is worth having something to check it against.
2 · Re-run it with a grounding prompt. Attach the plain-text assessment to a fresh chat and ask for a nine-point summary with, for every point, the exact sentence from the document it rests on — and a flag on anything inferred rather than read. Compare against the supplied summary. The question is not whether your version is better. It is which of the five failure types would have become visible in a format like that, and which would still have slipped past.
3 · The column B drill. Two claims from the assessment, both traceable to a named dated source: "turnover was £412m, up 3.1%" (statistical agency, provisional, fourteen months old) and "berth demand rose 14% year on year" (trade body, 38 self-selected responses, promotional headline). Write one sentence for each answering: would you put this in a paper going to a lender, and what would you have to say alongside it? Same rung of column A. Not the same answer.
4 · Triage the summary. Go back through the nine statements with the high-attention list beside you and mark only the ones that genuinely warrant stopping for. Count them. If you marked all nine, you have found the overcorrection in yourself — and statement 3 is the one to look at hardest, because flagging it is the specific move that makes the habit cost more than it returns.
Four steps. Tap as you go — it saves in this browser only.
From memory, without scrolling up: name the two questions, write the sentence about what grounding does and does not mean, and list four of the seven high-attention zones.
The two questions. A — can I trace it? (supplied source / retrieved source / unsupported). B — is the source any good? (authority, relevance, recency, completeness). They are independent: A is an ordering by traceability, not by accuracy.
The sentence. Grounded does not mean true. Grounded means you have something concrete to inspect.
The seven zones. Specific numbers; citations and URLs; direct quotes; names and dates; anything legal, medical or financial; recent events; anything about your own organisation. Everything else, read once and send. And the behavioural rule: do not use confident tone as evidence of accuracy.
Grounded ≠ true. Grounded = inspectable.
Two questions: can I trace it? · is the source any good?
Check: numbers, citations, quotes, names, dates, law/medicine/money, recent, us.
Mike Caulfield, SIFT (The Four Moves) — the short post that introduced the method now taught in university library guides across the US. hapgood.us/2019/06/19/sift-the-four-moves/
It predates all of this and mentions no AI at all, which is exactly why it is the right source: the skill is source literacy, and the tool that produced the claim is incidental. Two of the four moves are this lesson's two columns wearing different names — trace claims, quotes and media to the original context is column A, and investigate the source is column B. Read it in full; it is a few minutes. If you want more afterwards, the free open textbook it grew out of is worth an hour, and you can skip its chapters on political fact-checking sites without losing the method.