Year 6 · 45 minutes
Pip and the library that guessed
How can we help a model check, correct and learn from made-up answers?
You will learn to: Distinguish invented stories from unsupported factual claims; choose checking, correction or uncertainty; explain how feedback can improve future behaviour without guaranteeing truth.
A story about evidence
Pip’s very confident announcement
Pip, a young wallaby, helped at the bush library. Its new answer machine finished sentences beautifully. Pip asked it to announce the science fair.
“Bo won last year!” it said. Pip nearly read the announcement aloud. Wattle the wombat stopped him. “Which page tells us that?”
Making up a rocket adventure is fine when everyone knows it is fiction. Making up a winner in a factual announcement is a mistake.
Claim 1 of 4
The fair opens at 10 am.
What should Pip do before reading this aloud?
Fixing today’s answer. Learning for tomorrow.
A checker is like a library helper who raises a flag when something looks doubtful. It can miss a mistake or raise a false alarm. In the research, signals from inside a frozen model help score answers during training. The answering model learns from that feedback.
This announcement changes. Pip edits the text using evidence. The model’s stored numbers have not changed. The same mistake could appear tomorrow.
Try the real small-model experiment in Marimo: compare known colour facts, a learned checker and a policy that can answer or say “I need to check”.
Now investigate for yourself
Your Marimo research notebook
The story is an original analogy. The notebook uses a tiny colour-fact model, a learned checker and a small reward-trained output policy. It does not reproduce Gemma-3-12B, long biographies or Goodfire’s full RLFR pipeline.
Teacher notes & evidence task
Starting knowledge: Read a short evidence card and compare it with a claim. Spoken or drawn explanations are welcome.
- Read Pip’s story. Mark which details the library card actually supports.
- Choose whether Pip should keep, check, correct or withdraw each claim. Explain your choice using the card.
- Separate fixing this answer from changing future answers through training.
- In Marimo, observe errors from a small trained model. Train a fallible checker and compare it with the known fictional facts.
- Use the checker’s feedback to update a small answer policy. Measure both wrong answers and how often it answers.
Evidence to collect: Explain one correction, one appropriate ‘I need to check’, and why a high checker score is not proof.
Research basis: RLFR uses probes of a frozen model as reward signals. Its pipeline detects possible factual errors, proposes corrections or retractions, and trains from feedback. A monitor is fallible, not a truth machine.
Goodfire research article ↗ · Full research paper ↗ · Setup and teaching guidance