Inside a pretrained language model

Year 12 · Student investigation

Can restoring an internal state recover a language model’s answer preference?

Keep an individual record even when you investigate in pairs. Record the exact settings, source and data split.

Predict before seeing the answer

Which final-block patch should recover the clean score?

Two comparisons

What did you change? What stayed the same? What did you notice?

Challenge your first explanation

Compare six prompt pairs, layers and positions against no-op, zero and norm-matched random controls.

Make and explain

A reproducible patching report naming checkpoint, metric and scope.

Limit the claim

This activity can show _ . It cannot tell us _ . A new case I would test is ____ .

Use fictional data. Download or print your work before leaving.