Year 12 · Live trained model · 55 minutes
From attention to causal evidence
Can we recover an answer by restoring one internal state?
The first load downloads Python and may take a moment. Write a prediction, then open the lab. Controls, help, teacher notes and your evidence download are inside. If the school network blocks the runtime, download the Python notebook and run it with your teacher.
Investigate one research connection
Stanford Guest Lectures: AP293 (Fall 2025)
Compare representations, causal mechanisms and models of in-context learning. Record the source model/task, a baseline, the changed factor, a measurement and an alternative explanation. State exactly which part your notebook investigates and which part it does not reproduce.
This entry reviews the article and lecture outlines, not a full transcription of the videos.
Use the notebook below to collect the measurements. This writing remains in this page session until downloaded.
Teacher background and source method
Three interpretability lecture outlines
Introduces causal abstraction, circuits and learning in context; a teaching guide rather than one dataset.
Independent classroom adaptation; not a reproduction of the source model or complete method.
Read the source with a teacher ↗Opening the notebook page…
First use downloads Python and its libraries. A fresh session measured about 20–22 seconds and 14–17 MB during the initial audit; school networks vary. This hosted notebook computes on the browser CPU.
Your prediction and saved runs stay in this session. Download your evidence before leaving or refreshing. The hosted experiment runs on your browser CPU and does not connect to an H100.
Research connections · 8 archive entries
These lessons adapt ideas and methods. The original sources state their own model, data and validation scope.
Stanford Guest Lectures: AP293 (Fall 2025) ↗
This entry reviews the article and lecture outlines, not a full transcription of the videos.
Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Model ↗
A fast pipeline can still collect misaligned evidence. Classroom timings do not benchmark frontier hardware.
On Optimism for Interpretability ↗
The author explicitly presents optimism amid unsolved scientific questions.
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context ↗
Our tiny binding transformer is independently trained and does not reproduce the paper's nine-model circuit findings.
Open Problems in Mechanistic Interpretability ↗
A research roadmap is not evidence that these problems have been solved.
Replicating Circuit Tracing for a Simple Known Mechanism ↗
Replication found both similarities and differences. Our attention patching does not reproduce CLT circuit discovery.
The Circuits Research Landscape: Results and Perspectives ↗
The linked community guide was reviewed. Attribution graphs provide hypotheses with missing or uninterpretable components; interventions remain necessary.
Under the Hood of a Reasoning Model ↗
An SAE reveals partial patterns, not a complete transcript of reasoning. Effects do not generalise automatically.