Year 6 · Live trained model · 35 minutes

Kiki’s switch test

Which part changes the answer when we switch it off?

Open full-screen notebook ↗Download Python notebook ↓Teacher notes & worked example ↗Paper student journal ↗Research source map →

The first load downloads Python and may take a moment. Write a prediction, then open the lab. Controls, help, teacher notes and your evidence download are inside. If the school network blocks the runtime, download the Python notebook and run it with your teacher.

Investigate one research connection

Feature Steering for Reliable and Expressive AI Engineering

Turn one internal switch and look for side effects. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Steering is not guaranteed control or a source of new knowledge. Ember links are historical.

Use the notebook below to collect the measurements. This writing remains in this page session until downloaded.

Teacher background and source method

Historical Llama steering demonstrations

Changes selected internal feature strengths and inspects example outputs; demonstrations need independent task and side-effect checks.

Independent classroom adaptation; not a reproduction of the source model or complete method.

Read the source with a teacher ↗

Opening the notebook page…

First use downloads Python and its libraries. A fresh session measured about 20–22 seconds and 14–17 MB during the initial audit; school networks vary. This hosted notebook computes on the browser CPU.

Your prediction and saved runs stay in this session. Download your evidence before leaving or refreshing. The hosted experiment runs on your browser CPU and does not connect to an H100.

Research connections · 11 archive entries

These lessons adapt ideas and methods. The original sources state their own model, data and validation scope.

Feature Steering for Reliable and Expressive AI Engineering
Steering is not guaranteed control or a source of new knowledge. Ember links are historical.

On Optimism for Interpretability
The author explicitly presents optimism amid unsolved scientific questions.

Interpreting Language Model Parameters
Our SVD exercise is a contrast, not VPD. The paper studies a 67M-parameter model, not all frontier models.

Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
Our tiny binding transformer is independently trained and does not reproduce the paper's nine-model circuit findings.

Open Problems in Mechanistic Interpretability
A research roadmap is not evidence that these problems have been solved.

Painting With Concepts Using Diffusion Model Latents
This is image generation, not an LLM experiment; the classroom intervention is a transferable analogy.

Replicating Circuit Tracing for a Simple Known Mechanism
Replication found both similarities and differences. Our attention patching does not reproduce CLT circuit discovery.

Towards Scalable Parameter Decomposition
SPD evidence here is based on controlled models; the SVD notebook is a baseline contrast, not an SPD implementation.

The Circuits Research Landscape: Results and Perspectives
The linked community guide was reviewed. Attribution graphs provide hypotheses with missing or uninterpretable components; interventions remain necessary.

Understanding Memorization via Loss Curvature
The source uses curvature approximations including K-FAC; our exact quadratic toy is not a K-FAC replication. Edits have collateral costs.

Paper Summary: Interpreting Language Model Parameters
This explains the same work as Interpreting Language Model Parameters; it is not a second independent result.