Year 12 · Live trained probe on synthetic data · 55 minutes

Probe the signal, test the shortcut

Can a highly accurate probe be reading the wrong clue?

Open full-screen notebook ↗Download Python notebook ↓Teacher notes & worked example ↗Paper student journal ↗Research source map →

The first load downloads Python and may take a moment. Write a prediction, then open the lab. Controls, help, teacher notes and your evidence download are inside. If the school network blocks the runtime, download the Python notebook and run it with your teacher.

Investigate one research connection

Covariance-based Sequence Pooling

Compare mean and centred covariance features on sequences with matched means. Record the source model/task, a baseline, the changed factor, a measurement and an alternative explanation. State exactly which part your notebook investigates and which part it does not reproduce.

The source method uses second moments and approximations; our centred-covariance experiment is a teaching analogue. It does not recover sequence order.

Use the notebook below to collect the measurements. This writing remains in this page session until downloaded.

Teacher background and source method

NTv3 gene ontology and genomic-track tasks

Compares mean pooling with compressed second-order sequence statistics using downstream probes; regularisation matters when labels are scarce.

Independent classroom adaptation; not a reproduction of the source model or complete method.

Read the source with a teacher ↗

Opening the notebook page…

First use downloads Python and its libraries. A fresh session measured about 20–22 seconds and 14–17 MB during the initial audit; school networks vary. This hosted notebook computes on the browser CPU.

Your prediction and saved runs stay in this session. Download your evidence before leaving or refreshing. The hosted experiment runs on your browser CPU and does not connect to an H100.

Research connections · 4 archive entries

These lessons adapt ideas and methods. The original sources state their own model, data and validation scope.

Covariance-based Sequence Pooling
The source method uses second moments and approximations; our centred-covariance experiment is a teaching analogue. It does not recover sequence order.

Explaining 4.2 million genetic variants with state-of-the-art, interpretable predictions
Genomic predictions are hypotheses, not clinical conclusions. The linked preprint abstract and Goodfire report were reviewed; full preprint text was unavailable.

Using Interpretability to Identify a Novel Class of Alzheimer's Biomarkers
This is a small-cohort biomedical study, not a diagnostic classroom tool or an LLM result.

Deploying Interpretability to Production with Rakuten: SAE Probes for PII Detection
Results depend on data and baselines; SAEs are not always superior. Classroom fixtures contain no real personal information.