Year 6 · Live trained model · 35 minutes
Wattle’s clue microscope
Can one number tell us what a model is thinking?
The first load downloads Python and may take a moment. Write a prediction, then open the lab. Controls, help, teacher notes and your evidence download are inside. If the school network blocks the runtime, download the Python notebook and run it with your teacher.
Investigate one research connection
Announcing Our $50M Series A to Advance AI Interpretability Research
Separate a promise about a tool from evidence that it works. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.
Funding is not validation. The older Ember service is deprecated.
Use the notebook below to collect the measurements. This writing remains in this page session until downloaded.
Teacher background and source method
Company funding and strategy
Reports financing and plans; investment is not a model evaluation.
Independent classroom adaptation; not a reproduction of the source model or complete method.
Read the source with a teacher ↗Opening the notebook page…
First use downloads Python and its libraries. A fresh session measured about 20–22 seconds and 14–17 MB during the initial audit; school networks vary. This hosted notebook computes on the browser CPU.
Your prediction and saved runs stay in this session. Download your evidence before leaving or refreshing. The hosted experiment runs on your browser CPU and does not connect to an H100.
Research connections · 15 archive entries
These lessons adapt ideas and methods. The original sources state their own model, data and validation scope.
Announcing Our $50M Series A to Advance AI Interpretability Research ↗
Funding is not validation. The older Ember service is deprecated.
Intentionally Designing the Future of AI ↗
This is a research agenda; intentional design is not a solved capability.
Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Model ↗
A fast pipeline can still collect misaligned evidence. Classroom timings do not benchmark frontier hardware.
Understanding, Learning From, and Designing AI: Our Series B ↗
A company announcement provides strategy and context, not an independent benchmark.
Announcing Open-Source SAEs for Llama 3.3 70B and Llama 3.1 8B ↗
Open weights enable investigation; they do not certify feature labels or universal coverage.
Adversarial Examples Are Not Bugs, They Are Superposition ↗
Bidirectional causal evidence is strongest in toy models; real vision-model evidence is narrower, not a universal explanation.
Can SAEs Capture Neural Geometry? ↗
Low reconstruction error alone does not establish semantic completeness; our small SAE is not the paper's full evaluation.
Covariance-based Sequence Pooling ↗
The source method uses second moments and approximations; our centred-covariance experiment is a teaching analogue. It does not recover sequence order.
Interpreting Evo 2: Arc Institute's Next-Generation Genomic Foundation Model ↗
DNA models are not ordinary text LLMs. The 2025 report was updated to note Nature publication in March 2026.
Mapping the Latent Space of Llama 3.3 70B ↗
A two-dimensional map is a lossy view. The older Ember demo is deprecated.
Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train ↗
Our supervised colour task illustrates data effects; it is not a replication of contrastive-SAE post-training or DPO.
Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training ↗
The paper evaluates particular models and behaviours. Our checkpoint comparison does not implement its attribution algorithm.
Understanding Sparse Autoencoder Scaling in the Presence of Feature Manifolds ↗
The paper identifies regimes and mechanisms, not a claim that all larger SAEs get worse.
Understanding and Steering Llama 3 with Sparse Autoencoders ↗
Features can overlap or duplicate. Ember API and demo links are deprecated; these labs need neither.
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention ↗
The reported mechanism is studied in particular toy and language models; size is not a guarantee on every task.