Year 12 · Live trained SAE on synthetic data · 55 minutes
A sparse microscope with blind spots
Does a better reconstruction reveal every concept?
The first load downloads Python and may take a moment. Write a prediction, then open the lab. Controls, help, teacher notes and your evidence download are inside. If the school network blocks the runtime, download the Python notebook and run it with your teacher.
Investigate one research connection
Announcing Open-Source SAEs for Llama 3.3 70B and Llama 3.1 8B
Check checkpoint, layer and reconstruction quality before reusing an SAE. Record the source model/task, a baseline, the changed factor, a measurement and an alternative explanation. State exactly which part your notebook investigates and which part it does not reproduce.
Open weights enable investigation; they do not certify feature labels or universal coverage.
Use the notebook below to collect the measurements. This writing remains in this page session until downloaded.
Teacher background and source method
Llama 3.1 8B layer 19 and Llama 3.3 70B layer 50
Releases SAEs and evaluates sparsity, fidelity and judged steering. Check judge dependence and checkpoint/layer compatibility.
Independent classroom adaptation; not a reproduction of the source model or complete method.
Read the source with a teacher ↗Opening the notebook page…
First use downloads Python and its libraries. A fresh session measured about 20–22 seconds and 14–17 MB during the initial audit; school networks vary. This hosted notebook computes on the browser CPU.
Your prediction and saved runs stay in this session. Download your evidence before leaving or refreshing. The hosted experiment runs on your browser CPU and does not connect to an H100.
Research connections · 8 archive entries
These lessons adapt ideas and methods. The original sources state their own model, data and validation scope.
Announcing Open-Source SAEs for Llama 3.3 70B and Llama 3.1 8B ↗
Open weights enable investigation; they do not certify feature labels or universal coverage.
Adversarial Examples Are Not Bugs, They Are Superposition ↗
Bidirectional causal evidence is strongest in toy models; real vision-model evidence is narrower, not a universal explanation.
Uncovering Neural Geometry in Vision Models With Block-Sparse Featurizers ↗
The experiments concern vision and diffusion models; the classroom geometry is an analogy for representation, not an LLM replication.
Can SAEs Capture Neural Geometry? ↗
Low reconstruction error alone does not establish semantic completeness; our small SAE is not the paper's full evaluation.
Interpreting Evo 2: Arc Institute's Next-Generation Genomic Foundation Model ↗
DNA models are not ordinary text LLMs. The 2025 report was updated to note Nature publication in March 2026.
Mapping the Latent Space of Llama 3.3 70B ↗
A two-dimensional map is a lossy view. The older Ember demo is deprecated.
Understanding Sparse Autoencoder Scaling in the Presence of Feature Manifolds ↗
The paper identifies regimes and mechanisms, not a claim that all larger SAEs get worse.
Understanding and Steering Llama 3 with Sparse Autoencoders ↗
Features can overlap or duplicate. Ember API and demo links are deprecated; these labs need neither.