Wattle’s clue microscope · teacher notes

Year 6 · 35 minutes · Live trained model

Look at hidden unit activity across six fact cards. Suggest a label for a clue, then try to break your own explanation.

Student paper journal · Interactive notebook · Editable Python

Preparation and access

Try the starting controls, save two runs and reopen a downloaded project. Print the student journal before class. If Python cannot load, use the supplied example below for a prediction and evidence critique; rerunning changed inputs requires the live or native notebook. Do not present paper discussion as a new model experiment. Draw, point or explain aloud. A partner or adult may record your words.

Teaching sequence

Start with 5–8 minutes of prediction, use about half the lesson for paired comparisons, then reserve at least 10 minutes for a counterexample and an individual explanation. A unit is a measured number, not a feeling or a verified concept. Keep the block fixed while changing the unit. The heat map shows only the first 12 units; the selected unit may be outside that view.

Supplied starting example — reveal after prediction

Settings: {'layer': 0, 'neuron': 0, 'authored': ''}

Neuron 0 in block 0 across 6 examples. A large number is a clue to investigate, not a verified label.

example prompt supported activation model_answer
1 pip has red . wattle has blue . pip has red 0.0 red
2 pip has blue . wattle has red . pip has blue 0.0 blue
3 kiki has gold . bo has green . kiki has gold 0.0 gold
4 bo has gold . kiki has red . kiki has red 0.0 red
5 pip has gold . wattle has red . pip has gold 0.17858 red
6 pip has red . wattle has blue . wattle has blue 0.0 blue

The table shows up to eight rows. Inspect the notebook for all values, controls and denominators.

Method and scope: Hidden unit activation = ReLU(normalised residual × learned weights + bias). The heat map shows units 0–11 at each token of your current prompt (the authored prompt when supplied); the bars show your selected unit across examples.

Assessment and feedback

Assess four criteria, each 0–2: testable prediction; comparison identifying what stayed fixed; accurate use of exact evidence; counterexample and limited conclusion. 0 means absent or contradicted by the record, 1 means partly supported, and 2 means clear and supported. A surprising result earns no penalty.

Beginning response: ‘The picture looks right.’ Ask for a particular case. Developing response: names one number without its setting. Ask what comparison supports it. Secure response: states the setting and measured change and separates the notebook's result from the source paper's claim. Extending response: authors or reserves a new case, tests an alternative explanation, and revises the claim if needed. These are marking examples, not pupil data.

Extension: Download the code and add another fictional card. Test the label before changing it.

Research connections and separate classroom tasks

Announcing Our $50M Series A to Advance AI Interpretability Research

Read the source

System and data: Company funding and strategy

Method: Reports financing and plans; investment is not a model evaluation.

Task: Separate a promise about a tool from evidence that it works. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: Funding is not validation. The older Ember service is deprecated.

Intentionally Designing the Future of AI

Read the source

System and data: Research agenda with worked training examples

Method: Proposes observing per-example learning signals and changing what training generalises; separate aspirations from demonstrated interventions.

Task: Change the teaching examples, then test different examples. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: This is a research agenda; intentional design is not a solved capability.

Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Model

Read the source

System and data: Kimi K2 Thinking activation collection

Method: Collects billions of token-linked tensors for SAE training; batching, memory and provenance are engineering evidence, not a safety benchmark.

Task: Keep each clue attached to the sentence it came from. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: A fast pipeline can still collect misaligned evidence. Classroom timings do not benchmark frontier hardware.

Understanding, Learning From, and Designing AI: Our Series B

Read the source

System and data: Company financing and intentional-design strategy

Method: Reports funding and proposed applications; follow the underlying experiments for empirical evidence.

Task: Distinguish what has been demonstrated from what people hope to build. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: A company announcement provides strategy and context, not an independent benchmark.

Announcing Open-Source SAEs for Llama 3.3 70B and Llama 3.1 8B

Read the source

System and data: Llama 3.1 8B layer 19 and Llama 3.3 70B layer 50

Method: Releases SAEs and evaluates sparsity, fidelity and judged steering. Check judge dependence and checkpoint/layer compatibility.

Task: A labelled clue map can be inspected and questioned. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: Open weights enable investigation; they do not certify feature labels or universal coverage.

Adversarial Examples Are Not Bugs, They Are Superposition

Read the source

System and data: Known toy representations and ResNet18 vision models

Method: Varies overlap and adversarial robustness; bidirectional causal evidence is demonstrated in toy models, with a narrower direction tested in vision.

Task: Two clues sharing the same space can be confused. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: Bidirectional causal evidence is strongest in toy models; real vision-model evidence is narrower, not a universal explanation.

Can SAEs Capture Neural Geometry?

Read the source

System and data: Synthetic shapes and Llama 3.1 8B activations

Method: Varies SAE capacity and examines splitting, dilution and recovery of manifolds; reconstruction alone cannot measure concept completeness.

Task: A big clue map might split one idea into many labels. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: Low reconstruction error alone does not establish semantic completeness; our small SAE is not the paper's full evaluation.

Covariance-based Sequence Pooling

Read the source

System and data: NTv3 gene ontology and genomic-track tasks

Method: Compares mean pooling with compressed second-order sequence statistics using downstream probes; regularisation matters when labels are scarce.

Task: Two collections can have the same average but different paired patterns. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: The source method uses second moments and approximations; our centred-covariance experiment is a teaching analogue. It does not recover sequence order.

Interpreting Evo 2: Arc Institute's Next-Generation Genomic Foundation Model

Read the source

System and data: Evo 2 genomic activations

Method: Trains SAEs and compares discovered features with biological annotations; associated signals motivate hypotheses rather than clinical conclusions.

Task: Check whether a proposed clue appears in examples that should and should not match. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: DNA models are not ordinary text LLMs. The 2025 report was updated to note Nature publication in March 2026.

Mapping the Latent Space of Llama 3.3 70B

Read the source

System and data: An intermediate layer of Llama 3.3 70B

Method: Maps SAE feature relationships and demonstrates steering selected features; the displayed clusters are selected examples, not exhaustive coverage.

Task: Look for a clue, then deliberately find a counterexample to its label. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: A two-dimensional map is a lossy view. The older Ember demo is deprecated.

Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train

Read the source

System and data: Llama base models with Dolci and Tulu preference data

Method: Inspects predicted per-example learning effects before training, then compares intended and unintended signals in realistic preference datasets.

Task: Inspect the answer key before teaching from it. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: Our supervised colour task illustrates data effects; it is not a replication of contrastive-SAE post-training or DPO.

Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training

Read the source

System and data: OLMo 2 7B SFT/DPO, preference data and 120 held-out LMSYS prompts

Method: Matches behaviour-difference activation vectors to preference-pair vectors, then filters or swaps ranked data and retrains to test the attribution causally.

Task: Find which teaching examples might explain a repeated mistake. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: The paper evaluates particular models and behaviours. Our checkpoint comparison does not implement its attribution algorithm.

Understanding Sparse Autoencoder Scaling in the Presence of Feature Manifolds

Read the source

System and data: ReLU SAEs on circles and other known feature manifolds

Method: Studies how adding dictionary capacity can tile common manifolds and lower loss while leaving rare features undiscovered.

Task: More labels do not automatically mean all ideas are covered. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: The paper identifies regimes and mechanisms, not a claim that all larger SAEs get worse.

Understanding and Steering Llama 3 with Sparse Autoencoders

Read the source

System and data: Llama-3-8B and LMSYS-Chat-1M

Method: Trains an SAE, inspects features and demonstrates activation steering; feature labels and generated examples require counterexamples and evaluation.

Task: Invent a clue label, then try examples that challenge it. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: Features can overlap or duplicate. Ember API and demo links are deprecated; these labs need neither.

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

Read the source

System and data: Power-law toy tasks and OLMo models from 4M to 4B

Method: Compares per-task losses, representations and gradient interference across capacity, including infrequent tasks; average loss can conceal retention failures.

Task: A frequently practised skill can crowd out a less common one. Use the linked notebook to make two observations. Draw or describe one result and one thing this activity cannot tell us about the source system.

Boundary: The reported mechanism is studied in particular toy and language models; size is not a guarantee on every task.

Independent Brightlab adaptations; no Goodfire endorsement or full reproduction claim. Source methods reviewed 7 September 2026. Curriculum connections are selected planning links; confirm your school syllabus.