A complete map of the public research and blog archive we found on 6 September 2026.
Perspective · August 27, 2026
Useful safety research also needs reliable engineering and deployment.
Research question, system and method
What evidence would support or challenge this idea: Useful safety research also needs reliable engineering and deployment.
System and data: Engineering perspective; examples of safety infrastructure
How it was investigated: Reason through deployment, testing and operational failure modes; no new controlled model experiment.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: An engineering argument, not experimental proof of model safety.
Announcement · August 20, 2026
Funding and model access can help researchers test interpretability ideas.
Research question, system and method
What evidence would support or challenge this idea: Funding and model access can help researchers test interpretability ideas.
System and data: Research funding programme
How it was investigated: Describes eligibility and support; assess proposed experiments separately from the announcement.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A grants announcement supplies opportunities, not a new scientific result.
Announcement · April 17, 2025
Goodfire raised funding to develop interpretability tools.
Research question, system and method
What evidence would support or challenge this idea: Goodfire raised funding to develop interpretability tools.
System and data: Company funding and strategy
How it was investigated: Reports financing and plans; investment is not a model evaluation.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Funding is not validation. The older Ember service is deprecated.
Lecture guide
Three guest lectures cover causal interpretation, circuits and learning in context.
Research question, system and method
What evidence would support or challenge this idea: Three guest lectures cover causal interpretation, circuits and learning in context.
System and data: Three interpretability lecture outlines
How it was investigated: Introduces causal abstraction, circuits and learning in context; a teaching guide rather than one dataset.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: This entry reviews the article and lecture outlines, not a full transcription of the videos.
Explainer · Nov. 20, 2024
Changing selected internal activations can alter generated behaviour.
Research question, system and method
What evidence would support or challenge this idea: Changing selected internal activations can alter generated behaviour.
System and data: Historical Llama steering demonstrations
How it was investigated: Changes selected internal feature strengths and inspects example outputs; demonstrations need independent task and side-effect checks.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Steering is not guaranteed control or a source of new knowledge. Ember links are historical.
Announcement · October 9, 2025
Interpretability research combines experimental questions with software engineering.
Research question, system and method
What evidence would support or challenge this idea: Interpretability research combines experimental questions with software engineering.
System and data: Research fellowship programme
How it was investigated: Describes research directions and participation; does not test a scientific hypothesis.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A fellowship announcement is careers context, not a research finding.
Perspective · February 5, 2026
Interpretability could guide training by identifying what data teaches a model.
Research question, system and method
What evidence would support or challenge this idea: Interpretability could guide training by identifying what data teaches a model.
System and data: Research agenda with worked training examples
How it was investigated: Proposes observing per-example learning signals and changing what training generalises; separate aspirations from demonstrated interventions.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: This is a research agenda; intentional design is not a solved capability.
Engineering report · February 25, 2026
Large-model activation collection needs careful memory, batching and data handling.
Research question, system and method
What evidence would support or challenge this idea: Large-model activation collection needs careful memory, batching and data handling.
System and data: Kimi K2 Thinking activation collection
How it was investigated: Collects billions of token-linked tensors for SAE training; batching, memory and provenance are engineering evidence, not a safety benchmark.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A fast pipeline can still collect misaligned evidence. Classroom timings do not benchmark frontier hardware.
Announcement · September 9, 2025
Goodfire and Mayo Clinic announced work on genomic model interpretation.
Research question, system and method
What evidence would support or challenge this idea: Goodfire and Mayo Clinic announced work on genomic model interpretation.
System and data: Genomic research collaboration
How it was investigated: Describes planned interpretation of biological model representations; requires later independent scientific validation.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A collaboration announcement is not clinical validation. No pupil health data are used.
Perspective · July 17, 2025
Early progress motivates trying to understand and deliberately improve neural networks.
Research question, system and method
What evidence would support or challenge this idea: Early progress motivates trying to understand and deliberately improve neural networks.
System and data: Cross-domain interpretability perspective
How it was investigated: Builds an argument from existing examples and open problems; optimism is not a measured success rate.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The author explicitly presents optimism amid unsolved scientific questions.
Perspective · Dec. 23, 2024
Goodfire describes moderation, access controls and research collaboration for its tools.
Research question, system and method
What evidence would support or challenge this idea: Goodfire describes moderation, access controls and research collaboration for its tools.
System and data: Policies for interpretability tools
How it was investigated: Describes moderation, feature access and research processes; organisational policy and model capability are separate claims.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A stated safety process is not proof that all failure modes are covered.
Announcement
A funding update connects interpretability, scientific discovery and intentional model design.
Research question, system and method
What evidence would support or challenge this idea: A funding update connects interpretability, scientific discovery and intentional model design.
System and data: Company financing and intentional-design strategy
How it was investigated: Reports funding and proposed applications; follow the underlying experiments for empirical evidence.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A company announcement provides strategy and context, not an independent benchmark.
Announcement · July 30, 2025
A partnership proposes using interpretability for materials discovery.
Research question, system and method
What evidence would support or challenge this idea: A partnership proposes using interpretability for materials discovery.
System and data: Materials-discovery partnership
How it was investigated: Announces collaboration; a predicted material must still be independently evaluated.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The announcement is not a materials experiment or an LLM result.
Resource release · Jan. 10, 2025
Goodfire released sparse autoencoders for particular Llama checkpoints and layers.
Research question, system and method
What evidence would support or challenge this idea: Goodfire released sparse autoencoders for particular Llama checkpoints and layers.
System and data: Llama 3.1 8B layer 19 and Llama 3.3 70B layer 50
How it was investigated: Releases SAEs and evaluates sparsity, fidelity and judged steering. Check judge dependence and checkpoint/layer compatibility.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Open weights enable investigation; they do not certify feature labels or universal coverage.
Assurance report · May 22, 2026
Goodfire reports an audit of organisational security controls over a period of time.
Research question, system and method
What evidence would support or challenge this idea: Goodfire reports an audit of organisational security controls over a period of time.
System and data: Organisational security-control audit
How it was investigated: Reports independent control assurance over an audit period; does not evaluate the truth of generated answers.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: SOC 2 is not a certificate of LLM truthfulness or absence of harmful behaviour.
Engineering reflection · October 2, 2025
Research agents can accelerate experiments, but supervision and validation become bottlenecks.
Research question, system and method
What evidence would support or challenge this idea: Research agents can accelerate experiments, but supervision and validation become bottlenecks.
System and data: Scribe experiments on Evo 2 and GPT-2 examples
How it was investigated: Reports agent-assisted research tasks and failure cases; provenance checks are needed when agents can shortcut work.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The shared task suite is described as directional, not a rigorously audited benchmark.
LLM research · May 14, 2026
Llama's cyclic reasoning reuses an internal base-10 addition mechanism before mapping back to a cycle.
Research question, system and method
What evidence would support or challenge this idea: Llama's cyclic reasoning reuses an internal base-10 addition mechanism before mapping back to a cycle.
System and data: Llama 3.1 8B arithmetic and cyclic tasks
How it was investigated: Tracks layer/token representations, identifies a shared addition mechanism and checks it with causal interventions.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
For Year 12
Contrast an explanatory geometry with the arithmetic mechanism established by causal interventions.
Steer along a manifold →What the evidence does—and doesn’t—cover: The paper does not claim that Llama simply adds directly around a circle. Our calendar is a teaching model.
Linked primary research: Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts ↗Mixed-model research
Experiments connect overlapping representations with vulnerability to small input changes.
Research question, system and method
What evidence would support or challenge this idea: Experiments connect overlapping representations with vulnerability to small input changes.
System and data: Known toy representations and ResNet18 vision models
How it was investigated: Varies overlap and adversarial robustness; bidirectional causal evidence is demonstrated in toy models, with a narrower direction tested in vision.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Bidirectional causal evidence is strongest in toy models; real vision-model evidence is narrower, not a universal explanation.
Linked primary research: Adversarial Examples Are Not Bugs, They Are Superposition ↗LLM research
A Bayesian framework relates contextual evidence to likelihood updates and steering to prior shifts.
Research question, system and method
What evidence would support or challenge this idea: A Bayesian framework relates contextual evidence to likelihood updates and steering to prior shifts.
System and data: Many-shot context and activation-steering experiments
How it was investigated: Fits a Bayesian account that distinguishes likelihood updates from changes to priors; tests how context strength changes steering effects.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A fitted account of model behaviour does not establish that a model has human beliefs.
Linked primary research: Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering ↗Vision research · July 7, 2026
Block-sparse features represent concepts using subspaces rather than only single directions.
Research question, system and method
What evidence would support or challenge this idea: Block-sparse features represent concepts using subspaces rather than only single directions.
System and data: Synthetic manifolds, DINOv3 and SDXL
How it was investigated: Trains block-sparse featurizers, compares reconstruction and subspace coverage, and tests edits; multidimensional features are evaluated against direction-based baselines.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The experiments concern vision and diffusion models; the classroom geometry is an analogy for representation, not an LLM replication.
Linked primary research: Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds ↗Representation research · May 21, 2026
SAEs can split, dilute or compactly capture parts of concept manifolds.
Research question, system and method
What evidence would support or challenge this idea: SAEs can split, dilute or compactly capture parts of concept manifolds.
System and data: Synthetic shapes and Llama 3.1 8B activations
How it was investigated: Varies SAE capacity and examines splitting, dilution and recovery of manifolds; reconstruction alone cannot measure concept completeness.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Low reconstruction error alone does not establish semantic completeness; our small SAE is not the paper's full evaluation.
Linked primary research: Do Sparse Autoencoders Capture Concept Manifolds? ↗Genomic research · April 10, 2026
Second-order pooling can preserve co-activation information that a simple mean loses.
Research question, system and method
What evidence would support or challenge this idea: Second-order pooling can preserve co-activation information that a simple mean loses.
System and data: NTv3 gene ontology and genomic-track tasks
How it was investigated: Compares mean pooling with compressed second-order sequence statistics using downstream probes; regularisation matters when labels are scarce.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The source method uses second moments and approximations; our centred-covariance experiment is a teaching analogue. It does not recover sequence order.
Genomic research
Evo 2 representations support variant-effect predictions and evidence-grounded generated explanations.
Research question, system and method
What evidence would support or challenge this idea: Evo 2 representations support variant-effect predictions and evidence-grounded generated explanations.
System and data: Evo 2 embeddings and labelled genetic variants
How it was investigated: Fits variant-effect predictors and constructs evidence-based mechanistic hypotheses; abstract and research report were available for review.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Genomic predictions are hypotheses, not clinical conclusions. The linked preprint abstract and Goodfire report were reviewed; full preprint text was unavailable.
Linked primary research: Interpretable variant effect prediction from genomic foundation model representations | bioRxiv ↗LLM research
Repeated continuations from shared prefixes reveal how answer uncertainty changes along a generation.
Research question, system and method
What evidence would support or challenge this idea: Repeated continuations from shared prefixes reveal how answer uncertainty changes along a generation.
System and data: Llama-3-8B-Instruct and DeepSeek-R1-Distill-Llama-8B on tinyMMLU
How it was investigated: Resamples continuations at shared prefixes and compares uncertainty estimates across sample budgets, spacing and smoothing.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Our finite branching simulator is not a reasoning LLM. More samples reduce sampling noise, not model bias.
Linked primary research: Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation ↗Genomic research · January 28, 2026
Interpretability helped identify candidate fragment-length signals for a biomarker classifier.
Research question, system and method
What evidence would support or challenge this idea: Interpretability helped identify candidate fragment-length signals for a biomarker classifier.
System and data: Pleiades and cell-free DNA pilot cohorts
How it was investigated: Uses model interpretation to propose fragment-length signals, then tests a simpler classifier on an independent cohort with uncertainty intervals.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: This is a small-cohort biomedical study, not a diagnostic classroom tool or an LLM result.
Genomic research · Feb. 20, 2025
SAEs revealed biologically associated features in a genomic foundation model.
Research question, system and method
What evidence would support or challenge this idea: SAEs revealed biologically associated features in a genomic foundation model.
System and data: Evo 2 genomic activations
How it was investigated: Trains SAEs and compares discovered features with biological annotations; associated signals motivate hypotheses rather than clinical conclusions.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: DNA models are not ordinary text LLMs. The 2025 report was updated to note Nature publication in March 2026.
LLM research · May 5th 2026
VPD decomposes model weights into simpler components while preserving behaviour under its objectives.
Research question, system and method
What evidence would support or challenge this idea: VPD decomposes model weights into simpler components while preserving behaviour under its objectives.
System and data: A four-layer 67M language model trained on the Pile
How it was investigated: Learns parameter subcomponents and causal importance with VPD; evaluates reconstruction, sparsity, interventions and robustness against alternative decompositions.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
For Year 12
Compare exact reconstruction, component removal and behavioural tests of a decomposition.
Take the weights apart →What the evidence does—and doesn’t—cover: Our SVD exercise is a contrast, not VPD. The paper studies a 67M-parameter model, not all frontier models.
LLM research
The probability of a selected phrase can provide a low-cost signal of evaluation awareness.
Research question, system and method
What evidence would support or challenge this idea: The probability of a selected phrase can provide a low-cost signal of evaluation awareness.
System and data: Kimi K2.5 and Qwen3-32B; Fortress/Petri evaluation settings compared with WildChat conversations
How it was investigated: At sentence boundaries, scores a prespecified evaluation-awareness phrase from token probabilities and compares the monitor with explicit verbalisation and judge-based monitoring. Reported rollout savings are specific to these experiments; the score is a proxy for awareness, not a direct mental-state measurement.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Phrase choice and prompt framing matter. High AUROC does not establish intent or universal reliability.
Linked primary research: Linked primary research ↗LLM and world-model research · May 7, 2026
Following a fitted curved representation can control cyclic behaviour more effectively than a straight displacement.
Research question, system and method
What evidence would support or challenge this idea: Following a fitted curved representation can control cyclic behaviour more effectively than a straight displacement.
System and data: Llama 3.1 8B weekday behaviour and activations
How it was investigated: Fits related behavioural and activation geometry, then compares movement along fitted manifolds with linear edits.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Only selected fitted manifolds and tasks were tested; not every concept is circular.
Linked primary research: Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior ↗LLM explainer · Dec. 23, 2024
An SAE feature map provides a navigable view of patterns in Llama activations.
Research question, system and method
What evidence would support or challenge this idea: An SAE feature map provides a navigable view of patterns in Llama activations.
System and data: An intermediate layer of Llama 3.3 70B
How it was investigated: Maps SAE feature relationships and demonstrates steering selected features; the displayed clusters are selected examples, not exhaustive coverage.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A two-dimensional map is a lossy view. The older Ember demo is deprecated.
LLM research
Models combine lexical, positional and other mechanisms when retrieving bound entities.
Research question, system and method
What evidence would support or challenge this idea: Models combine lexical, positional and other mechanisms when retrieving bound entities.
System and data: Llama, Gemma and Qwen families, 2–72B, ten binding tasks
How it was investigated: Uses ablations to separate positional, lexical and reflexive retrieval; compares causal predictions as entity lists and context length change.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Our tiny binding transformer is independently trained and does not reproduce the paper's nine-model circuit findings.
Linked primary research: Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context ↗LLM research · August 21, 2025
Amplifying logit differences between checkpoints can expose changed, otherwise rare behaviour.
Research question, system and method
What evidence would support or challenge this idea: Amplifying logit differences between checkpoints can expose changed, otherwise rare behaviour.
System and data: Paired language-model checkpoints and controlled fine-tunes
How it was investigated: Amplifies output-score differences to expose changed behaviour; increased discovery frequency is not natural failure prevalence.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Amplified discovery is not an estimate of how often the original model fails.
Series index
A collection connects cyclic concepts, stories, sparse features and other neural geometry studies.
Research question, system and method
What evidence would support or challenge this idea: A collection connects cyclic concepts, stories, sparse features and other neural geometry studies.
System and data: Research collection across several model domains
How it was investigated: Organises related studies of representation geometry; the collection is not an independent experiment beyond its linked papers.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: This is an index, not an additional independent experiment.
Research survey
A survey organises unsolved problems in understanding, validating and applying interpretability.
Research question, system and method
What evidence would support or challenge this idea: A survey organises unsolved problems in understanding, validating and applying interpretability.
System and data: Mechanistic interpretability survey
How it was investigated: Organises unresolved conceptual, methodological and practical questions; evaluate the evidence for each underlying method separately.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A research roadmap is not evidence that these problems have been solved.
Linked primary research: Open Problems in Mechanistic Interpretability ↗Vision research · May 27, 2025
Interventions in diffusion-model features can edit concepts at selected image locations.
Research question, system and method
What evidence would support or challenge this idea: Interventions in diffusion-model features can edit concepts at selected image locations.
System and data: Image-patch representations and a BatchTopK SAE
How it was investigated: Decomposes image representations into sparse features and edits selected patches; visual examples and reconstruction are distinct from reliable semantic control.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
For Year 12
Consider fragmentation, spatial location and interactions between prompts and feature interventions.
Steer along a manifold →What the evidence does—and doesn’t—cover: This is image generation, not an LLM experiment; the classroom intervention is a transferable analogy.
Genomic research · August 27, 2025
A learned low-dimensional representation in Evo 2 reflects evolutionary relationships.
Research question, system and method
What evidence would support or challenge this idea: A learned low-dimensional representation in Evo 2 reflects evolutionary relationships.
System and data: Evo 2 representations of cross-species DNA
How it was investigated: Constructs data to distinguish evolutionary relationship from simple sequence similarity and compares distances along learned geometry.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
For Year 12
Ask whether a fitted representation generalises to held-out clades and controls for sequence similarity.
Steer along a manifold →What the evidence does—and doesn’t—cover: Biological geometry does not establish an equivalent map inside text LLMs.
LLM research
Logit Path Extrapolation estimates rare failures using a related model and an empirical trend.
Research question, system and method
What evidence would support or challenge this idea: Logit Path Extrapolation estimates rare failures using a related model and an empirical trend.
System and data: Qwen3-4B and an abliterated variant on HarmBench, with a 100,000-rollout reference
How it was investigated: Interpolates the paired models in logit space, measures compliance along the path, fits an empirical log-linear trend below 50% compliance and extrapolates to the original model. Requires a related variant and a suitable trend; the classroom Wilson interval uses a different, direct-sampling method.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The reported efficiency is setting-dependent. Our binomial experiment explains uncertainty; it does not implement the paper's extrapolation method.
Linked primary research: Linked primary research ↗LLM research · June 11, 2026
Contrasts between preferred and rejected training examples can forecast and alter learned behaviours.
Research question, system and method
What evidence would support or challenge this idea: Contrasts between preferred and rejected training examples can forecast and alter learned behaviours.
System and data: Llama base models with Dolci and Tulu preference data
How it was investigated: Inspects predicted per-example learning effects before training, then compares intended and unintended signals in realistic preference datasets.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Our supervised colour task illustrates data effects; it is not a replication of contrastive-SAE post-training or DPO.
Linked primary research: Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal ↗LLM research
Temporal feature analysis separates predictable context from new information over a sequence.
Research question, system and method
What evidence would support or challenge this idea: Temporal feature analysis separates predictable context from new information over a sequence.
System and data: Gemma-2-2B activations from Pile-Uncopyrighted
How it was investigated: Compares temporal feature analysis with ReLU, TopK and BatchTopK SAEs, testing predictable versus innovation components and event structure.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Our Bayesian story model is explicit and hand-specified; it is not Temporal Feature Analysis applied to an LLM.
Linked primary research: Priors in Time: Missing Inductive Biases for Language Model Interpretability ↗LLM research · April 29, 2026
Behavioural activation directions can rank training pairs for targeted filtering or relabelling.
Research question, system and method
What evidence would support or challenge this idea: Behavioural activation directions can rank training pairs for targeted filtering or relabelling.
System and data: OLMo 2 7B SFT/DPO, preference data and 120 held-out LMSYS prompts
How it was investigated: Matches behaviour-difference activation vectors to preference-pair vectors, then filters or swaps ranked data and retrains to test the attribution causally.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The paper evaluates particular models and behaviours. Our checkpoint comparison does not implement its attribution algorithm.
Linked primary research: Probe-Based Data Attribution: Discovering and Mitigating Undesirable Behaviors in LLM Post-Training ↗LLM application · October 28, 2025
SAE-based probes helped detect personal information under noisy and shifted text conditions.
Research question, system and method
What evidence would support or challenge this idea: SAE-based probes helped detect personal information under noisy and shifted text conditions.
System and data: Llama 3.1 8B sidecar, English/Japanese synthetic training and real production tests
How it was investigated: Compares activation, attention and frozen-SAE probes for token-level PII detection across synthetic-to-real shift, label noise and language settings.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Results depend on data and baselines; SAEs are not always superior. Classroom fixtures contain no real personal information.
LLM research · March 12, 2026
Probes can sometimes predict a model's eventual answer before its written reasoning ends.
Research question, system and method
What evidence would support or challenge this idea: Probes can sometimes predict a model's eventual answer before its written reasoning ends.
System and data: DeepSeek-R1 families and GPT-OSS-120B on MMLU-Redux and GPQA-Diamond
How it was investigated: Trains context-pooling probes to forecast eventual answers during reasoning and evaluates early-exit token cost versus benchmark accuracy.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Savings depend on task and paper version; the Goodfire post and later arXiv revision report different percentages. Our trajectories are simulated.
Linked primary research: Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought ↗LLM research · June 11, 2025
Cross-layer transcoders recovered a version of a greater-than mechanism in GPT-2 small.
Research question, system and method
What evidence would support or challenge this idea: Cross-layer transcoders recovered a version of a greater-than mechanism in GPT-2 small.
System and data: GPT-2 Small on a known greater-than task
How it was investigated: Trains cross-layer transcoders, generates attribution graphs and compares the recovered mechanism with a previously characterised circuit.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Replication found both similarities and differences. Our attention patching does not reproduce CLT circuit discovery.
LLM research · February 11, 2026
Frozen-model feature probes supplied rewards in a combined hallucination-reduction approach.
Research question, system and method
What evidence would support or challenge this idea: Frozen-model feature probes supplied rewards in a combined hallucination-reduction approach.
System and data: Gemma-3-12B-IT and LongFact++ with 999 held-out prompts
How it was investigated: Trains factuality/correction probes and uses a frozen model for rewards; compares RL, inline intervention and best-of-N contributions with independent labels.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The paper combines methods; its headline reduction is not attributable to feature rewards alone. Our reward pool is synthetic.
Linked primary research: Features as Rewards: Scalable Supervision forOpen-Ended Tasks via Interpretability ↗Representation research
Larger dictionaries can spend capacity tiling common manifolds while missing rare features.
Research question, system and method
What evidence would support or challenge this idea: Larger dictionaries can spend capacity tiling common manifolds while missing rare features.
System and data: ReLU SAEs on circles and other known feature manifolds
How it was investigated: Studies how adding dictionary capacity can tile common manifolds and lower loss while leaving rare features undiscovered.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The paper identifies regimes and mechanisms, not a claim that all larger SAEs get worse.
Linked primary research: Understanding sparse autoencoder scaling in the presence of feature manifolds ↗Materials research · April 1, 2026
An internal property probe guides accept/reject decisions during generated-material search.
Research question, system and method
What evidence would support or challenge this idea: An internal property probe guides accept/reject decisions during generated-material search.
System and data: Band-gap-conditioned MatterGen diffusion
How it was investigated: Uses an activation probe to accept or reject proposed denoising steps and evaluates targeting, stability, uniqueness and novelty.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A model's predicted material property is not a physical measurement; this is not text-LLM research.
Parameter research · June 27, 2025
Stochastic partial ablations encourage a decomposition into simpler weight components.
Research question, system and method
What evidence would support or challenge this idea: Stochastic partial ablations encourage a decomposition into simpler weight components.
System and data: Toy networks with known ground-truth mechanisms
How it was investigated: Learns parameter subcomponents using stochastic ablations and evaluates recovery of known computations; full-language-model scaling is a separate result.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: SPD evidence here is based on controlled models; the SVD notebook is a baseline contrast, not an SPD implementation.
Linked primary research: Stochastic Parameter Decomposition ↗LLM research · June 23, 2026
Story representations follow trajectories associated with genres and characters' emotions.
Research question, system and method
What evidence would support or challenge this idea: Story representations follow trajectories associated with genres and characters' emotions.
System and data: Llama 3.1 8B and SimpleStories
How it was investigated: Collects sentence-end activations, fits trajectories and compares emotional organisation with behavioural readouts and human valence/arousal judgements.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Representing a character's emotions does not mean the model has emotions. Classroom story probabilities are hand-specified.
Linked primary research: Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space ↗Research survey
A community guide compares approaches and evidence in circuit research.
Research question, system and method
What evidence would support or challenge this idea: A community guide compares approaches and evidence in circuit research.
System and data: Community synthesis of circuit methods and findings
How it was investigated: Compares probes, sparse decompositions, attribution graphs and causal testing across published work; synthesis is not a single replication.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: The linked community guide was reviewed. Attribution graphs provide hypotheses with missing or uninterpretable components; interventions remain necessary.
Linked primary research: The Circuits Research Landscape: Results and Perspectives ↗Perspective · May 7, 2026
Geometric structure may help explain how networks represent relationships across domains.
Research question, system and method
What evidence would support or challenge this idea: Geometric structure may help explain how networks represent relationships across domains.
System and data: Cross-domain neural-geometry perspective
How it was investigated: Connects structured data to learned representations and an unsupervised geometry-discovery pipeline; individual causal claims require their own experiments.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: A cross-domain research perspective does not prove every useful concept has an easily readable geometry.
LLM research · Apr. 15, 2025
SAEs expose patterns in DeepSeek R1; steering effects depend on strength and where it begins.
Research question, system and method
What evidence would support or challenge this idea: SAEs expose patterns in DeepSeek R1; steering effects depend on strength and where it begins.
System and data: DeepSeek R1 671B on custom reasoning and OpenR1-Math data
How it was investigated: Trains two SAEs and studies feature activations and steering timing/strength; partial features do not expose a complete reasoning transcript.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: An SAE reveals partial patterns, not a complete transcript of reasoning. Effects do not generalise automatically.
LLM research · September 24, 2024
SAEs extract sparse features from Llama activations and support targeted steering experiments.
Research question, system and method
What evidence would support or challenge this idea: SAEs extract sparse features from Llama activations and support targeted steering experiments.
System and data: Llama-3-8B and LMSYS-Chat-1M
How it was investigated: Trains an SAE, inspects features and demonstrates activation steering; feature labels and generated examples require counterexamples and evaluation.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Features can overlap or duplicate. Ember API and demo links are deprecated; these labs need neither.
LLM research · November 6, 2025
Aggregate loss curvature helps identify directions associated with memorisation and other knowledge.
Research question, system and method
What evidence would support or challenge this idea: Aggregate loss curvature helps identify directions associated with memorisation and other knowledge.
System and data: Language models and vision transformers with label noise
How it was investigated: Uses aggregate curvature approximations to identify and edit memorisation-associated weight directions, then checks collateral downstream effects.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
For Year 12
Compare a cheap direction under common-task curvature with its effect on a rare task.
Take the weights apart →What the evidence does—and doesn’t—cover: The source uses curvature approximations including K-FAC; our exact quadratic toy is not a K-FAC replication. Edits have collateral costs.
Linked primary research: From Memorization to Reasoning in the Spectrum of Loss Curvature ↗LLM research · May 4, 2026
Mentioning an evaluation is associated with different measured behaviour; selected interventions support causal effects.
Research question, system and method
What evidence would support or challenge this idea: Mentioning an evaluation is associated with different measured behaviour; selected interventions support causal effects.
System and data: Eight models, nineteen benchmarks; causal work on Kimi K2.5/Fortress
How it was investigated: Manually verifies verbalised awareness, compares matched cues and performs interventions; broad correlations and narrower causal evidence have different scope.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
What the evidence does—and doesn’t—cover: Correlations span more models than the causal intervention study. Silence about a test does not prove absence of awareness.
Paper explainer · May 5th 2026
A readable companion explains VPD's objectives and component-level model editing.
Research question, system and method
What evidence would support or challenge this idea: A readable companion explains VPD's objectives and component-level model editing.
System and data: Companion to the 67M VPD language-model study
How it was investigated: Explains component reconstruction and learned causal importance; refers to the same experiment as Interpreting Language Model Parameters.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
For Year 12
Compare faithful reconstruction with simplicity and robustness of component explanations.
Take the weights apart →What the evidence does—and doesn’t—cover: This explains the same work as Interpreting Language Model Parameters; it is not a second independent result.
Learning research
Larger models can reduce interference that overwrites rare tasks during shared training.
Research question, system and method
What evidence would support or challenge this idea: Larger models can reduce interference that overwrites rare tasks during shared training.
System and data: Power-law toy tasks and OLMo models from 4M to 4B
How it was investigated: Compares per-task losses, representations and gradient interference across capacity, including infrequent tasks; average loss can conceal retention failures.
Independent classroom adaptation; not a reproduction of the source model or complete method. Reviewed 2026-09-07.
For Year 12
Inspect frequency-weighted loss and rare-task retention rather than only average performance.
Take the weights apart →What the evidence does—and doesn’t—cover: The reported mechanism is studied in particular toy and language models; size is not a guarantee on every task.
Linked primary research: Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention ↗