Year 12 · Synthetic evaluation and selection · 55 minutes

Audit the claim, not the confidence

What would count as evidence that a monitor really works?

Open full-screen notebook ↗Download Python notebook ↓Teacher notes & worked example ↗Paper student journal ↗Research source map →

The first load downloads Python and may take a moment. Write a prediction, then open the lab. Controls, help, teacher notes and your evidence download are inside. If the school network blocks the runtime, download the Python notebook and run it with your teacher.

Investigate one research connection

AI Safety Still Needs Great Engineers

Audit data provenance, monitoring and reproducibility alongside model metrics. Record the source model/task, a baseline, the changed factor, a measurement and an alternative explanation. State exactly which part your notebook investigates and which part it does not reproduce.

An engineering argument, not experimental proof of model safety.

Use the notebook below to collect the measurements. This writing remains in this page session until downloaded.

Teacher background and source method

Engineering perspective; examples of safety infrastructure

Reason through deployment, testing and operational failure modes; no new controlled model experiment.

Independent classroom adaptation; not a reproduction of the source model or complete method.

Read the source with a teacher ↗

Opening the notebook page…

First use downloads Python and its libraries. A fresh session measured about 20–22 seconds and 14–17 MB during the initial audit; school networks vary. This hosted notebook computes on the browser CPU.

Your prediction and saved runs stay in this session. Download your evidence before leaving or refreshing. The hosted experiment runs on your browser CPU and does not connect to an H100.

Research connections · 15 archive entries

These lessons adapt ideas and methods. The original sources state their own model, data and validation scope.

AI Safety Still Needs Great Engineers
An engineering argument, not experimental proof of model safety.

Announcing Goodfire Research Grants
A grants announcement supplies opportunities, not a new scientific result.

Announcing Our $50M Series A to Advance AI Interpretability Research
Funding is not validation. The older Ember service is deprecated.

Announcing Goodfire’s Fellowship Program for Interpretability Research
A fellowship announcement is careers context, not a research finding.

Goodfire Announces Collaboration to Advance Genomic Medicine with AI Interpretability
A collaboration announcement is not clinical validation. No pupil health data are used.

Our Approach to Safety at Goodfire
A stated safety process is not proof that all failure modes are covered.

Understanding, Learning From, and Designing AI: Our Series B
A company announcement provides strategy and context, not an independent benchmark.

Partnering with Radical AI to Advance Materials Science With Interpretability
The announcement is not a materials experiment or an LLM result.

Announcing our SOC 2 Type II Certification
SOC 2 is not a certificate of LLM truthfulness or absence of harmful behaviour.

You and Your Research Agent: Lessons From Using Agents for Interpretability Research
The shared task suite is described as directional, not a rigorously audited benchmark.

Logits as a new monitor for evaluation awareness
Phrase choice and prompt framing matter. High AUROC does not establish intent or universal reliability.

Predicting Rare LLM Failures with 30× Fewer Rollouts
The reported efficiency is setting-dependent. Our binomial experiment explains uncertainty; it does not implement the paper's extrapolation method.

Features as Rewards: Using Interpretability to Reduce Hallucinations
The paper combines methods; its headline reduction is not attributable to feature rewards alone. Our reward pool is synthetic.

Using Self-Correcting Search to Accelerate Materials Discovery
A model's predicted material property is not a physical measurement; this is not text-LLM research.

Verbalized Eval Awareness Inflates Measured Safety
Correlations span more models than the causal intervention study. Silence about a test does not prove absence of awareness.