Year 12 · Synthetic evaluation and selection · 55 minutes
Audit the claim, not the confidence
What would count as evidence that a monitor really works?
The first load downloads Python and may take a moment. Write a prediction, then open the lab. Controls, help, teacher notes and your evidence download are inside. If the school network blocks the runtime, download the Python notebook and run it with your teacher.
Investigate one research connection
AI Safety Still Needs Great Engineers
Audit data provenance, monitoring and reproducibility alongside model metrics. Record the source model/task, a baseline, the changed factor, a measurement and an alternative explanation. State exactly which part your notebook investigates and which part it does not reproduce.
An engineering argument, not experimental proof of model safety.
Use the notebook below to collect the measurements. This writing remains in this page session until downloaded.
Teacher background and source method
Engineering perspective; examples of safety infrastructure
Reason through deployment, testing and operational failure modes; no new controlled model experiment.
Independent classroom adaptation; not a reproduction of the source model or complete method.
Read the source with a teacher ↗Opening the notebook page…
First use downloads Python and its libraries. A fresh session measured about 20–22 seconds and 14–17 MB during the initial audit; school networks vary. This hosted notebook computes on the browser CPU.
Your prediction and saved runs stay in this session. Download your evidence before leaving or refreshing. The hosted experiment runs on your browser CPU and does not connect to an H100.
Research connections · 15 archive entries
These lessons adapt ideas and methods. The original sources state their own model, data and validation scope.
AI Safety Still Needs Great Engineers ↗
An engineering argument, not experimental proof of model safety.
Announcing Goodfire Research Grants ↗
A grants announcement supplies opportunities, not a new scientific result.
Announcing Our $50M Series A to Advance AI Interpretability Research ↗
Funding is not validation. The older Ember service is deprecated.
Announcing Goodfire’s Fellowship Program for Interpretability Research ↗
A fellowship announcement is careers context, not a research finding.
Goodfire Announces Collaboration to Advance Genomic Medicine with AI Interpretability ↗
A collaboration announcement is not clinical validation. No pupil health data are used.
Our Approach to Safety at Goodfire ↗
A stated safety process is not proof that all failure modes are covered.
Understanding, Learning From, and Designing AI: Our Series B ↗
A company announcement provides strategy and context, not an independent benchmark.
Partnering with Radical AI to Advance Materials Science With Interpretability ↗
The announcement is not a materials experiment or an LLM result.
Announcing our SOC 2 Type II Certification ↗
SOC 2 is not a certificate of LLM truthfulness or absence of harmful behaviour.
You and Your Research Agent: Lessons From Using Agents for Interpretability Research ↗
The shared task suite is described as directional, not a rigorously audited benchmark.
Logits as a new monitor for evaluation awareness ↗
Phrase choice and prompt framing matter. High AUROC does not establish intent or universal reliability.
Predicting Rare LLM Failures with 30× Fewer Rollouts ↗
The reported efficiency is setting-dependent. Our binomial experiment explains uncertainty; it does not implement the paper's extrapolation method.
Features as Rewards: Using Interpretability to Reduce Hallucinations ↗
The paper combines methods; its headline reduction is not attributable to feature rewards alone. Our reward pool is synthetic.
Using Self-Correcting Search to Accelerate Materials Discovery ↗
A model's predicted material property is not a physical measurement; this is not text-LLM research.
Verbalized Eval Awareness Inflates Measured Safety ↗
Correlations span more models than the causal intervention study. Silence about a test does not prove absence of awareness.