An assistant that can show its working

Write answerability cases with expected decisions before running the policy. Include absent and contradictory evidence.

Hypothesis and criterion

Name the expected effect, metric and what would count against the hypothesis.

Method and reproducibility

Record fixtures, source, version, parameters, units, controls and how cases are split.

Paired evidence

Case and split Baseline result Changed result Interpretation

Counterexample and revised design

A high-scoring passage can be about the right topic while lacking the required fact; removing a source must change answerability.

Record your new case and rerun the original cases after redesign.

Individual defence

Explain the mechanism, one exact result and what would overturn your conclusion.

Remaining uncertainty and handover

This policy tests answerability logic, not natural-language entailment or generated answer correctness.

Next test: _ . Project filename: _ . Work that is mine and tools I used: ____ .

Keep your work

Use fictional data. Download a resumable project before changing devices. On a shared device, turn remembering off and clear your work when finished.