Write a fictional scenario, assign costs before testing, and defend the threshold with per-case evidence.
Name the expected effect, metric and what would count against the hypothesis.
Record fixtures, source, version, parameters, units, controls and how cases are split.
| Case and split | Baseline result | Changed result | Interpretation |
|---|---|---|---|
A no-alert policy can look accurate in a rare-target dataset while recall is zero and precision is undefined.
Record your new case and rerun the original cases after redesign.
Explain the mechanism, one exact result and what would overturn your conclusion.
Costs are scenario assumptions chosen before evaluation, not objective values of people.
Next test: _ . Project filename: _ . Work that is mine and tools I used: ____ .
Use fictional data. Download a resumable project before changing devices. On a shared device, turn remembering off and clear your work when finished.