Fit a calibration parameter on one split, choose it using validation, then reveal the untouched final split.
Name the expected effect, metric and what would count against the hypothesis.
Record fixtures, source, version, parameters, units, controls and how cases are split.
| Case and split | Baseline result | Changed result | Interpretation |
|---|---|---|---|
Temperature can improve Brier score without changing any predicted class; calibration can fail again on a shifted population.
Record your new case and rerun the original cases after redesign.
Explain the mechanism, one exact result and what would overturn your conclusion.
A few cases cannot establish population calibration. Final rows must remain hidden until the choice is locked.
Next test: _ . Project filename: _ . Work that is mine and tools I used: ____ .
Use fictional data. Download a resumable project before changing devices. On a shared device, turn remembering off and clear your work when finished.