Does the stated probability match unseen outcomes?
Fit a calibration parameter on one split, choose it using validation, then reveal the untouched final split.
Transform each logit by 1/temperature and report per-case squared probability error, grouped by split.
Cut these cards apart. Use the original example first, then let a partner supply a changed case. Keep sealed/final cards with the teacher until the learner commits a rule.
Probability: 0.8
Outcome: 1
Split: fit
My prediction and reason:
Probability: 0.8
Outcome: 0
Split: fit
My prediction and reason:
Probability: 0.6
Outcome: 1
Split: validation
My prediction and reason:
Probability: 0.4
Outcome: 0
Split: validation
My prediction and reason:
Probability: 0.9
Outcome: 1
Split: final
My prediction and reason:
Probability: 0.7
Outcome: 0
Split: final
My prediction and reason:
A few cases cannot establish population calibration. Final rows must remain hidden until the choice is locked.