Confidence meets its evidence

Does the stated probability match unseen outcomes?

Make your own investigation

Fit a calibration parameter on one split, choose it using validation, then reveal the untouched final split.

Transform each logit by 1/temperature and report per-case squared probability error, grouped by split.

Case cards

Cut these cards apart. Use the original example first, then let a partner supply a changed case. Keep sealed/final cards with the teacher until the learner commits a rule.

Case 1

Probability: 0.8

Outcome: 1

Split: fit

My prediction and reason:

Case 2

Probability: 0.8

Outcome: 0

Split: fit

My prediction and reason:

Case 3

Probability: 0.6

Outcome: 1

Split: validation

My prediction and reason:

Case 4

Probability: 0.4

Outcome: 0

Split: validation

My prediction and reason:

Case 5

Probability: 0.9

Outcome: 1

Split: final

My prediction and reason:

Case 6

Probability: 0.7

Outcome: 0

Split: final

My prediction and reason:

What this cannot establish

A few cases cannot establish population calibration. Final rows must remain hidden until the choice is locked.