Clinic 13
First-Model Defense
Defend a ten-class digit classifier with a baseline, a clear split, and an error inspection.
Situation¶
You trained a first classifier on scikit-learn's load_digits dataset: 1,797 small images with ten digit labels. For this exercise the validation set contains 360 images from a random split. The packet reports predictions for that held-out split and an empirical majority-class baseline.
Artifact Packet¶
These counts are illustrative, not output from a newly trained model in the browser. The runner calculates both accuracies from the counts. It does not substitute a binary AUC example for the ten-class task.
| validation rows | correct predictions | largest class count | model accuracy | majority-class accuracy |
|---|---|---|---|---|
| 360 | 350 | 37 | 0.9722 | 0.1028 |
Decision Prompt¶
- How does the model compare with the majority-class baseline?
- What error information is missing from an overall accuracy?
- What deployment assumptions does a random split make?
- What must be recorded so another learner can reproduce the baseline?
Strong Reasoning Looks Like¶
- report 350/360 correct and compare with the 37/360 majority baseline
- inspect a ten-by-ten confusion matrix and per-class precision/recall
- inspect the ten errors, including their true and predicted labels
- record the split seed, preprocessing fitted on training data, model settings, and runtime
Run The Clinic In Browser¶
The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.
Reference Reveal¶
Open after writing your note
The model clearly beats the majority baseline **on this illustrative split**: accuracy is about 0.9722 versus 0.1028. That is a useful first result, not evidence of production readiness. Request the actual confusion matrix, class supports, and misclassified images before making claims about which digits work. The aggregate counts cannot recover those details. Assess variation across suitable repeats or an appropriate uncertainty estimate; ten errors provide limited evidence about rare failure types. A random image split estimates performance for a similar image population. If deployment requires new writers or new scanners, obtain the corresponding group/source metadata and design validation around that shift. Do not infer the absence of writer dependence from the lack of a group field in a convenience loader.What To Do Next¶
After this clinic:
- open Honest Splits and Baselines — the systematic treatment of what you just defended
- open Public/Private Restraint — the grown-up version of this clinic for competition settings
- open Evaluation Metrics Deep Dive — for when accuracy is not the right primary
- rerun the code with
random_state0–19 and watch the score span; the banded view is the real answer to "is 0.972 good?"