Clinic 21
Embedding Reuse Or Retrain
Compare frozen representations and adaptation using validation evidence, repeatability, and the actual maintenance budget.
Situation¶
A biomedical abstract classifier has 5,200 labels and 80,000 unlabeled in-domain abstracts. It starts from an off-the-shelf sentence encoder. The team can continue masked-language-model pretraining using a compatible token encoder and MLM head, then evaluate its pooled sentence representations with a linear probe. MLM adaptation does not automatically preserve sentence-embedding quality.
The release policy requires repeated validation evidence before replacing the frozen baseline. The remaining budget is 96 GPU-hours; two more 36-hour adaptation runs fit if the first run is already complete. Quarterly label updates will be maintained by a review coordinator, while encoder adaptation needs an engineering owner.
Artifact Packet¶
F1 values are illustrative means where multiple seeds are reported, or single-run measurements otherwise. All candidates use the same validation protocol. Seed SD is not a confidence interval and does not capture all uncertainty from the sampled cases.
| approach | training hours | validation F1 | repeat evidence |
|---|---|---|---|
frozen_ots_linear |
1 | 0.78 | 5 seeds; SD 0.01 |
frozen_domain_refit_linear |
36 | 0.82 | 1 seed; uncertainty unknown |
finetune_ots |
4 | 0.73 | 5 seeds; SD 0.04 |
finetune_domain_refit |
40 | 0.79 | 1 seed; uncertainty unknown |
Decision Prompt¶
- What would you release now, and what is the strongest challenger?
- How would you measure uncertainty in the difference?
- What does a distribution-distance diagnostic tell you?
- Which operations must the next maintainer be able to repeat?
Strong Reasoning Looks Like¶
- keep the tested frozen baseline available while checking the stronger one-run challenger
- compare paired validation predictions and repeated training runs, with uncertainty appropriate to the sampling unit
- inspect relevant topic slices, calibration, and operational cost
- separate cached-feature head retraining from computing new embeddings and adapting the encoder
Run The Clinic In Browser¶
The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.
Reference Reveal¶
Open after writing your note
Keep **`frozen_ots_linear` as the provisional release** under the stated repeat-evidence policy. `frozen_domain_refit_linear` is the strongest challenger: spend the available budget testing whether its observed gain repeats. Promote it if the gain and slice behavior justify the cost and the maintenance owner accepts the workflow. The packet does not contain those additional results. The lower end-to-end scores do not, by themselves, prove overfitting; inspect training/validation curves, tuning, and repeat runs. There is no universal label-count boundary at which full fine-tuning becomes optimal. MMD or another distribution-distance measure can reveal representation differences. A large change does not establish task improvement; a small aggregate distance does not prove a measured gain is noise. Use held-out task performance to decide.What To Do Next¶
- open Vision and Text Encoders for the freeze-vs-adapt ladder
- open Self-Supervised and Representation Learning — the MLM-refit recipe in detail
- open Freeze Or Fine-Tune? — the adjacent classical-transfer clinic
- measure the linear-probe gap and MMD on your own domain; if the gap is < 1 point, ship the frozen OTS and save the compute