Skip to content

Clinic 08

Threshold Under Asymmetric Cost

A missed fraud costs 100 times as much as a false alarm. Compare operating points using their error costs.

Situation

A fraud system assigns scores that are not assumed calibrated. Missing a fraud costs $10,000; a false alarm costs $100. Correct decisions have zero additional cost in this exercise. Choose among the five tested thresholds using validation data, then evaluate the locked policy on untouched test data.

Artifact Packet

This fixed illustrative packet has a 5% fraud rate: 50 positives and 950 negatives per 1,000 transactions. Fractional counts represent rates scaled from a larger evaluation set. Every metric is calculated from the same TP, FN, FP, and TN counts.

threshold precision recall FPR accuracy cost per 1,000
0.50 0.706 0.610 0.013 0.9678 $196,270
0.30 0.478 0.790 0.045 0.9463 $109,320
0.15 0.265 0.910 0.133 0.8692 $57,630
0.10 0.172 0.950 0.240 0.7694 $47,810
0.05 0.089 0.980 0.527 0.4986 $60,040

Underlying confusion counts:

threshold TP FN FP TN
0.50 30.5 19.5 12.7 937.3
0.30 39.5 10.5 43.2 906.8
0.15 45.5 4.5 126.3 823.7
0.10 47.5 2.5 228.1 721.9
0.05 49 1 500.4 449.6

cost = 10,000 × FN + 100 × FP. Accuracy is (TP + TN) / 1,000.

Decision Prompt

  1. Which tested threshold minimizes cost?
  2. Why does the accuracy winner lose on cost?
  3. Why does lowering the threshold from 0.10 to 0.05 increase cost?
  4. What changes if the scores become calibrated probabilities?

Strong Reasoning Looks Like

  • compare the stated costs across the tested grid
  • distinguish empirical score tuning from a population rule for calibrated probabilities
  • check prevalence, calibration, review capacity, and cost changes before deployment

Run The Clinic In Browser

The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.

Validate Your Decision In Browser

The validator checks the original packet's grid choice; review the written explanation separately.

Reference Reveal

Open after writing your note Choose **0.10**, costing **$47,810 per 1,000**: $25,000 from missed fraud and $22,810 from false alarms. At 0.50, accuracy is 0.9678 but cost is $196,270. At 0.05, the extra false alarms outweigh the savings from fewer misses. These are conclusions about the tested grid, not a proof of a global optimum. If `p` is a calibrated fraud probability for the deployment population, predicting fraud costs `100(1-p)` in expectation and predicting legitimate costs `10,000p`. Under these assumptions, choose fraud when `p ≥ 100 / 10,100 ≈ 0.00990`. That probability rule does not apply directly to arbitrary model scores. Capacity limits, additional costs, or a changed population require a revised decision rule. Lower false-negative costs or higher false-positive costs favor a higher calibrated-probability threshold. Verify changes empirically on validation data.

What To Do Next

After this clinic:

  1. open Calibration and Thresholds
  2. run the matching threshold demo example
  3. use Imbalanced Triage and Review Budgets for the full cost-aware workflow