Skip to content

Clinic 20

IoU Threshold Or NMS Tune

Post-processing changes predictions. Keep the evaluator fixed and inspect the slices that must not regress.

Situation

A shelf detector uses greedy non-maximum suppression (NMS). You lower its suppression IoU from 0.50 to 0.30, suppressing overlapping boxes more aggressively, while keeping score filtering and the evaluator fixed. The deployment agreement requires dense-shelf AP to remain at least 0.45 on the designated validation set.

Distinguish NMS IoU, which controls which predictions survive, from evaluation IoU, which determines matches to ground truth. Tune post-processing on validation data, then lock the complete prediction pipeline for final testing.

Artifact Packet

All rows are fixed illustrative measurements using the same ground truth and AP-at-IoU-0.50 evaluation protocol. The runner calculates the differences. Slice AP values cannot generally be averaged by image count to recover overall AP.

slice images AP, NMS IoU 0.50 AP, NMS IoU 0.30 AP change
sparse 600 0.78 0.80 0.02
medium 900 0.63 0.70 0.07
dense 500 0.45 0.42 -0.03
edge 200 0.41 0.44 0.03

Decision Prompt

  1. Does changing NMS change the predictions or the evaluator?
  2. Would you accept the new setting under the dense-slice requirement?
  3. What could explain the dense-shelf regression?
  4. Which experiment would you run next with the remaining time?

Run The Clinic In Browser

The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.

Reference Reveal

Open after writing your note Keep **NMS IoU 0.50** for now: the proposed 0.30 setting drops dense-shelf AP from 0.45 to 0.42 and fails the stated requirement. Inspect overlapping distinct objects and duplicate detections to determine what the more aggressive suppression removes. A gain from NMS can be a real system improvement without changing model weights. Here the reason to reject it is the critical-slice regression, not that post-processing is an invalid improvement. Changing the evaluator's matching threshold would be a different experiment and would require reporting metrics under a common protocol for comparison. A smaller NMS change, Soft-NMS, or a model/data change are candidates to validate, not guaranteed fixes. Some slices improving while another worsens is not sufficient to claim Simpson's paradox. AP50 and AP averaged over several IoU thresholds measure different localization requirements; disagreement alone does not establish a broken evaluation harness.

What To Do Next

  1. open Object Detection Basics for the NMS-and-matching-IoU distinction in depth
  2. open Detection and Segmentation for the architectural knobs that move dense-slice performance
  3. open Reliability Slices — the discipline this clinic depends on
  4. sweep NMS IoU on your own eval set and plot per-slice AP; one chart answers most of this clinic