Object Detection Basics¶
What This Is¶
Object detection goes beyond classification — instead of asking "what is in the image?", it asks "what is in the image, and where is it?" The model outputs bounding boxes with class labels and confidence scores.
The key shift: classification produces one label per image; detection produces multiple boxes per image, each with its own label.
When You Use It¶
- locating and counting objects in images
- building surveillance, autonomous driving, or quality inspection systems
- when you need to know where things are, not just what they are
- when multiple objects of different classes appear in a single image
Detection vs Classification vs Segmentation¶
| Task | Output | Granularity |
|---|---|---|
| Classification | one label per image | image-level |
| Object Detection | bounding boxes + labels | box-level |
| Semantic Segmentation | class label per pixel | pixel-level |
| Instance Segmentation | object mask per instance | pixel-level, instance-aware |
How Detection Works¶
Detector taxonomies overlap. Useful axes include proposal stages, anchor use, and prediction matching:
One-Stage Detectors (YOLO, SSD, RetinaNet)¶
- predict dense boxes/classes without a separate learned proposal stage
- often chosen for latency-sensitive systems, but speed depends on model size, input size, hardware, and implementation
- may be anchor-based or anchor-free; not every current one-stage detector uses predefined anchors
Two-Stage Detectors (Faster R-CNN, Mask R-CNN)¶
- First stage proposes candidate regions
- Second stage classifies and refines each region
- can offer different accuracy/latency trade-offs; it is not universally more accurate or slower than every one-stage model
Set-Prediction Detectors (DETR Family)¶
- use learned object queries and one-to-one matching between predictions and ground truth during training
- avoid the traditional anchor design and usually avoid standard NMS
- can be compared with one- and two-stage systems only under the same data, resolution, hardware, and evaluation protocol
Key Concepts¶
Bounding Boxes¶
A detection output is typically: [x_min, y_min, x_max, y_max, class_id, confidence]
Intersection over Union (IoU)¶
IoU measures how much a predicted box overlaps with the ground truth:
IoU = Area of Overlap / Area of Union
- IoU = 1.0: perfect match
- IoU ≥ a declared threshold can qualify a localization match, but correctness also requires the class to match and each ground-truth instance to be matched at most once
- IoU = 0.0: no overlap at all
Non-Maximum Suppression (NMS)¶
Detectors often produce multiple overlapping boxes for the same object. Traditional NMS is normally applied per class so overlapping boxes from different classes do not suppress each other:
- Sort boxes by confidence
- Keep the top box
- Remove remaining boxes of the same class with IoU above the threshold
- Repeat for remaining boxes
Mean Average Precision (mAP)¶
A detection evaluation must pin its protocol:
- For each class, sort predictions across evaluation images by confidence.
- Match each prediction to the highest-IoU unmatched ground-truth box of that class in the same image, subject to the declared IoU threshold. Duplicate detections are false positives.
- Build the precision-recall curve and integrate it to get average precision (AP).
- Average AP across the declared class set. COCO-style mAP additionally averages over IoU thresholds from 0.50 to 0.95; AP50 alone is a different number.
Common Architectures¶
| Model | Type | Speed | Accuracy | Best For |
|---|---|---|---|---|
| YOLO family | one-stage; recent variants often anchor-free | benchmark | benchmark | latency/throughput candidates |
| SSD | one-stage, anchor-based | benchmark | benchmark | established compact baseline |
| RetinaNet | one-stage, anchor-based | benchmark | benchmark | focal-loss baseline |
| Faster R-CNN | two-stage | benchmark | benchmark | proposal-based baseline |
| DETR family | query-based set prediction | benchmark | benchmark | end-to-end matching without traditional NMS |
Anchor Boxes¶
Many influential detectors use predefined anchor boxes at multiple scales and aspect ratios: - Small anchors catch small objects - Large anchors catch large objects - The model predicts offsets from these anchors, not absolute coordinates
Anchor-free one-stage detectors predict locations without predefined anchor templates. DETR-family models instead use learned object queries and set matching.
Failure Pattern¶
Training a detector on images where objects are always centered and large, then deploying on images with small, occluded, or densely packed objects. The model never learned to detect those cases.
Another failure is reporting “mAP” without the class set, IoU thresholds, area ranges, and maximum-detections convention. AP50 and COCO-style AP@[.50:.95] are not interchangeable.
Common Mistakes¶
- confusing detection confidence with classification accuracy
- applying class-agnostic NMS accidentally, or applying NMS to a model whose evaluation recipe does not use it
- training on unbalanced classes without techniques like focal loss
- evaluating mAP at only one IoU threshold when the task requires precise localization
Practice¶
- Explain the difference between one-stage and two-stage detectors and when each is preferable.
- Compute IoU between two bounding boxes by hand.
- Describe what NMS does and why it is necessary.
- Explain why mAP is preferred over simple accuracy for detection tasks.
- Compare YOLO and Faster R-CNN for a specific use case and justify your choice.
Runnable Example¶
Longer Connection¶
Continue with Convolutional Neural Networks for the backbone architectures that power detectors, and Vision Augmentation and Shift Robustness for making detection robust to real-world conditions.