Density-Based Counting¶
When you have to count many overlapping, densely packed objects (crowds of people, cells under a microscope, chickens in a coop, trees from satellite imagery), object detection is often the wrong tool. Density-map regression is the right one.
What This Is¶
Instead of predicting bounding boxes, you predict a density map D with the same spatial shape as the input. Every pixel of D is a real number, and sum(D) = count. Each annotated point becomes a small Gaussian bump at that location; the model learns to regress the sum of those bumps.
Training target:
for each annotated point (x_i, y_i): add Gaussian_σ(x - x_i, y - y_i) to D
integral of D over the image ≈ number of annotated points
Loss: ||D_pred - D_target||² (MSE on the density map)
Inference count: n_hat = sum(D_pred)
This reframes "how many are there?" as a dense regression problem. The spatial structure of the density map also gives you where they are, for free.
When You Use It¶
- objects are dense, overlapping, or partially occluded (crowds, dense cell clusters, herds)
- bounding boxes are hard to define or annotate (fuzzy boundaries, highly variable scale)
- you want a count and a spatial distribution (heatmap of where objects are)
- the count can range from 0 to thousands per image — detection-then-count scales poorly at the high end
Do Not Use It When¶
- objects are sparse and well-separated — detection works fine and gives you boxes for free
- the task cares about individual identity (re-ID, tracking) — density maps do not preserve identity
- you need to report bounding boxes to a downstream consumer — density maps do not contain boxes
- the dataset is tiny (< 50 training images) — density regression needs spatial diversity to learn
Building The Target Map¶
For annotated points (x_1, y_1), ..., (x_N, y_N):
def build_density_target(points, H, W, sigma=4.0):
if not np.isfinite(sigma) or sigma <= 0:
raise ValueError("sigma must be positive and finite")
D = np.zeros((H, W), dtype=np.float32)
yy, xx = np.ogrid[:H, :W]
for (x, y) in points:
kernel = np.exp(-((yy - y)**2 + (xx - x)**2) / (2 * sigma**2))
D += kernel / kernel.sum() # unit discrete mass, including clipped edges
return D
The kernel width σ matters:
- fixed σ: simple; works when objects have similar size across the image
- adaptive σ: set σ proportional to the distance to the k-nearest neighbor point — tighter bumps in crowded regions, wider bumps in sparse ones (the MCNN / CSRNet approach)
Architecture: Encoder And Density Head¶
A convolutional encoder extracts spatial features, and a convolutional head predicts a one-channel density map. A pretrained encoder can be frozen for a first adaptation experiment or fine-tuned with the head. Freezing is a design choice; the original CSRNet is trained end to end.
The offline lab initializes three separate small CNNs from scratch and trains their encoders and heads together:
scalar model: CNN → global pooling → count
heatmap model: CNN → spatial head → center probabilities → connected components
density model: CNN → spatial head → density map → sum
The models share an architectural pattern, not weights. They are compared on the same held-out scenes. The heatmap comparator includes a learned head; connected-component counting is its post-processing step.
Loss And Metrics¶
Training loss. MSE on the density map is standard. Two common refinements:
- per-pixel weighting — weight foreground (nonzero density) pixels more than background, since most pixels are zero
- count-consistency term — add
|sum(D_pred) - sum(D_target)|to the loss to regularize the total count explicitly
Evaluation metrics. The task is usually scored on count, not density:
- MAE = mean absolute error on counts (lower is better)
- MSE = mean squared error on counts (penalizes big misses more)
- GAME (grid average mean error) = MAE computed on subdivisions of the image; penalizes getting the right count in the wrong place
What To Inspect¶
- predicted density map overlaid on the input image — does it light up where the objects are?
- predicted count vs true count scatter plot — is it unbiased, or systematically over/underestimating?
- count error binned by true count — the model might be good at small counts and bad at large ones (or vice versa)
- σ sweep — try σ ∈ {2, 4, 8, 16} and see how count MAE changes; there is usually a clear optimum
- background false-positive density — does the model put density on empty regions? That is pure overcount.
Failure Pattern¶
- target-width mismatch. A poorly chosen σ can make spatial regression harder. A correctly normalized Gaussian has unit mass at every width, so narrow bumps do not inherently undercount. Keep the target-construction rule consistent during an experiment and validate width choices without changing the count labels.
- resolution mismatch. When interpolating a density map for the same scene from an old grid to a new grid, multiply by
old_area / new_areato approximately preserve its sum: 256×256 to 1024×1024 requires a factor of 1/16. Check the discrete sum and renormalize if exact conservation is required. A new field of view or a different model output stride requires its own units analysis; there is no universal deployment area multiplier. - learning a smooth blob. The decoder learns a smooth low-frequency prior that roughly predicts average density everywhere. Count is approximately right, but the spatial distribution is garbage. Fix: inspect the density map, not just the count.
- class imbalance. Most pixels are zero; the model learns to predict zero everywhere. Fix: upsample crowded regions or weight the loss by density mass.
- point annotation error. Annotators dropped points; your targets undercount the real density. You cannot train past the annotation ceiling.
Quick Checks¶
- does
sum(D_target)equal the point count per image? (sanity-check the target builder) - is σ the same at train and eval?
- is the model's predicted count distribution centered on the true count distribution?
- does the density map look spatially correct, not just aggregate-correct?
- do the trainable parameters match the intended experiment? This lab trains encoder and head; a frozen-backbone variant must explicitly disable encoder gradients and fix its normalization statistics.
Practice¶
Run .venv/bin/python labs/density-map-counting/src/counting_workflow.py:
- generates a synthetic dataset of scenes with 5-80 small objects placed randomly (with overlap allowed)
- builds density targets with σ=4
- trains three separate CNNs end to end: scalar count regression, center-heatmap prediction followed by thresholding and connected-component counting, and density-map regression
- measures MAE per method and per count bin
- visualizes predicted vs true density maps and count-vs-count scatter
You should leave able to explain when density-map regression beats detection-and-count, why σ matters, and what to look at when your counts are right but your density map is wrong.
Runnable Example¶
From the repository root, run:
.venv/bin/python labs/density-map-counting/src/counting_workflow.py
Compare predicted-map mass with the true count, then inspect whether density is located near target points rather than merely summing to the right number.
Longer Connection¶
Density regression is one instance of a broader pattern: represent the target as a spatial distribution, not a discrete list. Instances:
- density maps — counting
- heatmap regression for keypoints (human pose estimation, landmark detection) — put a Gaussian at each keypoint, regress the heatmap
- segmentation masks — every pixel gets a class probability
- saliency maps — every pixel gets an importance score (see Saliency and Explainability)
All of these predict dense outputs over a spatial grid. Encoder-decoder models are a common starting point; freezing or fine-tuning the encoder is a separate experiment. Each task still needs suitable targets, losses, and evaluation.