Generative Adversarial Networks¶
What This Is¶
A Generative Adversarial Network (GAN) is a pair of neural networks trained against each other:
- a generator
G: z → x̂that maps random noise to synthetic samples - a discriminator
D: x → p(real)that tries to tell real samples from synthetic ones
Training alternates between improving the discriminator and improving the generator. Under the idealized objective, unlimited capacity, and successful optimization, the data distribution is an equilibrium; finite GAN training need not reach it (Goodfellow et al., 2014).
GANs drove major progress in image generation and remain useful when fast sampling or a particular adversarial objective matters. Diffusion and autoregressive models are strong alternatives, especially for broad conditional generation. This topic teaches when a GAN fits and how to spot its characteristic training failures.
When You Use It¶
- you need fast inference — GAN generation is a single forward pass, unlike diffusion's many
- you want sharp, high-frequency samples (textures, faces, specific photo styles)
- the task is narrow, the dataset is clean, and a validated GAN baseline meets the quality/latency trade-off
- you need a learned discriminator as part of a larger objective (perceptual losses, cycle losses for image translation)
Do Not Use It When¶
- the data is heterogeneous or long-tailed — GANs struggle with mode coverage
- you need controllable, text-conditional generation — diffusion is usually better
- training stability is a first-order concern — GAN training is famously touchy; prefer diffusion or VAEs if you cannot tolerate long debugging cycles
- likelihoods or density estimates are needed — GANs do not estimate them
The Minimax Objective¶
The classical GAN loss is:
min_G max_D E_{x ~ p_data}[ log D(x) ] + E_{z ~ p_z}[ log (1 - D(G(z))) ]
In plain terms: the discriminator wants to output 1 on real and 0 on fake; the generator wants the discriminator to output 1 on its fakes.
In practice, the log(1 - D(G(z))) term gives vanishing gradients when the discriminator is strong, so the generator is trained with the "non-saturating" loss instead:
L_G = -E[ log D(G(z)) ]
The non-saturating objective is a common practical choice because it provides a stronger early generator gradient, but other GAN objectives are also used.
A Minimal GAN In PyTorch¶
A barebones DCGAN-style skeleton:
import torch
from torch import nn
import torch.nn.functional as F
class Generator(nn.Module):
def __init__(self, z_dim=64, out_dim=784):
super().__init__()
self.net = nn.Sequential(
nn.Linear(z_dim, 256), nn.ReLU(),
nn.Linear(256, 512), nn.ReLU(),
nn.Linear(512, out_dim), nn.Tanh(),
)
def forward(self, z): return self.net(z)
class Discriminator(nn.Module):
def __init__(self, in_dim=784):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_dim, 512), nn.LeakyReLU(0.2),
nn.Linear(512, 256), nn.LeakyReLU(0.2),
nn.Linear(256, 1),
)
def forward(self, x): return self.net(x)
def gan_step(x_real, G, D, opt_G, opt_D, z_dim=64):
B = x_real.size(0)
# --- Discriminator ---
z = torch.randn(B, z_dim, device=x_real.device)
x_fake = G(z).detach()
d_real = D(x_real)
d_fake = D(x_fake)
loss_d = (F.binary_cross_entropy_with_logits(d_real, torch.ones_like(d_real))
+ F.binary_cross_entropy_with_logits(d_fake, torch.zeros_like(d_fake)))
opt_D.zero_grad(); loss_d.backward(); opt_D.step()
# --- Generator (non-saturating) ---
z = torch.randn(B, z_dim, device=x_real.device)
x_fake = G(z)
d_fake = D(x_fake)
loss_g = F.binary_cross_entropy_with_logits(d_fake, torch.ones_like(d_fake))
opt_G.zero_grad(); loss_g.backward(); opt_G.step()
return loss_d.item(), loss_g.item()
Points that matter:
- use
LeakyReLU(0.2)in the discriminator, plain ReLU in the generator — tradition that still holds up - the generator's last activation is
Tanhand images are normalized to[-1, 1]— matching the range is a common bug detach()the fake when trainingD— otherwise generator gradients are computed and stored unnecessarily during the discriminator update
Variants That Changed Things¶
| variant | what it changed | why it matters |
|---|---|---|
| DCGAN (Radford) | convolutional generator/discriminator, specific architecture choices | influential convolutional baseline |
| WGAN / WGAN-GP | Wasserstein objective; WGAN-GP replaces weight clipping with a gradient penalty | alternative training signal and Lipschitz-constraint strategies |
| Spectral Normalization | bounds discriminator Lipschitz constant | standard stability aid |
| Progressive Growing / StyleGAN | grow resolution over training, style-based generator | high-res face generation |
| BigGAN | large batch, class conditioning, truncation | high-res conditional generation |
| CycleGAN | unpaired image translation | practical for style transfer |
| Pix2Pix | paired image-to-image translation | practical for conditional mapping |
For a new GAN project, start from a reproduced baseline appropriate to the data, such as a spectral-normalized GAN or WGAN-GP, and change one stability choice at a time.
Mode Collapse¶
The defining GAN failure is mode collapse: the generator learns to produce one (or a few) samples that fool the discriminator reliably, and stops exploring the data distribution. Output diversity collapses; the model "wins" by being a narrow specialist.
Signatures:
- a small number of distinct samples across many
zdraws - loss dynamics change while sample diversity falls; the exact loss signature depends on the objective
- interpolation between
z_1andz_2produces nearly the same sample
Fixes (not guarantees):
- use the non-saturating loss, not the original saturating form
- add spectral normalization or gradient penalty to stabilize the discriminator
- use minibatch discrimination — the discriminator sees a batch at once and can penalize similarity across a batch
- use historical buffers — occasionally train the discriminator against past generator outputs, not just the latest
- consider unrolled GANs or other stabilization techniques
Evaluation¶
The same evaluation caveats as diffusion apply:
- FID (Fréchet Inception Distance) is a common distribution-level metric; its value depends on feature extractor, sample count, preprocessing, and implementation (Heusel et al., 2017)
- Inception Score is historical; do not rely on it alone
- precision and recall (Kynkäänniemi et al.) — separately estimate sample quality and distribution coverage; mode collapse can appear as low recall with reasonable precision (paper)
- human evaluation — catches task-relevant artifacts that the chosen feature metric may miss
- classifier-based diagnostics — a classifier trained on real-vs-fake reveals what axes the generator has not yet matched
For mode-coverage diagnostics, report a precision/recall-style measure alongside FID rather than asking FID alone to separate quality from coverage.
What To Inspect¶
- loss dynamics — interpret them under the chosen objective and compare them with fixed-sample panels; there is no universal healthy loss value or shape
- gradient norms — large or vanishing gradients can diagnose instability, but do not uniquely predict mode collapse
- sample diversity — fix a batch of random
zand decode every N iterations; watch whether the samples are still distinct - precision-recall curves on real vs. generated features — the honest coverage signal
- checkpoints over time — select with distribution metrics and fixed-sample panels rather than assuming the last checkpoint is best
Failure Pattern¶
The two classic failures:
- mode collapse — described above, most common
- training instability divergence — generator and discriminator losses oscillate wildly; training never settles. Often caused by: LR mismatch, overly aggressive discriminator, missing spectral norm or gradient penalty
A third, quieter failure: the discriminator becomes too good, too fast. Generator gradients can become unhelpful, especially with the saturating loss. The non-saturating objective, update balance, regularization, or a different GAN objective are experiments—not guaranteed fixes.
Common Mistakes¶
- using the saturating
log(1 - D(G(z)))loss in practice (use the non-saturating form) - matching learning rates between generator and discriminator when one should be lower (often two-time-scale, e.g.,
2e-4for D,1e-4for G with Adam) - stopping too early — GANs can go through long plateaus before gaining quality
- judging quality by visual cherry-picks without a quantitative metric
- comparing FID values computed with small or unequal sample counts without estimating sampling variability
- not setting up a fixed-
zeval panel for comparing checkpoints (plus separate fresh samples for coverage) - training on heavily class-imbalanced data without a plan — GANs magnify imbalance into mode collapse
Decision: GAN vs. Diffusion vs. VAE vs. AR¶
| option | samples | speed | training ease | likelihood |
|---|---|---|---|---|
| GAN | can be sharp; coverage varies | fast (single forward pass) | often sensitive | no explicit likelihood |
| diffusion | strong quality and coverage on many image tasks | iterative, but reducible by distillation/fewer-step samplers | substantial compute | commonly evaluated with a variational bound or related objective |
| VAE | reconstruction quality depends on decoder and objective | fast | comparatively direct objective | variational lower bound |
| autoregressive | strong conditional modeling | serial generation | stable likelihood training | tractable token likelihood |
A practical rule: if you need fast sampling and you can tolerate training pain, GAN. If you need controllable text-to-image at scale and can afford inference time, diffusion. If you need a latent space, VAE. If you need density, AR or diffusion.
Practice¶
- Train the minimal GAN above on MNIST with the non-saturating loss. Plot generator and discriminator loss side by side for 10 epochs.
- Swap to the saturating loss. Watch the generator gradient vanish. Compare loss curves.
- Add spectral normalization to the discriminator. Measure training stability across three seeds.
- Perturb the generator/discriminator update ratio or learning-rate ratio and measure whether diversity changes across several seeds; do not assume one setting deterministically causes collapse.
- Train once with WGAN-GP instead of the BCE loss. Compare a declared feature-space distance using equal real/generated sample counts and report its uncertainty; standard ImageNet-Inception FID is not automatically meaningful for MNIST digits.
- Run precision-and-recall on the same trained model against the training data. Identify whether the model has a precision problem, a recall problem, or both.
Runnable Example¶
Run the offline four-mode lab from the repository root:
.venv/bin/python labs/gan-workflow/src/gan_workflow.py
Compare vanilla GAN, LSGAN, and WGAN-GP with the same data. Inspect mode coverage and sample geometry together; discriminator loss by itself is not the result.
Longer Connection¶
GANs stand next to the other generative topics in the academy:
- Autoencoders and VAEs — a likelihood-based latent-variable alternative with different reconstruction and sampling trade-offs
- Diffusion Models — an iterative generative alternative with different quality, coverage, and latency trade-offs
- Attention and Transformers — modern GAN generators and discriminators often use attention; StyleGAN-XL and derivatives push this
For the decision frame—GAN, diffusion, VAE, or autoregressive—the relevant axes include inference latency, training stability, coverage, likelihood needs, conditioning, and the data domain. Read Baseline-First Task Solving before picking; a simpler method may already meet the bar.