Skip to content

Generative Adversarial Networks

What This Is

A Generative Adversarial Network (GAN) is a pair of neural networks trained against each other:

  • a generator G: z → x̂ that maps random noise to synthetic samples
  • a discriminator D: x → p(real) that tries to tell real samples from synthetic ones

Training alternates between improving the discriminator and improving the generator. Under the idealized objective, unlimited capacity, and successful optimization, the data distribution is an equilibrium; finite GAN training need not reach it (Goodfellow et al., 2014).

GANs drove major progress in image generation and remain useful when fast sampling or a particular adversarial objective matters. Diffusion and autoregressive models are strong alternatives, especially for broad conditional generation. This topic teaches when a GAN fits and how to spot its characteristic training failures.

When You Use It

  • you need fast inference — GAN generation is a single forward pass, unlike diffusion's many
  • you want sharp, high-frequency samples (textures, faces, specific photo styles)
  • the task is narrow, the dataset is clean, and a validated GAN baseline meets the quality/latency trade-off
  • you need a learned discriminator as part of a larger objective (perceptual losses, cycle losses for image translation)

Do Not Use It When

  • the data is heterogeneous or long-tailed — GANs struggle with mode coverage
  • you need controllable, text-conditional generation — diffusion is usually better
  • training stability is a first-order concern — GAN training is famously touchy; prefer diffusion or VAEs if you cannot tolerate long debugging cycles
  • likelihoods or density estimates are needed — GANs do not estimate them

The Minimax Objective

The classical GAN loss is:

min_G  max_D   E_{x ~ p_data}[ log D(x) ]  +  E_{z ~ p_z}[ log (1 - D(G(z))) ]

In plain terms: the discriminator wants to output 1 on real and 0 on fake; the generator wants the discriminator to output 1 on its fakes.

In practice, the log(1 - D(G(z))) term gives vanishing gradients when the discriminator is strong, so the generator is trained with the "non-saturating" loss instead:

L_G = -E[ log D(G(z)) ]

The non-saturating objective is a common practical choice because it provides a stronger early generator gradient, but other GAN objectives are also used.

A Minimal GAN In PyTorch

A barebones DCGAN-style skeleton:

import torch
from torch import nn
import torch.nn.functional as F

class Generator(nn.Module):
    def __init__(self, z_dim=64, out_dim=784):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(z_dim, 256), nn.ReLU(),
            nn.Linear(256, 512), nn.ReLU(),
            nn.Linear(512, out_dim), nn.Tanh(),
        )
    def forward(self, z): return self.net(z)

class Discriminator(nn.Module):
    def __init__(self, in_dim=784):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(in_dim, 512), nn.LeakyReLU(0.2),
            nn.Linear(512, 256), nn.LeakyReLU(0.2),
            nn.Linear(256, 1),
        )
    def forward(self, x): return self.net(x)

def gan_step(x_real, G, D, opt_G, opt_D, z_dim=64):
    B = x_real.size(0)
    # --- Discriminator ---
    z = torch.randn(B, z_dim, device=x_real.device)
    x_fake = G(z).detach()
    d_real = D(x_real)
    d_fake = D(x_fake)
    loss_d = (F.binary_cross_entropy_with_logits(d_real, torch.ones_like(d_real))
            + F.binary_cross_entropy_with_logits(d_fake, torch.zeros_like(d_fake)))
    opt_D.zero_grad(); loss_d.backward(); opt_D.step()

    # --- Generator (non-saturating) ---
    z = torch.randn(B, z_dim, device=x_real.device)
    x_fake = G(z)
    d_fake = D(x_fake)
    loss_g = F.binary_cross_entropy_with_logits(d_fake, torch.ones_like(d_fake))
    opt_G.zero_grad(); loss_g.backward(); opt_G.step()
    return loss_d.item(), loss_g.item()

Points that matter:

  • use LeakyReLU(0.2) in the discriminator, plain ReLU in the generator — tradition that still holds up
  • the generator's last activation is Tanh and images are normalized to [-1, 1] — matching the range is a common bug
  • detach() the fake when training D — otherwise generator gradients are computed and stored unnecessarily during the discriminator update

Variants That Changed Things

variant what it changed why it matters
DCGAN (Radford) convolutional generator/discriminator, specific architecture choices influential convolutional baseline
WGAN / WGAN-GP Wasserstein objective; WGAN-GP replaces weight clipping with a gradient penalty alternative training signal and Lipschitz-constraint strategies
Spectral Normalization bounds discriminator Lipschitz constant standard stability aid
Progressive Growing / StyleGAN grow resolution over training, style-based generator high-res face generation
BigGAN large batch, class conditioning, truncation high-res conditional generation
CycleGAN unpaired image translation practical for style transfer
Pix2Pix paired image-to-image translation practical for conditional mapping

For a new GAN project, start from a reproduced baseline appropriate to the data, such as a spectral-normalized GAN or WGAN-GP, and change one stability choice at a time.

Mode Collapse

The defining GAN failure is mode collapse: the generator learns to produce one (or a few) samples that fool the discriminator reliably, and stops exploring the data distribution. Output diversity collapses; the model "wins" by being a narrow specialist.

Signatures:

  • a small number of distinct samples across many z draws
  • loss dynamics change while sample diversity falls; the exact loss signature depends on the objective
  • interpolation between z_1 and z_2 produces nearly the same sample

Fixes (not guarantees):

  • use the non-saturating loss, not the original saturating form
  • add spectral normalization or gradient penalty to stabilize the discriminator
  • use minibatch discrimination — the discriminator sees a batch at once and can penalize similarity across a batch
  • use historical buffers — occasionally train the discriminator against past generator outputs, not just the latest
  • consider unrolled GANs or other stabilization techniques

Evaluation

The same evaluation caveats as diffusion apply:

  • FID (Fréchet Inception Distance) is a common distribution-level metric; its value depends on feature extractor, sample count, preprocessing, and implementation (Heusel et al., 2017)
  • Inception Score is historical; do not rely on it alone
  • precision and recall (Kynkäänniemi et al.) — separately estimate sample quality and distribution coverage; mode collapse can appear as low recall with reasonable precision (paper)
  • human evaluation — catches task-relevant artifacts that the chosen feature metric may miss
  • classifier-based diagnostics — a classifier trained on real-vs-fake reveals what axes the generator has not yet matched

For mode-coverage diagnostics, report a precision/recall-style measure alongside FID rather than asking FID alone to separate quality from coverage.

What To Inspect

  • loss dynamics — interpret them under the chosen objective and compare them with fixed-sample panels; there is no universal healthy loss value or shape
  • gradient norms — large or vanishing gradients can diagnose instability, but do not uniquely predict mode collapse
  • sample diversity — fix a batch of random z and decode every N iterations; watch whether the samples are still distinct
  • precision-recall curves on real vs. generated features — the honest coverage signal
  • checkpoints over time — select with distribution metrics and fixed-sample panels rather than assuming the last checkpoint is best

Failure Pattern

The two classic failures:

  1. mode collapse — described above, most common
  2. training instability divergence — generator and discriminator losses oscillate wildly; training never settles. Often caused by: LR mismatch, overly aggressive discriminator, missing spectral norm or gradient penalty

A third, quieter failure: the discriminator becomes too good, too fast. Generator gradients can become unhelpful, especially with the saturating loss. The non-saturating objective, update balance, regularization, or a different GAN objective are experiments—not guaranteed fixes.

Common Mistakes

  • using the saturating log(1 - D(G(z))) loss in practice (use the non-saturating form)
  • matching learning rates between generator and discriminator when one should be lower (often two-time-scale, e.g., 2e-4 for D, 1e-4 for G with Adam)
  • stopping too early — GANs can go through long plateaus before gaining quality
  • judging quality by visual cherry-picks without a quantitative metric
  • comparing FID values computed with small or unequal sample counts without estimating sampling variability
  • not setting up a fixed-z eval panel for comparing checkpoints (plus separate fresh samples for coverage)
  • training on heavily class-imbalanced data without a plan — GANs magnify imbalance into mode collapse

Decision: GAN vs. Diffusion vs. VAE vs. AR

option samples speed training ease likelihood
GAN can be sharp; coverage varies fast (single forward pass) often sensitive no explicit likelihood
diffusion strong quality and coverage on many image tasks iterative, but reducible by distillation/fewer-step samplers substantial compute commonly evaluated with a variational bound or related objective
VAE reconstruction quality depends on decoder and objective fast comparatively direct objective variational lower bound
autoregressive strong conditional modeling serial generation stable likelihood training tractable token likelihood

A practical rule: if you need fast sampling and you can tolerate training pain, GAN. If you need controllable text-to-image at scale and can afford inference time, diffusion. If you need a latent space, VAE. If you need density, AR or diffusion.

Practice

  1. Train the minimal GAN above on MNIST with the non-saturating loss. Plot generator and discriminator loss side by side for 10 epochs.
  2. Swap to the saturating loss. Watch the generator gradient vanish. Compare loss curves.
  3. Add spectral normalization to the discriminator. Measure training stability across three seeds.
  4. Perturb the generator/discriminator update ratio or learning-rate ratio and measure whether diversity changes across several seeds; do not assume one setting deterministically causes collapse.
  5. Train once with WGAN-GP instead of the BCE loss. Compare a declared feature-space distance using equal real/generated sample counts and report its uncertainty; standard ImageNet-Inception FID is not automatically meaningful for MNIST digits.
  6. Run precision-and-recall on the same trained model against the training data. Identify whether the model has a precision problem, a recall problem, or both.

Runnable Example

Run the offline four-mode lab from the repository root:

.venv/bin/python labs/gan-workflow/src/gan_workflow.py

Compare vanilla GAN, LSGAN, and WGAN-GP with the same data. Inspect mode coverage and sample geometry together; discriminator loss by itself is not the result.

Longer Connection

GANs stand next to the other generative topics in the academy:

  • Autoencoders and VAEs — a likelihood-based latent-variable alternative with different reconstruction and sampling trade-offs
  • Diffusion Models — an iterative generative alternative with different quality, coverage, and latency trade-offs
  • Attention and Transformers — modern GAN generators and discriminators often use attention; StyleGAN-XL and derivatives push this

For the decision frame—GAN, diffusion, VAE, or autoregressive—the relevant axes include inference latency, training stability, coverage, likelihood needs, conditioning, and the data domain. Read Baseline-First Task Solving before picking; a simpler method may already meet the bar.