Encoder-Decoder And Machine Translation¶
What This Is¶
An encoder-decoder model maps an input sequence to an output sequence through two modules: an encoder that reads the whole input into a set of contextual representations, and a decoder that generates the output token by token, conditioned on the encoder states and on what it has already produced.
The practical lesson is that encoder-decoder is not just a translation architecture. It is a natural fit for sequence-to-sequence tasks such as translation, summarization, generated question answering over context, and image captioning. Decoder-only and other architectures can solve some of the same tasks, so architecture remains an empirical choice.
When You Use It¶
- machine translation
- summarization — long article → short summary
- question-answering — context + question → answer span or generated answer
- image captioning — vision encoder → text decoder
- speech recognition — audio encoder → text decoder
- any task where the output length is not determined by the input length and must be generated one token at a time
Core Structure¶
source tokens → encoder → contextual representations (one per source position)
target tokens so far → decoder → next-token distribution
attends to encoder states + its own past
Two forms:
- RNN encoder-decoder — early systems used a fixed vector (Sutskever et al.); later attention-based systems aligned decoder states to encoder states (Bahdanau et al.)
- Transformer encoder-decoder (Vaswani et al.) — self-attention within each side, cross-attention from decoder to encoder; widely used sequence-to-sequence architecture
See Attention and Transformers for the attention mechanics that make both work.
Teacher Forcing — How Training Works¶
During training the decoder is given the ground-truth previous tokens as input, not its own predictions. This is teacher forcing. The loss at each position is cross-entropy against the true next token.
input to decoder: <BOS> y_1 y_2 ... y_{T-1}
target: y_1 y_2 y_3 ... y_T
Teacher forcing lets a Transformer compute losses for all target positions in parallel; an RNN decoder still processes recurrent steps sequentially. Exclusive teacher forcing also creates a train/inference mismatch called exposure bias: at inference, the decoder conditions on its own earlier predictions. Scheduled sampling and sequence-level objectives are possible interventions, but each changes optimization and must be evaluated.
See Language Modeling Fundamentals for the underlying next-token-prediction framing.
Decoding At Inference Time¶
The decoder is autoregressive: it predicts one token, appends it to the context, and predicts the next. Strategies to pick each token:
| Strategy | Rule | Good for |
|---|---|---|
| Greedy | argmax at every step | fast, deterministic baseline; can miss better sequence-level hypotheses |
| Beam search | keep top-k partial hypotheses |
translation, summarization — the historical default |
| Sampling | draw from the distribution | open-ended generation |
| Top-k sampling | sample from top k tokens, renormalized |
creative writing, chat |
| Top-p / nucleus | sample from smallest set with total prob ≥ p |
modern default for open-ended text |
| Temperature | scale logits by 1/T before softmax |
lower T = sharper, higher T = more random |
Beam search remains a useful translation baseline, while sampling is appropriate when output diversity matters. Tune beam width, length penalty, temperature, and sampling filters on task-specific evaluation rather than treating one decoder as universally best. See Prompting and Tool Use.
Evaluation Metrics¶
Sequence generation has no single right answer, so evaluation is harder than classification.
| Metric | Measures | Use for |
|---|---|---|
| BLEU | n-gram precision against references | translation — standard baseline |
| chrF | character-level F-score | translation for morphologically rich languages |
| ROUGE | n-gram recall against references | summarization |
| BERTScore | embedding-space similarity to references | semantic similarity beyond n-grams |
| COMET | learned metric trained on human judgments | translation — strong learned metric; pin the model/version (original paper) |
| perplexity | exponentiated average token negative log-likelihood | intrinsic fit when tokenization and evaluation setup match |
No metric is a substitute for reading outputs. A BLEU of 42 says the n-grams overlap with the reference; it does not say the sentence is correct.
Length Control And /¶
Two special tokens hold the sequence-generation plumbing together:
<BOS>(beginning of sequence) — primes the decoder<EOS>(end of sequence) — the decoder emits this to stop; without it the decoder generates forever- practical systems also impose a
max_new_tokenssafety cap
At inference, generation stops when the decoder emits <EOS> or the cap is hit. Rare <EOS> emission can come from data formatting, masking/label bugs, length bias, decoding settings, or a mismatch in training examples.
What To Inspect¶
- decoder cross-attention patterns as a debugging signal for masks and indexing—not a requirement that every head align with human word correspondences
- whether
<EOS>is produced at a reasonable rate - output length distribution vs. reference length distribution — a systematic shortening or lengthening is a training-signal bug
- whether the error pattern lives in translation quality (word choice) or in fluency (syntax)
- BLEU and a small read of 20 examples — never ship on metric alone
- teacher-forced token loss plus free-running sequence metrics; their difference is not a single directly comparable "loss gap"
Failure Pattern¶
Shipping a model that rarely emits <EOS> and relying on max_new_tokens to truncate. Inspect target construction, loss masking, decoding length bias, and training length distribution before choosing an intervention.
A second failure pattern is selecting greedy or beam decoding by convention without comparing task quality, length bias, and latency on held-out translations.
A third failure pattern is comparing two models on BLEU alone and picking the higher one. Estimate uncertainty with paired resampling or repeated evaluation where appropriate, and inspect a fixed output diff set.
A fourth failure pattern is evaluating generation quality on the training distribution. Exposure bias shows only at inference on fresh inputs, where the decoder must condition on its own output.
Quick Checks¶
- Is teacher forcing wired in training, and free-running wired at inference?
- Is the decoder producing
<EOS>at a reasonable rate? - Was the decoding strategy selected on held-out task quality, length behavior, and latency?
- Is the cross-attention alignment sensible on a sample?
- Is the output length distribution matched to the reference length distribution?
Practice¶
- Implement a tiny RNN encoder-decoder on a character-level reversal task (input → reversed).
- Replace the RNN decoder with a greedy transformer decoder and compare sample quality.
- Add beam search with
k = 1, 4, 16and compare BLEU on a small translation set. - Lower decoding temperature from 1.0 to 0.2 on an open-ended generation task and observe the sample distribution.
- Compute teacher-forced token loss and free-running sequence metrics on the same validation set. Explain what each reveals and why they are not directly comparable.
- Explain why cross-attention is what makes the decoder "look at" the encoder.
- Describe one case where BLEU disagrees with human judgment.
- State why
<EOS>is a trained decision and not a post-processing step. - Explain the relationship between temperature and top-p and when they stack.
- Describe one systematic error pattern you would expect in a low-resource language pair.
Runnable Example¶
Run the matching lab from the repository root:
.venv/bin/python labs/encoder-decoder-translation/src/encoder_decoder_workflow.py
Inspect the teacher-forcing curve, decoded sequences, and EOS positions. A falling training loss is not enough if free-running decoding repeats or never terminates.
Longer Connection¶
Encoder-decoder sits next to:
- Attention and Transformers — the architecture that powers the modern encoder-decoder stack
- Language Modeling Fundamentals — the next-token-prediction framing shared with decoder-only models
- Tokenization Mechanics — the sub-word tokenizers that keep open-vocabulary sequence generation tractable
- Text Generation and Language Models — the broader generation workflow including decoder-only models
- Recurrent Networks and Sequences — the sequence models the modern architecture replaced
Encoder-decoder is one powerful shape for sequence-in, sequence-out problems. The decisions—architecture, attention pattern, decoding strategy, tokenization, and metric—are where the real engineering lives.