
EE 641 - Unit 4B
Fall 2026
Sequences and Memory · Filters and State · The Recurrent Unit · Training Through the Loop · Gated Units · Recurrent Units as Components
[Encoder-Decoder] K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder–decoder for statistical machine translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, 2014, pp. 1724–1734.
[Seq2Seq] I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in Neural Information Processing Systems, 2014, pp. 3104–3112.
[BLEU] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: A method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 2002, pp. 311–318.
[Exposure] S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Advances in Neural Information Processing Systems, 2015, pp. 1171–1179.
[Decoding] P. Koehn and R. Knowles, “Six challenges for neural machine translation,” in Proceedings of the First Workshop on Neural Machine Translation, 2017, pp. 28–39.
[Length] K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder–decoder approaches,” in Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation, 2014, pp. 103–111.
[Captioning] O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3156–3164.
[Alignment] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in International Conference on Learning Representations, 2015.

One pair, used throughout
Three mismatches in one sentence
Across languages
Outputs of a unit reading the source
Against the pair
Required output

Object to build
\[p(\mathbf{y} \mid \mathbf{x}), \qquad \mathbf{x} = (x_1, \ldots, x_S),\ \mathbf{y} = (y_1, \ldots, y_T)\]
Chain rule, exact
\[p(\mathbf{y} \mid \mathbf{x}) = \prod_{t=1}^{T} p(y_t \mid y_1, \ldots, y_{t-1}, \mathbf{x})\]
On the example
Two conditioning inputs in every factor
Both are sequences of unbounded length

Reader
Writer
Reading the source, writing the target, choosing the words at test time, and training the pair from data are the four parts of the model.
Parallel corpus
Loss
\[L = -\sum_{n} \log p(\mathbf{y}^{(n)} \mid \mathbf{x}^{(n)}) = -\sum_{n}\sum_{t} \log p(y^{(n)}_t \mid y^{(n)}_{<t}, \mathbf{x}^{(n)})\]
Not given
Scale
Precision of each n-gram order
\[p_n = \frac{\text{candidate } n\text{-grams also in the reference, each counted at most as often as it appears there}}{\text{candidate } n\text{-grams}}\]
Score
\[\text{BLEU} = \text{BP} \cdot \exp\!\left(\tfrac{1}{4}\sum_{n=1}^{4} \log p_n\right), \qquad \text{BP} = \min\!\left(1,\ e^{\,1 - r/c}\right)\]
Published anchors, WMT’14 English to French
Against the reference “le chat noir s’est assis sur le tapis”
| Candidate | \(p_1\) | \(p_2\) | \(p_3\) | \(p_4\) | BP | BLEU |
|---|---|---|---|---|---|---|
| le chat noir s’est assis sur le tapis | 8/8 | 7/7 | 6/6 | 5/5 | 1 | 100 |
| le chat noir est assis sur le tapis | 7/8 | 5/7 | 3/6 | 1/5 | 1 | 50.0 |
| le chat noir s’est | 4/4 | 3/3 | 2/2 | 1/1 | 0.37 | 36.8 |
| un chat noir était assis sur un tapis | 5/8 | 2/7 | 0/6 | 0/5 | 1 | 0 |
Limits

Reader
\[\mathbf{h}^{\text{enc}}_s = f\!\left(\mathbf{h}^{\text{enc}}_{s-1},\ \mathbf{E}\,x_s\right), \quad s = 1, \ldots, S, \qquad \mathbf{c} = \mathbf{h}^{\text{enc}}_S\]
Against tagging
Size
Writer’s input
Memory demand
Measured, with the writer’s inputs given


One change to the input
Effect
Published result
Two readers, one summary
\[\overrightarrow{\mathbf{h}}_s = f_\rightarrow(\overrightarrow{\mathbf{h}}_{s-1}, \mathbf{E}x_s), \qquad \overleftarrow{\mathbf{h}}_s = f_\leftarrow(\overleftarrow{\mathbf{h}}_{s+1}, \mathbf{E}x_s)\]
\[\mathbf{c} = \left[\overrightarrow{\mathbf{h}}_S;\ \overleftarrow{\mathbf{h}}_1\right] \in \mathbb{R}^{2H}\]
Unchanged
Keeping every state

Captioning
Same writer, same training
Summary

One step of the writer
One connection added to the recurrent unit
Cost per word
Once, as the initial state
\[\mathbf{h}_0 = \mathbf{c}, \qquad \mathbf{h}_t = f(\mathbf{h}_{t-1}, \mathbf{E}'y_{t-1})\]
At every step, as an input
\[\mathbf{h}_t = f\!\left(\mathbf{h}_{t-1}, [\mathbf{E}'y_{t-1};\ \mathbf{c}]\right)\]

Same limit in both

Teacher forcing
\[L = -\sum_{t=1}^{T} \log p(y_t \mid y_{<t}, \mathbf{x})\]
Effect on the gradient
One connection differs
Same writer, two prefixes
| Step | Trained on (teacher forcing) | Generating, after one error |
|---|---|---|
| 1 | <sos> |
<sos> |
| 2 | le | le |
| 3 | le chat | le chien |
| 4 | le chat noir | le chien noir |
| 5 | le chat noir s’est | le chien noir … |
Exposure bias
Longer outputs, more exposure
Same mechanism wherever a model generates

Measured
Two treatments
| Token | Where | Role |
|---|---|---|
<sos> |
first writer input | starts the loop with no word chosen yet |
<eos> |
last target word | the writer ends the sentence by emitting it |
<pad> |
after <eos> in a batch |
fills shorter targets to the batch length, masked out of the loss |
<pad> is never predicted and never scoredLength set by the writer
<eos> is chosen, so the target length is not fixed in advance<eos> is never emitted<eos>, so emitting it at the right point is part of what is learnedOn the example
<eos> in place
Test-time search
\[\hat{\mathbf{y}} = \arg\max_{\mathbf{y}} \prod_{t} p(y_t \mid y_{<t}, \mathbf{x}) = \arg\max_{\mathbf{y}} \sum_t \log p(y_t \mid y_{<t}, \mathbf{x})\]
No word can be chosen alone
Cost
Found
Missed

Hypotheses kept open
Procedure, width \(k\)
<sos>, score 0<eos> or reached the maximum length, then return the bestCost
Found
Length and the score
<eos>, scores an early <eos> higherLength normalization
\[\text{score}(\mathbf{y}) = \frac{1}{T^{\alpha}} \sum_{t=1}^{T} \log p(y_t \mid y_{<t}, \mathbf{x})\]
Width
Beyond the search

Measured
Cause
Source of the loss

Two ragged sides
<pad>Cost of the padding
Sorting by length
Target side: mask the loss
\[L = -\frac{1}{\sum_{n,t} m_{n,t}} \sum_{n}\sum_{t=1}^{T_{\max}} m_{n,t}\, \log p(y_{n,t} \mid y_{n,<t}, \mathbf{x}_n), \qquad m_{n,t} = [t \le T_n]\]
<pad> after <eos>, and the loss of a short pair is diluted by its paddingSource side: take the summary at the right step

State carry-over, the general form
\[\mathbf{h}_{n,t} = m_{n,t}\, f(\mathbf{h}_{n,t-1}, \mathbf{x}_{n,t}) + (1 - m_{n,t})\, \mathbf{h}_{n,t-1}\]
What packing does
h_n holds each row’s state at its own last real step, and the mask on the state is no longer needed
Cost
| Value | |
|---|---|
| Reader and writer | LSTM, 4 layers each, 1000 cells per layer |
| Embeddings | 1000-dimensional, source and target |
| Vocabulary | 160,000 source words, 80,000 target words, the rest mapped to <unk> |
| Parameters | 384 million, of which 64 million are recurrent weights |
| Training data | 12 million pairs, 348 million French words, 304 million English words |
| Source order | reversed |
| Optimization | SGD without momentum, learning rate 0.7, halved every half epoch after epoch 5, 7.5 epochs |
| Batches | 128 pairs, sorted by length within windows |
| Gradient clipping | norm clipped at 5 |
| Hardware and time | 8 GPUs, one layer per GPU, about 10 days |
Parameters
Time
| System, WMT’14 English to French | BLEU |
|---|---|
| Single LSTM, source in order, beam 12 | 26.2 |
| Single LSTM, source reversed, beam 12 | 30.6 |
| Ensemble of 5 reversed LSTMs, beam 2 | 33.0 |
| Ensemble of 5 reversed LSTMs, beam 12 | 34.8 |
| Phrase-based baseline of the task | 33.3 |
| Best published WMT’14 result that year | 37.0 |
Row by row
Against the baselines
Metric

Measured
Published
Same model, longer input
Demands on the summary
Two demands, on both sides
Remedies
Limit

Per-word accuracy
Measured
Capacity

Per step
Fixed summary
Open question

Measured at \(S = 24\)
| \(H\) | Recurrent parameters | Mean per-word accuracy |
|---|---|---|
| 64 | 50,176 | 0.854 |
| 128 | 165,888 | 0.972 |
| 256 | 593,920 | 0.994 |
Cost
Unchanged by width

One summary per target word
\[\mathbf{c}_t = \sum_{s=1}^{S} a_{t,s}\,\mathbf{h}^{\text{enc}}_s, \qquad a_{t,s} = \frac{\exp(e_{t,s})}{\sum_{s'} \exp(e_{t,s'})}, \qquad e_{t,s} = \left(\mathbf{h}^{\text{dec}}_t\right)^{\!\top} \mathbf{h}^{\text{enc}}_s\]
Added

Measured at \(S = 24\), \(H = 64\)
Learned weights
Two paths from a target word back to the source
\[\frac{\partial L_t}{\partial \mathbf{h}^{\text{enc}}_s} = \underbrace{a_{t,s}\,\frac{\partial L_t}{\partial \mathbf{c}_t}}_{\text{through the read}} + \ \underbrace{\frac{\partial L_t}{\partial \mathbf{h}^{\text{enc}}_S}\,\frac{\partial \mathbf{h}^{\text{enc}}_S}{\partial \mathbf{h}^{\text{enc}}_s}}_{\text{through the recurrence}} \;+\; \ldots\]
Training the weights
Cost per sentence pair
Unchanged
Read over every position, everywhere