ConvS2S
2017Convolutional sequence-to-sequence architecture using stacked CNNs with gated linear units for both encoder and decoder instead of recurrence, offering greater training parallelism than RNN-based translation models (Gehring, Auli, Grangier, Yarats & Dauphin, "Convolutional Sequence to Sequence Learning," arXiv:1705.03122, 2017). The other convolutional architecture the Transformer directly benchmarks against.
Originators
- Gehring, J.
- Auli, M.
- Grangier, D.
- Yarats, D.
- Dauphin, Y.N.
Landmark Paper
Checked 2026-09-19 — interim signal only, see docs/BASIC_ROADMAP.md Phase 10
Connections
- competes with Transformerbasis: reasoned
Table 2 of Vaswani et al. 2017 benchmarks against Gehring et al.'s ConvS2S (ref [9]), reporting the Transformer (big) at 28.4 BLEU / 2.3e19 FLOPs vs. ConvS2S Ensemble's 26.36 BLEU / 7.7e19 FLOPs on English-to-German -- better quality at roughly a third of the training cost.