← Back to Basic Science

ConvS2S

2017
Computer Science (theoretical)Machine Learning TheoryFrameworkfoundational

Convolutional sequence-to-sequence architecture using stacked CNNs with gated linear units for both encoder and decoder instead of recurrence, offering greater training parallelism than RNN-based translation models (Gehring, Auli, Grangier, Yarats & Dauphin, "Convolutional Sequence to Sequence Learning," arXiv:1705.03122, 2017). The other convolutional architecture the Transformer directly benchmarks against.

Originators

  • Gehring, J.
  • Auli, M.
  • Grangier, D.
  • Yarats, D.
  • Dauphin, Y.N.

Landmark Paper

W2613904329 ↗
Not retracted (OpenAlex)

Checked 2026-09-19 — interim signal only, see docs/BASIC_ROADMAP.md Phase 10

Connections

  • competes with Transformer
    basis: reasoned

    Table 2 of Vaswani et al. 2017 benchmarks against Gehring et al.'s ConvS2S (ref [9]), reporting the Transformer (big) at 28.4 BLEU / 2.3e19 FLOPs vs. ConvS2S Ensemble's 26.36 BLEU / 7.7e19 FLOPs on English-to-German -- better quality at roughly a third of the training cost.