ByteNet
2016Dilated-convolution sequence transduction model computing hidden representations in parallel across all input and output positions, reducing the operation count needed to relate distant positions to logarithmic in sequence length (Kalchbrenner et al., "Neural Machine Translation in Linear Time," arXiv:1610.10099, 2016). One of two convolutional architectures the Transformer explicitly compares itself against and outperforms.
Originators
- Kalchbrenner, N.
- Espeholt, L.
- Simonyan, K.
- van den Oord, A.
- Graves, A.
- Kavukcuoglu, K.
Landmark Paper
Checked 2026-09-19 — interim signal only, see docs/BASIC_ROADMAP.md Phase 10
Connections
- competes with Transformerbasis: reasoned
Table 2 of Vaswani et al. 2017 benchmarks against Kalchbrenner et al.'s ByteNet (ref [18]), reporting 23.75 BLEU on English-to-German vs. the Transformer base model's 27.3; Section 2 contrasts ByteNet's O(log n) path length for distant positions against the Transformer's O(1).