← Back to Basic Science

Dropout

2014
Computer Science (theoretical)Machine Learning TheoryAlgorithmfoundational

Regularization technique that randomly deactivates a fraction of network units on each training step to prevent co-adaptation and overfitting (Srivastava, Hinton, Krizhevsky, Sutskever & Salakhutdinov, "Dropout: A Simple Way to Prevent Neural Networks from Overfitting," JMLR 15(1), 2014). Became a near-universal default regularizer for training large neural networks.

Originators

  • Srivastava, N.
  • Hinton, G.E.
  • Krizhevsky, A.
  • Sutskever, I.
  • Salakhutdinov, R.

Landmark Paper

W2095705004 ↗
Not retracted (OpenAlex)

Checked 2026-09-19 — interim signal only, see docs/BASIC_ROADMAP.md Phase 10

Connections

  • is component of Transformer
    basis: reasoned

    Section 5.4 of Vaswani et al. 2017 specifies residual dropout (P_drop=0.1) applied to every sub-layer output and to the embedding sums, citing Srivastava et al.'s Dropout (ref [33]) directly.