Implicit regularization by stochastic gradient descent is the primary explanation for why over-parameterized neural networks generalize
Weak-link patterns
Brief §8's named failure modes, computed from this claim's own paper roles and status -- not a single weakness score. Patterns B and E aren't implemented yet (need dependency-graph machinery, docs/BASIC_ROADMAP.md Phase 11).
stored status is 'unknown_conflicting'
Open questions
Pattern G. Human-written, never generated -- deliberately no importance or priority score (docs/BASIC_ROADMAP.md's own precedent against inventing one, matching innovation_considerations).
Is SGD's implicit regularization the primary mechanism behind generalization in over-parameterized networks, or one of several redundant mechanisms (architecture, data augmentation, explicit regularization) that jointly suffice?
Motivated by: Pattern C (conflicting evidence -- direct contradiction plus a conceptual-replication failure)
Evidence map
Position shows direction, circle area shows citation count. Click a paper for the reason its role was assigned.
Evidence dimensions
Eight independent 0–5 judgments, each with a written rationale. Deliberately never summed into one score.
Real theoretical results exist for restricted settings (separable data, linear models), but nothing establishes implicit SGD regularization as THE primary explanation in the deep non-convex case the claim is about.
Sub-results are examined repeatedly, but there is no single canonical experiment to reproduce.
No same-paradigm independent confirmation that this mechanism is the primary one; supporting work explores different facets rather than re-testing one result.
Theory (margin/implicit-bias proofs), optimization experiments (batch size, minima sharpness) and large-scale measure comparisons have all been brought to bear - and they disagree.
The theory is rigorous but proved under assumptions far from practice; the empirical work is solid but correlational.
The account's central promise - a generalization measure that predicts held-out performance - is what Jiang et al. tested directly and largely did not find.
No consensus. Implicit bias, flat minima, margin, norm-based capacity and NTK/lazy-training accounts all remain live.
Many well-articulated competing mechanisms, none ruled out. The lowest score on this dimension in the corpus, and appropriately so.
Scored by human.
Why this status
The derivation rule, evaluated top to bottom, first match wins. This is the rule itself — not a confidence score standing in for one.
- 1Unknown — unexplored
5 paper rows (needs <3) and evidence_strength=2 (needs <=1); both required
- 2Unknown — conflicting evidencematched
(a) 1 failure vs 0 success rows -> no; (b) 0 review rows — opposite-verdict test is a human judgment, not a count; (c) 1 conceptual failure vs 1 conceptual success, 1 original -> fires (balanced)
- 3Unknown — underdeterminednot reached
alternative_explanations=1 (needs <=1) AND evidence_strength=2 (needs >=3)
- 4Establishednot reached
independent_replication=1 (>=4), evidence_strength=2 (>=4), consensus=2 (>=4)
- 5Strong, domain-limitednot reached
evidence_strength=2 (>=4) AND independent_replication=1 (>=3), plus a written scope limit (human judgment)
- 6Active consensus, incompletenot reached
consensus=2 (>=4) AND (methodological_quality=3 <=3 OR alternative_explanations=1 <=2)
- 7Plausible, under active investigationnot reached
independent_replication=1 (needs ==2) AND evidence_strength=2 (needs 2-3), plus recent activity (human judgment)
- 8Speculativenot reached
3/5 rows are empirical (needs 0) AND evidence_strength=2 (needs <=1)
- 9Preliminarymatchednot reached
4 non-theoretical row(s) exist AND independent_replication=1 (needs <=1)
Paired claim
A status that belongs to a relationship between two claims, not to either one alone.
observation is 'established' (established family) and mechanism is 'unknown_conflicting' (not settled). This is a property of the pair — both claims keep the status their own evidence earned, and neither was overwritten.
This claim is the mechanism · paired with the observation
EstablishedHeavily over-parameterized neural networks generalize well to held-out data despite having enough capacity to fit random labels
Why these are linked · reasonedPilot 5 (rubric v6). Claim B proposes implicit SGD regularization as the explanation for the phenomenon asserted by Claim A. Definitionally true of the two statements as written - B names A's effect as the thing it explains - so confidence_basis is 'reasoned' rather than 'empirical': no co-citation check can establish that one sentence is a proposed mechanism for another, which is a semantic relation between claims rather than a co-occurrence fact about terms (BASIC_METHODOLOGY Tier 1 §10 allows reasoned for definitionally-true claims, with a written rationale). This is the corpus's first claim-to-claim edge and the link that makes phenomenon_established_mechanism_uncertain derivable at all.
Recorded rationale
v5 rule 2 clause (c1): 1 conceptual_replication_failure (Jiang 2020) vs 1 conceptual_replication_success (Keskar 2016), within 1 with at least one of each. A contradiction row (Dinh 2017) exists but clause (a) needs one success-side row too, and there is none. Naive citation-only guess was 'preliminary' - the FIRST claim in the corpus not naively guessed 'established', and the first where the naive guess was too DISMISSIVE rather than too generous; see Pilot 5 log Findings 2 and 4 (OpenAlex fragments this field's arXiv/proceedings records, so counts measure indexing rather than influence). Pilot 5; prior familiarity disclosed HIGH.