The central rule of statistical modelling taught in every introductory course is – A model that is too complex for its data will overfit. It will memorise the training examples rather than learning the underlying pattern, and it will fail when presented with new data. Keep the model simple. Do not use more parameters than necessary.
Modern deep learning violates this rule systematically and successfully.
Neural networks used for image recognition, language generation, protein structure prediction and other practical usecases, routinely have millions of parameters trained on datasets with far fewer examples. By classical statistical theory, these networks should overfit catastrophically. They should memorise training data and generalise poorly. Instead, they generalise with remarkable accuracy across tasks that no one explicitly programmed them to solve!
Why this works is not fully understood? The gap between what deep learning does empirically and what existing mathematical theory can explain is one of the most consequential open problems in contemporary science.
What Classical Theory Predicts
Statistical learning theory addresses a fundamental question. If a model is trained on a finite sample of data drawn from some unknown distribution, how well can it be expected to perform on new data from the same distribution? The difference between training performance and test performance is called the generalization gap.
The classical answer to controlling this gap is the bias-variance tradeoff. A model with few parameters has high bias (meaning: it cannot represent complex patterns in the data) but low variance (meaning: its predictions are stable across different training sets). A model with many parameters has low bias but high variance (meaning: it fits the training data well but fluctuates unpredictably when the training set changes). The optimal model sits at the minimum of the combined error, which for a fixed dataset corresponds to a specific level of complexity.
The mathematical tools developed to formalise this intuition, like- VC dimension, Rademacher complexity, PAC-learning bounds, all arrive at the same qualitative conclusion: complexity must be controlled1. Most of these existing theories fail when applied to modern deep networks. They are both unintelligent and non-predictive in realistic settings. This gap between theory and practice is largest for overparameterised models, which in theory have the capacity to overfit their training sets, but often do not in practice.
The networks simply do not behave as the theory says they should.
The Double Descent Phenomenon
In 2019, Mikhail Belkin and colleagues at the University of California published a result that formalised what practitioners had observed empirically for years2. Double descent is the phenomenon where a model’s error rate on the test set initially decreases with the number of parameters, then peaks, then decreases again. This phenomenon has been considered surprising, as it contradicts assumptions about overfitting in classical machine learning. The increase usually occurs near the interpolation threshold, where the number of parameters is the same as the number of training data points.

(A) The classical U-shaped risk curve arising from the bias–variance trade-off.
(B) The double-descent risk curve, which incorporates the U-shaped risk curve (i.e., the “classical” regime) together with the observed behavior from using high-capacity function classes (i.e., the “modern” interpolating regime), separated by the interpolation threshold. The predictors to the right of the interpolation threshold have zero training risk.3
The interpolation threshold is the point at which a model becomes large enough to fit the training data perfectly and to achieve zero training error. Classical theory predicts that performance beyond this point should deteriorate as the model begins memorising noise. What Belkin et al. observed was that performance deteriorates up to the interpolation threshold, peaks in error, and then as the model grows further, performance begins to improve again.
Double descent provides an indication that even though models that pass through every training data point are indeed overfitted, the structure of the resulting network forces the interpolation to be smooth and results in superior generalisation to unseen data.
This second descent, occurring in what is called the overparameterised regime, is not a marginal effect. It has been documented across linear models, fully connected networks, convolutional networks, and transformers. Even with fixed data, transformers like GPT-2/3/4 get better at generalising as the model size grows.
Implicit Bias of Gradient Descent
When a neural network is trained, the parameters are updated iteratively using gradient descent — a procedure that adjusts each parameter in the direction that reduces the training loss. For overparameterised networks, many different parameter configurations can achieve zero training error. The network has, in a sense, too many solutions available. Which solution gradient descent finds depends on the details of the procedure itself.
Research since 2017 has established that gradient descent has an implicit bias, meaning, it preferentially converges to solutions with specific mathematical properties even when not explicitly directed to do so. For linear models trained with gradient descent on separable data, the implicit bias is toward the maximum-margin solution, meaning the classifier that separates the training classes with the largest possible gap.
When a neural network is trained, gradient descent has to choose between an enormous number of different parameter configurations that all achieve the same low error on training data. Some of those configurations would fail on new data because they have memorised the specific quirks of the training set rather than learning the underlying pattern. Others capture genuine structure that transfers to new examples.
Gradient descent does not choose randomly between these options. It consistently prefers the second type, solutions that are smoother and more structured. It does this not because any rule instructs it to but because of how the mathematics of the algorithm works. It steers the network away from solutions that would generalise poorly, without any explicit penalty being written into the training objective.
For deep nonlinear networks, the same tendency is observed consistently in practice, but the precise mathematical account of why it happens remains an active area of research.
Grokking
A more recent phenomenon is grokking, a term introduced in a 2022 paper from researchers at OpenAI.4
In a standard training run, a neural network’s performance on training data and test data improve together, then stabilise. In grokking, something different happens: the network first achieves near-perfect performance on the training data, then enters a prolonged period where training performance remains high but test performance does not improve. Then, after many additional training steps test performance suddenly and dramatically improves.
The network appears to have learned two distinct solutions sequentially: first a memorisation strategy that works on training data only, then a genuine generalisation strategy that works on new data. The transition between them is abrupt.
The grokking phenomenon refers to the sudden and substantial improvement in a model’s performance following a prolonged period of stagnant or even regressive learning. It is possible to provide insights into grokking by leveraging existing theoretical foundations of machine learning, in particular concepts from statistical learning theory, such as norm-based and stability-based generalisation bounds.
The phenomenon has been documented primarily on tasks with clean mathematical structure — modular arithmetic, group theory problems — which may limit how directly it generalises to the messy real-world tasks where deep learning is most widely deployed. Whether grokking is a feature of network training more broadly, or a specific property of certain task structures, is not yet established.
What Networks Represent Internally
A separate and equally difficult question concerns not whether networks generalise, but what they actually learn when they do.
A trained neural network is a composition of linear transformations and nonlinear activation functions applied across many layers. As data moves through a neural network, each layer produces a set of numbers describing that data from the network’s perspective. Those sets of numbers exist in spaces too large to visualise, and their values do not correspond to anything we can directly read or interpret.
When a network trained to recognise images correctly identifies a photograph of a dog, the mathematical operations producing that classification do not correspond in any obvious way to the features we would use to identify a dog.
This opacity creates a practical problem. A network may perform well on a test set drawn from the same distribution as its training data while behaving unpredictably on inputs that differ from that distribution in ways that would be trivial for a human to handle. Small, carefully constructed perturbations to input images which are imperceptible to human observers can cause networks to misclassify with high confidence. Networks trained on photographs taken in one context may fail on photographs taken in slightly different conditions.
The field of mechanistic interpretability attempts to reverse-engineer what computations networks perform internally. Researchers have identified circuits within networks that perform identifiable operations such as detecting curves, recognising specific syntactic patterns in text, or computing indirect object identification. Whether these results scale to the large networks used in practice, and whether full mechanistic understanding of a large network is computationally feasible at all, are open questions.
The State of the Theory
A comprehensive theory of learning, especially one that explains the surprising empirical puzzles raised by deep learning remains incomplete. An eventual such theory, explaining why and how deep networks work and what their limitations are, could enable the development of even more powerful learning approaches, guide the use of these algorithms in safety-critical and high-stakes settings, and inform our understanding of human intelligence.
The current situation is that deep learning works reliably in practice, often exceeding human performance on well-defined tasks, while the mathematical account of ‘why it works’ remains incomplete. The classical framework (bias-variance tradeoff, VC dimension, explicit regularisation) does not adequately explain the behaviour of large overparameterised networks. The emerging frameworks (implicit bias of gradient descent, the geometry of overparameterised solution spaces, double descent, grokking) each capture part of the picture without providing a unified account.
This is not unusual in the history of science. Thermodynamics was an effective engineering theory for decades before statistical mechanics provided its microscopic foundation. Quantum mechanics was used to calculate atomic spectra before its mathematical structure was fully axiomatised. Technologies often precede their theoretical justification.
What is unusual is the scale at which deep learning is being deployed in consequential settings, like in the medical diagnosis, legal systems, financial decisions, autonomous vehicles even while the theoretical understanding of its behaviour remains incomplete. A network that generalises well on average may fail in specific, unpredictable ways that a complete theory would identify in advance.
That theory does not exist yet.
Further Readings:
- Nakkiran, P. et al. (2021). Deep Double Descent: Where Bigger Models and More Data Hurt. Journal of Statistical Mechanics, 2021. https://arxiv.org/abs/1912.02292
- Zhang, C. et al. (2017). Understanding Deep Learning Requires Rethinking Generalization. ICLR 2017. https://arxiv.org/abs/1611.03530
- Soudry, D. et al. (2018). The Implicit Bias of Gradient Descent on Separable Data. Journal of Machine Learning Research, 19(1), 2822–2878.
- Petersen, P. & Zech, J. (2024). Mathematical Theory of Deep Learning. arXiv. https://arxiv.org/abs/2407.18384
- MIT: Statistical Learning Theory and Applications
- https://www.alignmentforum.org/posts/FRv7ryoqtvSuqBxuT/understanding-deep-double-descent
- Cover image:Artistic illustration of a neuromorphic system of waveguides carrying light.© Clara Wanjura
Footnotes:
- https://dl.acm.org/doi/10.1145/3716554.3716617 ↩︎
- M. Belkin, D. Hsu, S. Ma, & S. Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off, Proc. Natl. Acad. Sci. U.S.A. 116 (32) 15849-15854, https://doi.org/10.1073/pnas.1903070116 (2019). ↩︎
- https://www.pnas.org/doi/10.1073/pnas.1903070116#fig01 ↩︎
- https://arxiv.org/abs/2201.02177 ↩︎

Leave a Reply