Machine Learning Specialization (Stanford / DeepLearning.AI)
A three-course beginner program from DeepLearning.AI and Stanford Online covering supervised, unsupervised and neural network methods in Python.
Overfitting happens when a model learns its training data too closely, including the noise, so it scores very well on data it has seen and poorly on new data. Underfitting happens when a model is too simple to capture the real pattern, so it performs poorly everywhere. The goal is a model that generalizes: one that performs about as well on new data as on its training data.
Every supervised model has to strike this balance, and it's one of the most common interview topics in machine learning. The good news is that you can diagnose it with a simple comparison of two numbers.
Three students prepare for an exam using last year's past papers.
Past papers are the training data. The real exam is new data. What counts is the real exam.
A team builds a model to predict customer churn. They train decision trees of increasing depth (how many questions the tree may ask) and measure accuracy on the training data and on a separate test set the model never saw:
| Tree depth | Training accuracy | Test accuracy | Diagnosis |
|---|---|---|---|
| 2 | 78% | 77% | Underfitting: both scores low |
| 6 | 86% | 84% | Good fit: both high, small gap |
| Unlimited | 100% | 71% | Overfitting: perfect on training, big drop on test |
The unlimited tree kept splitting until it had a rule for nearly every individual customer, including random quirks that won't repeat. It "knows" the training data perfectly and has learned little that transfers to new customers.
| Pattern | Training score | Test / validation score | What it means |
|---|---|---|---|
| Underfitting | Low | Low (similar) | Model too simple, or features too weak |
| Good fit | High | High (close to training) | Generalizes well |
| Overfitting | Very high | Much lower | Memorizing noise |
The gap between training and test performance is your main warning sign. That's why you must always hold out data the model doesn't train on, which is standard practice in supervised learning.
Learning curves give a fuller picture. They plot training and validation scores as you add more training data. If the curves converge at a low score, the model is underfitting. If there's a persistent gap with a high training score, it's overfitting, and more data may help.
The textbook framing uses two terms:
As a model gets more complex, bias usually falls and variance rises. The best model sits in the middle. You don't need the math to use this idea, but it helps to know the vocabulary. See how much math ML really needs.
max_depth and min_samples_leaf for trees, dropout for neural networks.Sometimes a model scores well on both training and test data and then fails in production. That's often data leakage, not overfitting in the usual sense: a feature that contains information the model won't have at prediction time. Examples are a "refund issued" flag in a model predicting complaints, or statistics calculated on the full dataset before splitting. If results look too good to be true, check for leakage first.
Overfitting is a core topic in any ML foundation course. The Machine Learning Specialization (Stanford / DeepLearning.AI) explains bias, variance and regularization very clearly (see our full review). Machine Learning A-Z takes a more hands-on, code-first approach. If you're choosing between them, read Machine Learning A-Z vs the ML Specialization. Browse all Machine Learning courses.
Is overfitting or underfitting worse? Overfitting is often more dangerous, because the model looks good during development and fails quietly on new data. Underfitting is obvious straight away because scores are low everywhere.
How much gap between training and test accuracy is too much? There's no universal threshold. It depends on the problem and the data size. A gap of a couple of percentage points is usually fine. A gap of 10+ points, as in our unlimited tree example, clearly signals overfitting.
Can more data fix underfitting? Usually not. If the model is too simple, more of the same data won't help. You need a more flexible model or better features. More data mainly helps overfitting.
Do large language models overfit? They can memorize parts of their training data, and fine-tuning on a small dataset can overfit quickly. That's why fine-tuning uses validation sets and early stopping, just like classic ML.
Underfitting means the model learned too little. Overfitting means it memorized noise. Compare training and test performance to diagnose which one you have. Then add flexibility or features to fix underfitting, and add data, simplify, regularize or cross-validate to fix overfitting. The aim isn't the highest training score but the best performance on data the model has never seen.
A three-course beginner program from DeepLearning.AI and Stanford Online covering supervised, unsupervised and neural network methods in Python.
One of Udemy's best-known ML courses, covering regression, classification and clustering in both Python and R with template-based coding.