Overfitting vs Underfitting: How to Spot and Fix Both

Overfitting vs Underfitting: How to Spot and Fix Both

Overfitting happens when a model learns its training data too closely, including the noise, so it scores very well on data it has seen and poorly on new data. Underfitting happens when a model is too simple to capture the real pattern, so it performs poorly everywhere. The goal is a model that generalizes: one that performs about as well on new data as on its training data.

Every supervised model has to strike this balance, and it's one of the most common interview topics in machine learning. The good news is that you can diagnose it with a simple comparison of two numbers.

An Analogy: Three Students

Three students prepare for an exam using last year's past papers.

  • Student A skims the material and learns a few general rules. They score 60% on the practice papers and 58% on the real exam. That's underfitting: too little learned.
  • Student B memorizes the exact answers to every past-paper question. They score 100% on the practice papers and 55% on the real exam, because the questions changed. That's overfitting: they memorized instead of understanding.
  • Student C works through the concepts behind the questions. They score 85% on practice and 82% on the real exam. That's a good fit.

Past papers are the training data. The real exam is new data. What counts is the real exam.

A Worked Example: Tuning a Decision Tree

A team builds a model to predict customer churn. They train decision trees of increasing depth (how many questions the tree may ask) and measure accuracy on the training data and on a separate test set the model never saw:

Tree depth Training accuracy Test accuracy Diagnosis
2 78% 77% Underfitting: both scores low
6 86% 84% Good fit: both high, small gap
Unlimited 100% 71% Overfitting: perfect on training, big drop on test

The unlimited tree kept splitting until it had a rule for nearly every individual customer, including random quirks that won't repeat. It "knows" the training data perfectly and has learned little that transfers to new customers.

How to Diagnose: Compare Two Numbers

Pattern Training score Test / validation score What it means
Underfitting Low Low (similar) Model too simple, or features too weak
Good fit High High (close to training) Generalizes well
Overfitting Very high Much lower Memorizing noise

The gap between training and test performance is your main warning sign. That's why you must always hold out data the model doesn't train on, which is standard practice in supervised learning.

Learning curves give a fuller picture. They plot training and validation scores as you add more training data. If the curves converge at a low score, the model is underfitting. If there's a persistent gap with a high training score, it's overfitting, and more data may help.

The Bias–Variance Trade-off

The textbook framing uses two terms:

  • Bias: error from wrong or overly simple assumptions. High bias leads to underfitting. A straight line fitted to a curved relationship has high bias.
  • Variance: error from sensitivity to the particular training sample. High variance leads to overfitting. Train the unlimited tree on a slightly different sample and you get a very different tree.

As a model gets more complex, bias usually falls and variance rises. The best model sits in the middle. You don't need the math to use this idea, but it helps to know the vocabulary. See how much math ML really needs.

How to Fix Underfitting

  1. Use a more flexible model, for example gradient boosting instead of linear regression, or a deeper tree.
  2. Add better features. Often the model isn't too simple, it just lacks the right information. "Days since last purchase" may predict churn far better than anything in the raw data.
  3. Reduce regularization if you applied too much.
  4. Train longer for models trained iteratively, like neural networks.

How to Fix Overfitting

  1. Get more training data. It's the most reliable fix when it's possible.
  2. Simplify the model: limit tree depth, reduce the number of features, use fewer layers.
  3. Regularization: add a penalty for complexity. L1 (lasso) and L2 (ridge) for linear models, max_depth and min_samples_leaf for trees, dropout for neural networks.
  4. Cross-validation: evaluate on several different splits of the data so you don't tune to one lucky test set.
  5. Early stopping: stop training when validation performance stops improving.
  6. Ensembles: random forests average many trees, which reduces variance.

A Hidden Cause: Data Leakage

Sometimes a model scores well on both training and test data and then fails in production. That's often data leakage, not overfitting in the usual sense: a feature that contains information the model won't have at prediction time. Examples are a "refund issued" flag in a model predicting complaints, or statistics calculated on the full dataset before splitting. If results look too good to be true, check for leakage first.

Common Mistakes

  1. Evaluating on training data. A 99% training score means nothing on its own.
  2. Tuning on the test set. If you try 50 settings and pick the best test score, the test set has become part of training. Use a separate validation set (or cross-validation) for tuning, and touch the test set once at the end.
  3. Assuming complex models always win. On small, tabular business datasets, simpler models often generalize better.
  4. Chasing tiny accuracy gains. An improvement of 0.3% on one test split is often noise.

What to Learn Next

Overfitting is a core topic in any ML foundation course. The Machine Learning Specialization (Stanford / DeepLearning.AI) explains bias, variance and regularization very clearly (see our full review). Machine Learning A-Z takes a more hands-on, code-first approach. If you're choosing between them, read Machine Learning A-Z vs the ML Specialization. Browse all Machine Learning courses.

Frequently Asked Questions

Is overfitting or underfitting worse? Overfitting is often more dangerous, because the model looks good during development and fails quietly on new data. Underfitting is obvious straight away because scores are low everywhere.

How much gap between training and test accuracy is too much? There's no universal threshold. It depends on the problem and the data size. A gap of a couple of percentage points is usually fine. A gap of 10+ points, as in our unlimited tree example, clearly signals overfitting.

Can more data fix underfitting? Usually not. If the model is too simple, more of the same data won't help. You need a more flexible model or better features. More data mainly helps overfitting.

Do large language models overfit? They can memorize parts of their training data, and fine-tuning on a small dataset can overfit quickly. That's why fine-tuning uses validation sets and early stopping, just like classic ML.

Bottom Line

Underfitting means the model learned too little. Overfitting means it memorized noise. Compare training and test performance to diagnose which one you have. Then add flexibility or features to fix underfitting, and add data, simplify, regularize or cross-validate to fix overfitting. The aim isn't the highest training score but the best performance on data the model has never seen.

Enjoyed this article?

Share it with your network

Listings related to Overfitting vs Underfitting: How to Spot and Fix Both