Unsupervised Learning in Python
Hands-on Python course on k-means, hierarchical clustering, t-SNE, PCA and NMF with scikit-learn, ending in a music artist recommender.
Supervised learning trains a model on examples where the correct answer (the label) is already known, so it can predict that answer for new cases. Examples are predicting churn or estimating a price. Unsupervised learning works on data without labels and finds structure on its own, such as grouping similar customers or flagging unusual transactions. The key question is simple: do you have a known outcome to learn from?
Almost every machine learning course starts with this distinction, and for good reason. It decides which algorithms you can use, how you measure success, and how much data preparation you'll need. This guide follows one company through both approaches.
A telecom company has 50,000 mobile customers. For each one it has 12 months of data: plan type, monthly bill, data usage, number of support calls, contract length and tenure. Management asks two questions:
The first question needs supervised learning. The second needs unsupervised learning.
For the churn question, the company looks back at last year's data. It knows which customers left and which stayed. That known outcome is the label.
The model sees thousands of examples like this:
| Tenure (months) | Monthly bill | Support calls | Contract | Churned? |
|---|---|---|---|---|
| 4 | $38 | 5 | Monthly | Yes |
| 36 | $52 | 0 | 2-year | No |
| 11 | $61 | 3 | Monthly | Yes |
It learns which combinations of inputs (called features) tend to lead to churn. Then it scores current customers, whose outcome isn't known yet, with a churn probability. The retention team calls the highest-risk customers first.
The two kinds of supervised problems:
Common algorithms: linear and logistic regression, decision trees, random forests, gradient boosting (XGBoost, LightGBM), and neural networks.
How you measure success: since you know the right answers for past data, you hold some of it back and check the predictions against it. You measure accuracy, precision and recall for classification, or average error for regression. That makes supervised learning easy to evaluate objectively. It's also how you catch a model that has memorized its training data instead of learning general patterns. See Overfitting vs Underfitting.
The second question has no "correct" label. Nobody tagged customers as "type A" or "type B". So the company runs a clustering algorithm (k-means is the classic one) on the usage and billing data, and asks it to find natural groups.
It comes back with four segments. An analyst then interprets and names them:
| Segment | Share | What defines it | Business idea |
|---|---|---|---|
| Heavy streamers | 22% | Very high data use, mid bill | Offer unlimited-data upgrade |
| Budget loyalists | 35% | Low bill, long tenure, few calls | Protect with loyalty perks |
| Business users | 15% | High bill, weekday usage, roaming | Push business bundles |
| At-risk newcomers | 28% | Short tenure, many support calls | Improve onboarding |
The algorithm found the groups. Humans gave them meaning. That's typical of unsupervised learning.
Other unsupervised tasks:
How you measure success: this is the hard part. There's no answer key, so you rely on statistical measures of how well-separated the groups are, plus the real test: are the segments useful for decisions? Two analysts can reasonably produce different segmentations of the same data.
| Supervised | Unsupervised | |
|---|---|---|
| Needs labels? | Yes | No |
| Question it answers | "What will the outcome be?" | "What patterns exist?" |
| Typical tasks | Classification, regression | Clustering, anomaly detection, dimensionality reduction |
| Telecom example | Predict who will churn | Discover customer segments |
| Evaluation | Objective: compare to known answers | Partly subjective: usefulness, cluster quality |
| Main cost | Getting good labels | Interpreting results |
| Share of business use | Most production ML | Common in exploration and marketing |
Semi-supervised learning. You have a few labeled examples and many unlabeled ones. For example, 500 transactions manually checked for fraud out of 5 million. Methods combine the two so you don't have to label everything.
Self-supervised learning. The data creates its own labels. Large language models are trained this way: hide the next word, predict it, and repeat across enormous amounts of text. It's technically supervised, but no human labeling is needed, which is what made modern generative AI possible.
Reinforcement learning. There are no fixed labels. An agent learns by trial and error from rewards: a game score, a completed delivery, or human ratings of answers. It's used in robotics, recommendation tuning and fine-tuning chatbots.
In practice they're often combined:
Supervised models are the engine of predictive analytics, one of the four types of data analytics. Unsupervised methods often support the descriptive and diagnostic side, revealing structure you can then explain and act on.
Most ML courses start with supervised learning, because it's easier to evaluate. Supervised Learning with scikit-learn (DataCamp) is a hands-on start in Python, and Unsupervised Learning in Python covers clustering and dimensionality reduction. For a full foundation, see our Machine Learning Specialization review and the 6-month ML engineer roadmap. Browse all Machine Learning courses.
Is supervised or unsupervised learning more common? Supervised learning, in production systems. Most business ML predicts a known outcome such as churn, fraud, demand or price. Unsupervised learning is used heavily in exploration, segmentation and anomaly detection.
Is clustering supervised or unsupervised? Unsupervised. Clustering groups data points by similarity without any predefined labels.
Is ChatGPT supervised or unsupervised? Its pre-training is self-supervised: it learns by predicting the next word in huge amounts of text. It's then refined with supervised examples and reinforcement learning from human feedback.
Which should I learn first? Supervised learning. Its concepts (train/test split, overfitting, evaluation metrics) are the foundation for everything else, and it's what most job interviews focus on.
Supervised learning predicts a known kind of outcome from labeled examples, and it's objective to evaluate and dominant in production. Unsupervised learning finds structure in unlabeled data, and needs human judgment to interpret. Ask whether you have a labeled outcome to learn from. The answer tells you which approach you need, and many real projects use both.
Hands-on Python course on k-means, hierarchical clustering, t-SNE, PCA and NMF with scikit-learn, ending in a music artist recommender.
Hands-on 4-hour DataCamp course teaching classification, regression, tuning, and pipelines in scikit-learn with real datasets in Python.