Data Science & AI Glossary: Key Terms Explained in Plain English
This glossary defines the data science, analytics and AI terms that come up most in courses, job postings and meetings, each in a sentence or two of plain English. Where we've published a full guide on a term, the entry links to it. We update the page as new guides go live.
Terms are grouped by topic and listed alphabetically within each group. Use Ctrl+F (Cmd+F on Mac) to jump to a specific term.
Analytics & Statistics
A/B test. An experiment that randomly splits users into groups and shows each group a different version (of a page, an email or a price) to measure which performs better. Randomization is what lets you draw conclusions about cause and effect. See What Is A/B Testing?.
Causation. A relationship where one variable actually produces a change in another. Proving it usually needs an experiment. See Correlation vs Causation.
Confounder. A hidden third variable that influences two others and creates a misleading correlation between them, like hot weather driving both ice cream sales and drownings.
Correlation. A measure of how strongly two variables move together, from −1 to +1. It doesn't tell you why they move together. See Correlation vs Causation.
Descriptive analytics. Analysis that summarizes what happened: totals, trends, breakdowns. It's the bulk of everyday analyst work. See The 4 Types of Data Analytics.
Diagnostic analytics. Analysis that explains why something happened, using drill-downs, segmentation and funnel analysis. See The 4 Types of Data Analytics.
KPI (key performance indicator). One of the few metrics a team has agreed matter most for judging success. All KPIs are metrics, but most metrics aren't KPIs. See KPI vs Metric.
Mean, median and mode. Three ways to describe a "typical" value: the arithmetic average, the middle value, and the most frequent value. In skewed data like salaries, the median is usually the more honest summary. See Mean vs Median vs Mode.
Predictive analytics. Using historical data to estimate future outcomes, like forecasts, churn risk or demand. See The 4 Types of Data Analytics.
Prescriptive analytics. Analysis that recommends an action by comparing options' expected outcomes. See The 4 Types of Data Analytics.
Selection bias. A distortion that happens when the people or records in your data aren't representative, for example because customers chose whether to join the group you're measuring.
Simpson's paradox. A pattern that shows up in every subgroup but reverses when the groups are combined, usually because the groups are different sizes. See the worked example in Correlation vs Causation.
Statistical significance. An indication that a result is unlikely to be due to chance alone, often judged with a p-value. It doesn't mean the effect is large or important. See P-Values Explained.
SQL & Data Analysis
Aggregate function. A function that turns many rows into one value, such as SUM, COUNT, AVG, MIN or MAX. Usually paired with GROUP BY.
Anti-join. A query pattern that finds rows in one table with no match in another, typically a LEFT JOIN plus WHERE ... IS NULL. See SQL Joins Explained.
CTE (common table expression). A named, temporary result defined with WITH at the top of a query, used to break complex SQL into readable steps. See The Complete SQL Roadmap.
Data cleaning. Fixing or removing incorrect, duplicate, inconsistent or missing data before analysis. It often takes more time than the analysis itself. See our Data Cleaning Checklist.
Foreign key. A column in one table that refers to the primary key of another, such as orders.customer_id pointing to customers.customer_id. It's what joins are usually built on.
INNER JOIN. A join that keeps only rows with a match in both tables. See SQL Joins Explained.
LEFT JOIN. A join that keeps every row from the left table and fills in NULLs where the right table has no match. It's the most-used join in analysis. See SQL Joins Explained.
NULL. A marker meaning "no value" or "unknown". It isn't zero or an empty string, and comparisons with NULL don't return true, which causes many subtle SQL bugs.
pandas. The most widely used Python library for working with tabular data, built around the DataFrame, a table you can filter, group and reshape in code. See Pandas for Beginners.
Primary key. A column (or set of columns) that uniquely identifies each row in a table.
Window function. A SQL function that calculates across related rows without collapsing them into one, such as ROW_NUMBER, RANK, LAG or running totals. See SQL Window Functions Explained.
Business Intelligence
Calculated column. In Power BI, a DAX column computed row by row when data refreshes and stored in the model. It's best for values you'll filter or group by. See What Is DAX?.
CALCULATE. The most important DAX function. It evaluates an expression inside a modified filter context. See What Is DAX?.
Dashboard. A single view that combines key charts and metrics for monitoring. See How to Build Your First Power BI Dashboard.
DAX (Data Analysis Expressions). The formula language for calculations in Power BI, Power Pivot and Analysis Services. See What Is DAX?.
Dimension table. A table of descriptive attributes, such as products, customers or dates, used to filter and group facts. See Star Schema vs Snowflake Schema.
Fact table. A table of measurable events, such as sales, orders or clicks, usually with many rows and foreign keys to dimension tables. See Star Schema vs Snowflake Schema.
Filter context. In DAX, the set of filters active for a particular cell or visual. It's why the same measure shows different values in different places. See What Is DAX?.
Measure. In Power BI, a DAX calculation computed on demand for whatever the visual is showing, such as total sales, margin % or year-over-year growth. See What Is DAX?.
Power Query. The data preparation tool in Power BI and Excel, used to connect to, clean and reshape data before it loads into the model.
Star schema. A data model with one central fact table connected to surrounding dimension tables. It's the recommended structure for Power BI and most data warehouses. See Star Schema vs Snowflake Schema.
Machine Learning
Classification. An ML task that predicts a category, such as churn vs no churn, or spam vs not spam. See Classification vs Regression.
Deep learning. Machine learning using neural networks with many layers. It excels with images, audio and text. See AI vs Machine Learning vs Deep Learning.
Feature. An input variable a model uses to make predictions, like customer age or number of purchases in the last 90 days.
Machine learning. The part of AI where systems learn patterns from data instead of following hand-written rules. See AI vs Machine Learning vs Deep Learning.
Model. The output of training: a learned function that turns inputs (features) into predictions.
Overfitting. When a model learns the training data too closely, noise included, and performs poorly on new data. See Overfitting vs Underfitting.
Precision and recall. Two metrics for classification models. Precision is the share of flagged cases that are truly positive, and recall is the share of true positives the model catches. On imbalanced data they're far more informative than accuracy. See Precision vs Recall.
Regression. An ML task that predicts a number, such as price, demand or revenue. It's also the name of a family of statistical methods. See Classification vs Regression.
Supervised learning. Training a model on examples where the correct answer (label) is known. See Supervised vs Unsupervised Learning.
Training data / test data. The data a model learns from, and separate held-back data used to check how well it performs on examples it hasn't seen.
Unsupervised learning. Finding structure in data without labels, for example grouping customers into segments by clustering. See Supervised vs Unsupervised Learning.
Generative AI & LLMs
AI agent. A system in which an LLM plans steps, calls tools (search, databases, code) and acts on the results to complete a goal. See Generative AI vs Agentic AI.
Artificial intelligence (AI). The broad field of making computers perform tasks that normally need human intelligence. See AI vs Machine Learning vs Deep Learning.
Context window. The maximum number of tokens a model can consider in one request, including instructions, documents, history and the reply. See Tokens, Context Windows and Temperature.
Embedding. A list of numbers representing the meaning of a piece of text (or an image), so that similar meanings end up close together. Embeddings power semantic search and RAG. See What Are Embeddings and Vector Databases?.
Fine-tuning. Further training a pre-trained model on your own examples to change its behavior or style. See RAG vs Fine-Tuning vs Prompt Engineering.
Generative AI. AI that creates new content, such as text, images, code or audio. See AI vs Machine Learning vs Deep Learning.
Hallucination. When a language model produces confident but false or made-up information. See What Are AI Hallucinations?.
LangChain / LlamaIndex. Popular frameworks for building LLM applications, especially retrieval-based ones. See LangChain vs LlamaIndex.
Large language model (LLM). A deep learning model trained on huge amounts of text to predict and generate language. It's the technology behind ChatGPT, Claude and Gemini. See How Large Language Models Work.
MCP (Model Context Protocol). An open standard for connecting AI applications to external data and tools through reusable "servers" that expose tools, resources and prompts. See What Is MCP?.
Prompt engineering. Designing instructions and context that get reliable, useful output from an LLM. See Mastering Prompt Engineering: A 30-Day Study Plan.
RAG (retrieval-augmented generation). An approach where relevant documents are retrieved and given to the LLM with the question, so answers are grounded in trusted sources. See What Is RAG?.
Temperature. A setting that controls how random an LLM's output is. Lower means more focused and repeatable. See Tokens, Context Windows and Temperature.
Token. The unit of text an LLM reads, writes and bills by, roughly ¾ of an English word. See Tokens, Context Windows and Temperature.
Transformer. The neural network architecture, introduced in 2017, that almost all modern LLMs are built on. See How Large Language Models Work.
Vector database. A database designed to store embeddings and quickly find the ones most similar to a query. It's a core component of RAG systems. See What Are Embeddings and Vector Databases?.
Data Engineering
Data lake / lakehouse. A data lake stores raw data of any type cheaply in object storage. A lakehouse adds warehouse features (tables, reliable updates, fast SQL) on top of that storage using open table formats. See Data Warehouse vs Data Lake vs Lakehouse.
Data pipeline. An automated sequence of steps that moves data from sources to where it's used, cleaning and transforming it along the way.
Data warehouse. A central database optimized for analytics, holding cleaned, structured data from many sources. See Data Warehouse vs Data Lake vs Lakehouse.
dbt (data build tool). A tool for writing, testing and documenting SQL transformations inside the warehouse. See What Is dbt?.
ETL / ELT. Extract-Transform-Load and Extract-Load-Transform: two orders for moving and preparing data. ELT, where you load raw data first and transform it inside the warehouse, is the modern default. See ETL vs ELT.
MLOps. Practices and tools for deploying, monitoring and maintaining machine learning models in production. See 5 Best MLOps Courses.
Orchestration. Scheduling and coordinating pipeline tasks and their dependencies, using tools like Airflow, Dagster or Prefect. See Airflow vs Dagster vs Prefect.
Where to Go From Here
If you're new to the field, the data analyst roadmap puts many of these terms in the order you'll need them. If you're heading toward AI, start with AI vs Machine Learning vs Deep Learning vs Generative AI. To compare courses across every topic here, browse Data Science & Machine Learning courses.