Generative AI with Large Language Models (DeepLearning.AI)
A hands-on course covering how LLMs are built, fine-tuned, and deployed, taught by AWS AI practitioners through DeepLearning.AI on Coursera.
Tokens are the chunks of text a language model reads and writes, often parts of words. A rough rule for English is that 1 token is about ¾ of a word. The context window is the maximum number of tokens a model can consider at once, including your instructions, documents, chat history and its own reply. Temperature controls how random the output is: low values give focused, consistent answers, and higher values give more varied ones.
These three settings explain a lot of confusing LLM behavior. They're why a long chat seems to "forget" earlier instructions, why the same prompt gives different answers, and why an API bill can be larger than expected. Once you understand them, you'll prompt better, choose models better and estimate costs properly.
Language models don't see letters or whole words. Before processing, text is split into tokens using a vocabulary the model learned during training. Common words are usually a single token. Rarer words are split into pieces.
For example, a sentence like "Data analysts love spreadsheets" might be split into something like:
Data · analysts · love · spread · sheets
The exact split depends on the model's tokenizer, and different vendors use different ones.
Useful rules of thumb for English:
Two practical consequences:
API pricing is per token, usually quoted per million tokens, with separate prices for input (what you send) and output (what the model writes). Output tokens usually cost several times more than input tokens.
Worked example. You want to summarize 500 customer support tickets. Each ticket plus your instructions is about 800 tokens, and each summary is about 150 tokens. Using an illustrative price of $3 per million input tokens and $15 per million output tokens:
| Tokens | Price per million | Cost | |
|---|---|---|---|
| Input | 500 × 800 = 400,000 | $3 | $1.20 |
| Output | 500 × 150 = 75,000 | $15 | $1.13 |
| Total | ≈ $2.33 |
Prices vary widely between models, and a smaller model can cost a fraction of this. The method stays the same: estimate tokens in and out, multiply by the price, then test on a small sample before running the full batch.
The context window is the total number of tokens the model can take into account in a single request. Everything has to fit in it:
Context windows have grown fast. Many current models accept hundreds of thousands of tokens, and some over a million. Check each model's documentation, because limits differ, and there's usually a separate, smaller cap on how long the reply can be.
What happens when you hit the limit? In chat apps, older parts of the conversation are dropped or summarized, which is why a long chat seems to "forget" instructions from the start. Through the API, an input that's too long is typically rejected with an error.
A bigger window isn't a cure-all:
When a model writes, it predicts a probability for every possible next token. Temperature changes how it chooses among them:
The exact range depends on the provider. Some APIs use 0–1, others 0–2, and some newer reasoning models don't let you change temperature at all.
| Task | Suggested temperature | Why |
|---|---|---|
| Extracting fields from invoices | Low | You want the same answer every time |
| Writing SQL or code | Low | Correctness over variety |
| Summarizing reports | Low to medium | Faithful to the source |
| Drafting marketing copy | Medium to high | Variety is useful |
| Brainstorming ideas | High | You want options you wouldn't expect |
A common misconception: temperature 0 is not guaranteed to be perfectly deterministic. Outputs can still differ slightly between runs, because of how the model is run on the provider's hardware. For repeatable pipelines, combine low temperature with strict output formats and validation.
The same principles apply across ChatGPT, Claude and Gemini, even though limits and prices differ. See our practical comparison for data work.
These settings are the foundation for prompt engineering and for building LLM applications. For a structured path, see our 30-day prompt engineering study plan. For a short, hands-on course on using LLM APIs, including settings like these, try ChatGPT Prompt Engineering for Developers. To understand how LLMs work under the hood, Generative AI with Large Language Models goes deeper. If you're still placing LLMs in the bigger picture, start with AI vs Machine Learning vs Deep Learning vs Generative AI. Browse all Generative AI & LLM courses.
How many words is 1,000 tokens? Roughly 750 English words. Other languages, and text with lots of numbers or code, usually give fewer words per token.
Does a bigger context window make a model smarter? No. It lets the model see more at once, but it doesn't improve its reasoning. Sending focused, relevant context often beats sending everything.
Why does ChatGPT give different answers to the same question? Because output is sampled with some randomness, controlled by temperature. The chat apps use settings tuned for natural conversation, and you usually can't change them there. Through the API you can.
What temperature should I use for data analysis? Low, around 0 to 0.3 where the API allows it. You want consistent, faithful output, not creativity.
Tokens are the units LLMs read, write and bill by. The context window is how many tokens fit into a single request. Temperature controls how predictable the output is. Keep prompts lean, match temperature to the task, and do the token math before large jobs. Those three habits explain most "why did the AI do that?" moments and prevent most unexpected bills.
A hands-on course covering how LLMs are built, fine-tuned, and deployed, taught by AWS AI practitioners through DeepLearning.AI on Coursera.
A free DeepLearning.AI short course teaching developers to write effective prompts and build simple applications on the OpenAI API using.