Tokens, Context Windows and Temperature: LLM Settings Explained

Tokens, Context Windows and Temperature: LLM Settings Explained

Tokens are the chunks of text a language model reads and writes, often parts of words. A rough rule for English is that 1 token is about ¾ of a word. The context window is the maximum number of tokens a model can consider at once, including your instructions, documents, chat history and its own reply. Temperature controls how random the output is: low values give focused, consistent answers, and higher values give more varied ones.

These three settings explain a lot of confusing LLM behavior. They're why a long chat seems to "forget" earlier instructions, why the same prompt gives different answers, and why an API bill can be larger than expected. Once you understand them, you'll prompt better, choose models better and estimate costs properly.

Tokens: How Models Read Text

Language models don't see letters or whole words. Before processing, text is split into tokens using a vocabulary the model learned during training. Common words are usually a single token. Rarer words are split into pieces.

For example, a sentence like "Data analysts love spreadsheets" might be split into something like:

Data · analysts · love · spread · sheets

The exact split depends on the model's tokenizer, and different vendors use different ones.

Useful rules of thumb for English:

  • 1 token ≈ 4 characters ≈ ¾ of a word
  • 100 tokens ≈ 75 words
  • A 1,500-word article ≈ 2,000 tokens

Two practical consequences:

  1. Non-English text often uses more tokens per word. Many tokenizers were trained mostly on English, so languages like Vietnamese, with diacritics and different word structures, often need more tokens for the same meaning. That affects both cost and how much fits in the context window.
  2. Numbers and code tokenize unpredictably. A long number may be split into several tokens. That's one reason LLMs sometimes make arithmetic mistakes that a calculator never would.

Why Tokens Matter: Cost

API pricing is per token, usually quoted per million tokens, with separate prices for input (what you send) and output (what the model writes). Output tokens usually cost several times more than input tokens.

Worked example. You want to summarize 500 customer support tickets. Each ticket plus your instructions is about 800 tokens, and each summary is about 150 tokens. Using an illustrative price of $3 per million input tokens and $15 per million output tokens:

Tokens Price per million Cost
Input 500 × 800 = 400,000 $3 $1.20
Output 500 × 150 = 75,000 $15 $1.13
Total ≈ $2.33

Prices vary widely between models, and a smaller model can cost a fraction of this. The method stays the same: estimate tokens in and out, multiply by the price, then test on a small sample before running the full batch.

Context Windows: The Model's Working Memory

The context window is the total number of tokens the model can take into account in a single request. Everything has to fit in it:

  • System instructions
  • The full conversation so far
  • Any documents or data you paste in
  • The model's reply

Context windows have grown fast. Many current models accept hundreds of thousands of tokens, and some over a million. Check each model's documentation, because limits differ, and there's usually a separate, smaller cap on how long the reply can be.

What happens when you hit the limit? In chat apps, older parts of the conversation are dropped or summarized, which is why a long chat seems to "forget" instructions from the start. Through the API, an input that's too long is typically rejected with an error.

A bigger window isn't a cure-all:

  • Cost: every token you send is billed, every time. Pasting a 100-page report into each question multiplies your cost.
  • Attention: research has found that models can pay less attention to information buried in the middle of very long inputs. Put key instructions at the start or end.
  • Relevance: sending only the relevant passages often gives better answers than sending everything. That's the idea behind RAG (retrieval-augmented generation): retrieve the right chunks instead of stuffing the whole library into the prompt.

Temperature: Focused vs Creative

When a model writes, it predicts a probability for every possible next token. Temperature changes how it chooses among them:

  • Low temperature (around 0–0.3): strongly favors the most likely tokens. Output is focused, consistent and repeatable.
  • Medium (around 0.5–0.8): a balance. Typical default settings sit around here.
  • High (around 1 and above): gives less likely tokens more of a chance. Output is more varied and creative, and more likely to drift or make things up.

The exact range depends on the provider. Some APIs use 0–1, others 0–2, and some newer reasoning models don't let you change temperature at all.

Task Suggested temperature Why
Extracting fields from invoices Low You want the same answer every time
Writing SQL or code Low Correctness over variety
Summarizing reports Low to medium Faithful to the source
Drafting marketing copy Medium to high Variety is useful
Brainstorming ideas High You want options you wouldn't expect

A common misconception: temperature 0 is not guaranteed to be perfectly deterministic. Outputs can still differ slightly between runs, because of how the model is run on the provider's hardware. For repeatable pipelines, combine low temperature with strict output formats and validation.

Other Settings You'll See

  • Max tokens / max output tokens: caps the length of the reply. If it's set too low, answers get cut off mid-sentence.
  • Top-p (nucleus sampling): another way to control randomness. The model only samples from the smallest set of tokens whose combined probability reaches p. Providers usually recommend adjusting temperature or top-p, not both.
  • Stop sequences: strings that tell the model to stop writing, which is useful for structured outputs.
  • System prompt: standing instructions that apply to the whole conversation. They use tokens on every request too.

What This Means in Practice

  1. Keep prompts lean. Remove boilerplate and irrelevant data. It lowers cost and often improves quality.
  2. Restate key instructions in long chats, or start a new chat when you change tasks.
  3. Match temperature to the task. Use low for analysis and code, and higher for ideas and drafts.
  4. Estimate cost before batch jobs, using the token math above.
  5. Check work involving numbers. Tokenization is one reason LLMs struggle with exact arithmetic. For data work, ask the model to write code that does the calculation, as we recommend in How to Use ChatGPT for Data Analysis.

The same principles apply across ChatGPT, Claude and Gemini, even though limits and prices differ. See our practical comparison for data work.

What to Learn Next

These settings are the foundation for prompt engineering and for building LLM applications. For a structured path, see our 30-day prompt engineering study plan. For a short, hands-on course on using LLM APIs, including settings like these, try ChatGPT Prompt Engineering for Developers. To understand how LLMs work under the hood, Generative AI with Large Language Models goes deeper. If you're still placing LLMs in the bigger picture, start with AI vs Machine Learning vs Deep Learning vs Generative AI. Browse all Generative AI & LLM courses.

Frequently Asked Questions

How many words is 1,000 tokens? Roughly 750 English words. Other languages, and text with lots of numbers or code, usually give fewer words per token.

Does a bigger context window make a model smarter? No. It lets the model see more at once, but it doesn't improve its reasoning. Sending focused, relevant context often beats sending everything.

Why does ChatGPT give different answers to the same question? Because output is sampled with some randomness, controlled by temperature. The chat apps use settings tuned for natural conversation, and you usually can't change them there. Through the API you can.

What temperature should I use for data analysis? Low, around 0 to 0.3 where the API allows it. You want consistent, faithful output, not creativity.

Bottom Line

Tokens are the units LLMs read, write and bill by. The context window is how many tokens fit into a single request. Temperature controls how predictable the output is. Keep prompts lean, match temperature to the task, and do the token math before large jobs. Those three habits explain most "why did the AI do that?" moments and prevent most unexpected bills.

Enjoyed this article?

Share it with your network

Listings related to Tokens, Context Windows and Temperature: LLM Settings Explained