26 - LLM Fundamentals¶
Previous: 25 - Hugging Face | Index: All guides | Next: 27 - LLM APIs
Quick reference for how large language models work: tokens, next-token prediction, training, context windows, sampling, reasoning, limitations, costs and how to choose a model. No code needed to understand this guide; it is the mental foundation for every other AI guide.
Last verified: 2026-09-27. For newer changes, check the Official docs links in the Introduction.
Introduction¶
Before you start¶
You should know: nothing technical is required. If you have read 24 - PyTorch, the idea of a trained neural network will feel familiar, but this guide explains what you need in plain terms.
The problem it solves: LLMs can seem like magic or like a search engine, and both views lead to mistakes: trusting invented facts, being surprised by the bill, or not understanding why an answer changes between runs. A few core ideas (tokens, the context window, next-token prediction, sampling) explain almost every behaviour you will see.
Before LLMs: language software was built task by task: one model for translation, another for sentiment, another for spam, each trained on its own labelled data, and chatbots followed hand-written rules. The Transformer architecture (2017) and training ever larger models on huge amounts of text showed that one general model could do all these tasks from instructions alone, which led to ChatGPT (2022) and Claude.
Think of it like: an extremely well-read autocomplete. It has read a huge library and learned how text continues; ask it something and it writes the most fitting continuation. That makes it fluent and knowledgeable, and also explains why it can sound confident while being wrong.
What is an LLM?¶
A large language model (LLM) is a neural network trained on huge amounts of text to do one thing: predict the next token (piece of a word) given all the text before it. Repeating that prediction over and over produces sentences, code, summaries and answers. Models like Claude, GPT, Gemini, Llama, Qwen and Mistral are all LLMs. Because predicting text well requires modelling grammar, facts, reasoning patterns and intent, these models end up able to follow instructions, write code, translate, analyse and use tools.
Mental model: the autocomplete that learned to think¶
Input (the "context"): "The capital of France is"
|
v
+-----------------------------------------------+
| LLM: billions of learned weights |
| looks at ALL tokens in the context at once |
| (attention) and scores every possible next |
| token |
+-----------------------------------------------+
|
v
Probabilities: " Paris" 0.92 " a" 0.03 " the" 0.02 " Lyon" 0.001 ...
|
pick one (sampling) -> " Paris"
|
append it, repeat until a stop token or max length
v
Output: "The capital of France is Paris."
Three consequences of this design explain most LLM behaviour:
- Everything is text in context. The model only knows what is in its weights (training) plus what you put in the prompt. It has no memory between API calls unless you send the history again.
- It generates plausible text, not verified truth. When it does not know, it can still produce a fluent, confident, wrong answer (hallucination). Give it sources (RAG) and tools to ground answers.
- The prompt steers the probabilities. Clear instructions, examples and context shift which tokens become likely. That is why prompt engineering works.
The layers of an AI application¶
+------------------------------------------------------------------+
| Your app / UI (Streamlit, web app, chat, CLI) |
+------------------------------------------------------------------+
| Orchestration (agents, tool loops, RAG, workflows)|
+------------------------------------------------------------------+
| Context you provide (system prompt, history, documents, |
| tool results, examples) |
+------------------------------------------------------------------+
| Model (Claude, GPT, open models) |
| accessed via an API (hosted) or run locally (Ollama / vLLM) |
+------------------------------------------------------------------+
Most of your engineering work happens in the middle two layers: deciding what goes into the context and what to do with the output.
Key terms¶
| Term | Meaning |
|---|---|
| Token | Chunk of text (word piece) the model reads and writes; ~4 characters of English |
| Context window | Max tokens the model can consider at once (prompt + output) |
| Prompt | Everything you send: system instructions, messages, documents |
| Completion / output | Tokens the model generates |
| Inference | Running the model to get output |
| Parameters / weights | Learned numbers in the model (billions) |
| Pretraining | Learning to predict text on a huge corpus |
| Fine-tuning / post-training | Further training to follow instructions, be helpful and safe |
| Temperature | Randomness of token choice |
| Hallucination | Confident but false output |
| Grounding | Basing answers on provided sources / tools |
| Knowledge cutoff | Date after which the model has no training data |
| Multimodal | Can take images / PDFs / audio as input |
| Reasoning / thinking | Model spends tokens thinking before answering |
| Latency | Time until the (first) response |
Where it fits: the concepts behind 27 - LLM APIs, 28 - Prompt Engineering, 30 - Embeddings, 31 - RAG, 32 - AI Agents. Implementation details: 24 - PyTorch, 25 - Hugging Face.
Official docs¶
Where to read the latest, authoritative documentation:
| Resource | Link |
|---|---|
| Claude documentation | https://platform.claude.com/docs |
| Claude models overview | https://platform.claude.com/docs/en/about-claude/models/overview |
| Context windows | https://platform.claude.com/docs/en/build-with-claude/context-windows |
| Claude pricing | https://platform.claude.com/docs/en/about-claude/pricing |
| OpenAI platform docs | https://platform.openai.com/docs |
| Hugging Face LLM course | https://huggingface.co/learn |
Contents¶
- Tokens
- Next-Token Prediction
- The Transformer and Attention (Intuition)
- How LLMs Are Trained
- Context Window
- Messages and Roles
- Sampling: Temperature, top_p, top_k
- Reasoning / Thinking Models
- Embeddings (Meaning as Numbers)
- Multimodal Models
- Hallucinations and Grounding
- What LLMs Are Good and Bad At
- Cost and Latency
- Closed vs Open Models
- Choosing a Model
- Ways to Adapt a Model to Your Task
- LLM Application Patterns
- Glossary of AI Buzzwords
- Common Misconceptions
- Try It
1. Tokens¶
The units an LLM reads and writes. A tokenizer splits text into pieces from a fixed vocabulary (common words are one token, rare words several).
Use it for estimating cost, fitting text into the context window, understanding odd behaviour (spelling, counting letters).
"Tokenization is surprisingly useful!"
-> ["Token", "ization", " is", " surprisingly", " useful", "!"] 6 tokens
| Rule of thumb (English) | Value |
|---|---|
| 1 token | ~4 characters, ~0.75 words |
| 100 tokens | ~75 words |
| 1 page of text | ~500 to 700 tokens |
| 1,000 lines of code | ~10,000 tokens |
Other languages, numbers and code usually need more tokens per word. Each provider has its own tokenizer; use its token-counting endpoint for exact numbers. Because models see tokens, not letters, tasks like "count the r's in strawberry" can be surprisingly hard.
2. Next-Token Prediction¶
The one operation an LLM performs. Given the tokens so far, output a probability for every token in the vocabulary; pick one; append; repeat.
Use it for understanding streaming, stop reasons and why output length drives cost and time.
- Generation is sequential: output token 500 needs tokens 1 to 499 first. Long outputs take longer.
- Streaming simply shows each token as soon as it is picked.
- Generation stops at a stop token (the model decides it is done), a stop sequence you define, or max tokens (a hard cut; the answer may be incomplete).
- The model is stateless between requests; "memory" in chat apps is the app re-sending history.
3. The Transformer and Attention (Intuition)¶
The neural network architecture behind all modern LLMs (2017, "Attention Is All You Need"). Each token looks at every other token in the context and decides which ones matter for it (attention); many stacked layers refine the meaning step by step.
Use it for knowing why context matters and why long contexts cost more.
"The trophy did not fit in the suitcase because it was too big."
^^
Attention for "it" -> strongly on "trophy" (big thing that does not fit), weakly on "suitcase"
- Each token becomes a vector (embedding); layers of attention + feed-forward networks transform these vectors.
- Attention compares tokens pairwise, so compute grows with context length; long prompts cost more and can be slower.
- Models can miss details buried in very long contexts; put key instructions and questions where they are easy to find (see 28 - Prompt Engineering).
4. How LLMs Are Trained¶
The stages that turn a random network into a helpful assistant. Pretraining for knowledge and language; post-training for following instructions, helpfulness and safety.
Use it for understanding knowledge cutoffs, why models refuse some requests, and what fine-tuning can and cannot do.
1. PRETRAINING trillions of tokens of text and code -> "base model": great at continuing text,
(months, huge GPUs) objective: predict next token knows a lot, does not follow instructions well
2. SUPERVISED curated (instruction, ideal answer) pairs -> follows instructions, chat format
FINE-TUNING (SFT)
3. PREFERENCE / RL humans / AI rank answers; reinforcement -> more helpful, honest, safe;
(RLHF, RLAIF, learning on feedback and verifiable tasks better at reasoning and tool use
constitutional AI)
- Knowledge cutoff: the model knows nothing after its training data ends; give recent facts in the prompt or via search tools.
- Your data is not learned from API calls in normal use; to use private knowledge, put it in the context (RAG) or fine-tune.
5. Context Window¶
The maximum number of tokens the model can handle in one request (input + output). Everything the model "sees" must fit: system prompt, tools, history, documents, and the answer it writes.
Use it for designing chat history handling, RAG chunk counts and document processing.
+------------------------------ context window -------------------------------+
| system prompt | tool definitions | conversation history | documents | query | -> output tokens
+-----------------------------------------------------------------------------+
- Current frontier models offer very large windows (hundreds of thousands to 1M tokens), smaller open models often 8K to 128K.
- Bigger context = more cost per call and sometimes weaker attention to details. Send what is relevant, not everything.
- Long conversations: summarise or trim old turns, or use provider features (compaction, context editing).
- Max output tokens is a separate, smaller limit on how much the model can write in one response.
6. Messages and Roles¶
How chat APIs structure the context. A list of messages, each with a role; the system prompt sets behaviour for the whole conversation.
Use it in every API call and chat app.
| Role | Contains |
|---|---|
system |
Instructions, persona, rules, background (set by the developer) |
user |
What the user / your app asks |
assistant |
What the model answered before (sent back as history) |
| tool results | Output of tools the model asked to call (see 29 - Tool Use) |
system: You are a support agent for Acme. Answer only about Acme products.
user: My order hasn't arrived.
assistant: I'm sorry! Could you share your order number?
user: A-1042 <- the model sees ALL of this every time
7. Sampling: Temperature, top_p, top_k¶
How the next token is chosen from the probabilities. Temperature reshapes the distribution; top_p / top_k cut off unlikely tokens before sampling.
Use it for tuning consistency vs variety (on models and APIs that expose these settings).
Probabilities for next word: "Paris" 0.70 "France" 0.15 "a" 0.10 "banana" 0.05
temperature 0 -> always "Paris" (deterministic-ish, factual tasks)
temperature 0.7 -> usually "Paris" (balanced default)
temperature 1.5 -> sometimes "banana" (creative, risky)
| Setting | Low value | High value |
|---|---|---|
temperature |
Focused, repeatable | Varied, creative, more errors |
top_p |
Only the most likely tokens | Wider choice |
top_k |
Choose among few tokens | Choose among many |
Some newer reasoning models fix these internally and do not accept them; you steer with instructions and effort settings instead.
8. Reasoning / Thinking Models¶
Models that generate internal reasoning before the final answer. The model spends extra tokens working through the problem step by step; you can often control how much (effort / thinking settings).
Use it for maths, coding, planning, multi-step analysis, agents. Not needed for simple lookups or classification.
- More thinking = usually better answers on hard problems, but more tokens (cost) and more latency.
- You pay for thinking tokens even if they are hidden or summarised.
- "Think step by step" in the prompt gave non-reasoning models some of this benefit; reasoning models do it natively.
9. Embeddings (Meaning as Numbers)¶
A vector (list of numbers) that represents the meaning of a text. An embedding model maps text to a point in a high-dimensional space; similar meanings land close together.
Use it for semantic search, RAG retrieval, clustering, recommendations, deduplication.
"How do I reset my password?" -> [0.12, -0.40, 0.88, ...] \
"I forgot my login" -> [0.10, -0.38, 0.90, ...] } close together (similar meaning)
"Shipping takes three days" -> [-0.70, 0.22, 0.05, ...] far away
Embedding models are separate, smaller models from chat LLMs. Details: 30 - Embeddings and Vector DBs.
10. Multimodal Models¶
Models that accept images, PDFs, audio or video as input (and some that produce images / audio). Non-text inputs are converted into tokens / embeddings the model can attend to alongside text.
Use it for reading invoices and screenshots, charts, handwriting, diagrams, voice apps.
Images and PDF pages cost tokens too (roughly proportional to resolution / page count).
11. Hallucinations and Grounding¶
Fluent, confident output that is false or made up (fake citations, wrong numbers, invented APIs). The model generates plausible text; without a source, plausible can be wrong.
Use it for designing any app where correctness matters.
| Technique | Why it helps |
|---|---|
| Provide sources in the prompt (RAG) | Model answers from given text, not memory |
| Ask for quotes / citations | Answers are checkable |
| Allow "I don't know" | Removes pressure to invent |
| Tools for facts (search, calculator, DB) | Real data instead of guesses |
| Structured outputs + validation | Catch malformed or impossible values |
| Evals and human review | Measure and catch errors (35) |
| Stronger model / more reasoning | Fewer errors on hard tasks |
12. What LLMs Are Good and Bad At¶
Realistic expectations. Strong at language and pattern tasks; weak where exactness, fresh facts or hidden state matter.
Use this when deciding whether an LLM is the right tool, and where to add tools or code.
| Good at | Weak at (add tools / code) |
|---|---|
| Summarising, rewriting, translating | Exact arithmetic on big numbers (use code) |
| Extracting structured data from messy text | Facts after the cutoff (use search / RAG) |
| Classifying, tagging, routing | Counting characters / exact string ops |
| Writing and explaining code | Guaranteed determinism |
| Brainstorming, drafting | Knowing your private data (put it in context) |
| Reasoning over provided documents | Long-running state without memory design |
| Following multi-step instructions | Tasks needing verified truth without sources |
13. Cost and Latency¶
How LLM usage is priced and what makes it slow. You pay per input token and per output token (output is several times more expensive); latency grows with output length and reasoning.
Use it for budgeting, choosing models, optimising apps.
cost of one call = input_tokens x input_price + output_tokens x output_price (prices per 1M tokens)
Example: 3,000 input + 500 output tokens at $5 / $25 per 1M
= 3,000 x 0.000005 + 500 x 0.000025 = $0.015 + $0.0125 = $0.0275
| Lever | Saves |
|---|---|
| Prompt caching (reuse the same long prefix) | Large share of repeated input cost and latency |
| Shorter prompts / fewer retrieved chunks | Input cost |
| Ask for concise output, set sensible max tokens | Output cost and latency |
| Smaller / cheaper model for simple steps | Cost and latency |
| Lower reasoning effort for easy tasks | Thinking tokens |
| Batch API for offline jobs | Often ~50% cheaper |
| Streaming | Perceived latency (first words appear fast) |
14. Closed vs Open Models¶
Hosted proprietary models vs models whose weights you can download. Closed models are used through a provider's API; open-weight models run on your hardware or a host of your choice.
Use it for balancing quality, privacy, cost and control.
| Closed / hosted (Claude, GPT, Gemini) | Open-weight (Llama, Qwen, Mistral, Gemma, DeepSeek) | |
|---|---|---|
| Access | API key, pay per token | Download weights (Hugging Face, Ollama) |
| Quality | Usually the strongest | Good and improving; best big ones need big GPUs |
| Setup | None | You run and scale it (or use a host) |
| Data privacy | Provider's terms; enterprise options, regional hosting | Fully in your control |
| Cost pattern | Per token | Per hour of hardware |
| Customisation | Prompting, some fine-tuning | Full fine-tuning, quantization |
15. Choosing a Model¶
Picking the right model for each task. Start with a strong model to prove the task works, measure with evals, then optimise cost / latency.
Use it in every new feature.
- Prototype with a top-tier model so model weakness is not the problem.
- Build a small eval set (35) with real examples.
- Try cheaper / faster models or lower effort; keep the cheapest that passes your quality bar.
- Consider hard constraints: data residency, offline, context length, multimodal input, tool use support.
| Task | Typical choice |
|---|---|
| Complex reasoning, coding, agents | Frontier model with reasoning |
| High-volume classification / extraction | Smaller, faster model (or fine-tuned small model) |
| Sensitive data, offline | Open model locally (36) or enterprise cloud hosting |
| Search / RAG retrieval | Embedding model + chat model |
16. Ways to Adapt a Model to Your Task¶
The options for making a general model good at your specific job, from cheapest to most effort. Most problems are solved by the first three; fine-tuning is for specific cases.
Use it for deciding how to improve quality.
effort / cost ->
1. Prompt engineering clear instructions, examples, output format (minutes) [28]
2. Context / RAG give the model the right documents at runtime (days) [31]
3. Tools / agents let the model fetch data and take actions (days) [29] [32]
4. Fine-tuning change model weights with your examples (weeks) [37]
5. Train from scratch almost never for applications (months, $$$)
| Problem | Usually solved by |
|---|---|
| "It doesn't follow my format" | Prompting, structured outputs |
| "It doesn't know our documents" | RAG |
| "It needs live data / to act" | Tools |
| "It must adopt a very specific style / narrow skill cheaply at scale" | Fine-tuning |
17. LLM Application Patterns¶
The common shapes of LLM apps, from simple to complex. Add complexity only when the simpler pattern is not enough.
Use it for designing a new AI feature.
| Pattern | How it works | Example |
|---|---|---|
| Single call | Prompt in, answer out | Summarise an email |
| Structured extraction | Prompt + schema -> JSON | Invoice fields |
| Chain / workflow | Fixed sequence of LLM calls and code | Draft -> critique -> revise |
| Routing | Classify first, then pick the handler | Support ticket routing |
| RAG | Retrieve documents, then answer with them | Chat with company docs |
| Tool use | Model calls functions you define | Look up order status |
| Agent | Model decides steps in a loop until done | Research and write a report |
| Multi-agent | Several agents with roles cooperate | Planner + workers + reviewer |
18. Glossary of AI Buzzwords¶
Quick definitions of terms you will hear. One line each; follow the links for depth.
Use it for reading docs, job posts, blog posts.
| Term | Meaning |
|---|---|
| Foundation model | Large general-purpose pretrained model |
| Prompt engineering | Designing inputs to get better outputs (28) |
| Context engineering | Deciding everything that goes into the context (instructions, docs, tools, memory) |
| Few-shot / zero-shot | With / without examples in the prompt |
| Chain of thought | Reasoning step by step before answering |
| Structured output | Output constrained to a schema (JSON) |
| Function calling / tool use | Model requests to run your functions (29) |
| RAG | Retrieval-augmented generation (31) |
| Vector database | Stores embeddings for similarity search (30) |
| Agent | LLM in a loop that uses tools to reach a goal (32) |
| MCP | Model Context Protocol: standard way to connect tools / data to AI apps (34) |
| Guardrails | Checks on inputs / outputs for safety and policy (38) |
| Evals | Systematic tests of model / app quality (35) |
| LLM-as-judge | Using a model to grade outputs |
| Fine-tuning / LoRA | Training a model further / cheap way to do it (37) |
| Quantization | Storing weights with fewer bits to save memory (36) |
| Distillation | Training a small model to imitate a big one |
| Inference server | Software serving a model over an API (vLLM, Ollama) |
| Prompt injection | Malicious text that tries to override your instructions (38) |
| Grounding | Tying answers to sources / data |
| Latency / TTFT | Response time / time to first token |
| Throughput | Tokens or requests processed per second |
19. Common Misconceptions¶
Beliefs that lead to bad designs. Each has a short correction.
Use it for sanity check when an AI feature misbehaves.
| Misconception | Reality |
|---|---|
| "The model remembers our previous chats" | Only if your app sends the history (or uses a memory feature) |
| "It learned from my API calls" | Standard API use does not train the model on your data (check provider terms) |
| "It looked it up online" | Only if you gave it a search / fetch tool |
| "Temperature 0 means always correct" | It means more consistent, not more truthful |
| "Bigger context always helps" | Irrelevant text can distract and costs more |
| "Fine-tuning teaches it our documents" | Fine-tuning shapes behaviour; RAG is better for facts that change |
| "It says it's sure, so it's right" | Confidence in text is not reliability; verify |
| "Agents are always better" | They are slower, costlier and less predictable; use the simplest pattern that works |
20. Try It¶
Short exercises to practise this guide. Try each task yourself first, then open the solution.
Use it right after reading the guide, or later as a quick self-test.
Exercise 1: Estimate a bill¶
A 2,000-word document goes in and a 300-word answer comes out at $5 / $25 per million tokens. Roughly what does one call cost?
Solution
Tokens ~ words / 0.75: input ~ 2,700 tokens, output ~ 400 tokens. Cost ~ 2,700 x 5 / 1,000,000 + 400 x 25 / 1,000,000 = $0.0135 + $0.01 = about $0.024 per call, so about $23.50 per day at 1,000 calls. Do this estimate before launch.
Exercise 2: Pick the technique¶
Prompting, RAG, tools or fine-tuning? (a) answer from 5,000 policy PDFs, (b) show today's order status, (c) always reply in a strict JSON format, (d) cheap high-volume classification in a fixed style.
Solution
(a) RAG, (b) tools (live data), (c) structured outputs / prompting, (d) prompting first; fine-tune a small model only if volume and cost justify it.
Exercise 3: The forgetful chatbot¶
Your chatbot forgets what the user said two messages ago. Why?
Solution
The API is stateless: the model only sees what you send in messages. Your app must store the history and send it with every request (and trim or summarise it when it gets long).
Previous: 25 - Hugging Face | Index: All guides | Next: 27 - LLM APIs