Fundamentals

What Is a Large Language Model, Actually? Tokens, Parameters, and Context Windows Explained

What Is a Large Language Model, Actually? Tokens, Parameters, and Context Windows Explained

An earlier article in this series placed large language models inside a nesting doll — machine learning contains deep learning contains LLMs — which is true but leaves the actual object underspecified. What is an LLM, mechanically, once you stop treating it as a black box? Three concepts do almost all the explanatory work: tokens, parameters, and the context window.

Tokens: the model doesn’t read words

An LLM doesn’t process text as words or characters — it processes tokens, chunks of text produced by a tokenizer that sit somewhere between characters and whole words. Common words often become a single token; rarer words get split into multiple sub-word pieces (“unbelievable” might become “un” + “believ” + “able”). This matters practically: pricing for API access to commercial LLMs is usually billed per token, not per word or character, and a model’s stated “context length” (say, 128,000 tokens) is measured in tokens, which for English text works out to roughly 75,000–100,000 words — but considerably fewer for languages like Chinese, where tokenization is denser per character.

The model’s actual job, stripped to its core, is: given a sequence of tokens, predict a probability distribution over what the next token is likely to be. Generate text by repeatedly sampling from that distribution, appending the result, and predicting again. That’s it — that single mechanical loop, run by a large enough model trained on enough data, is the entire generative process behind everything from a one-line answer to a 2,000-word essay.

Parameters: the numbers that got adjusted during training

A parameter is a single numerical weight inside the neural network that gets adjusted during training. When you see a model described as having “70 billion parameters” or “1.8 trillion parameters” (widely reported, though never officially confirmed, for GPT-4), that’s the count of these individual adjustable numbers — not a count of “facts” or “rules” the model knows in any explicit, inspectable sense. Training is the process of nudging all of these numbers, gradually, so that the model gets progressively better at its one job: predicting the next token correctly across a training set that, for frontier models, runs into the trillions of tokens scraped from text, code, and other sources.

More parameters generally means more capacity to capture complex patterns — which is part of why larger models tend to perform better — but parameter count alone doesn’t fully determine quality. How much data the model was trained on, how that data was curated, and the specific training techniques used all matter too. A smaller model trained well on carefully chosen data can outperform a larger model trained carelessly; parameter count is one input among several, not a single scoreboard.

It’s worth being precise about what “knowing” means here: none of a model’s training data is stored verbatim anywhere retrievable — there’s no database row for “Paris is the capital of France” that the model looks up. Instead, the statistical relationship between those tokens got baked into the numerical weights during training, spread diffusely across billions of parameters in a way no one can point to and read directly. This is why LLM behavior is genuinely hard to fully explain or predict from first principles, even by the people who built the model — the “knowledge” is real in its effects but isn’t stored as inspectable facts.

The context window: working memory, not long-term memory

The context window is the maximum number of tokens a model can consider at once — both the conversation history and any documents you’ve provided, plus the response it’s generating. This is best understood as working memory, not storage: once a conversation exceeds the context window, earlier content has to be dropped or summarized to make room, and the model has no access to it unless it’s explicitly reintroduced.

This distinction matters because it’s routinely confused with training. Training is where a model’s general capabilities get baked into its parameters, over weeks of computation on enormous datasets, before it’s ever released. The context window is what happens after that, at the moment you’re using the model — it doesn’t change the model’s weights at all. When you tell a model something in conversation, it hasn’t “learned” it in the training sense; it’s just present in this session’s working memory, gone once the context window fills up or the conversation ends. This is also the direct mechanical reason “prompt engineering” works at all — you’re not modifying the model, you’re controlling what’s present in its limited working memory at generation time.

What this architecture is actually good and bad at

Being precise about tokens, parameters, and context windows explains, rather than just asserts, a few things people often get hand-wavy about. LLMs are genuinely strong at tasks where the next plausible token is well-constrained by patterns in training data — fluent writing, code that follows common patterns, translation, summarization. They’re structurally weak at anything requiring information that either wasn’t in training data, changed after training data was collected, or requires precise lookup rather than plausible generation — which is exactly why modern AI products pair LLMs with the tool-calling and MCP-based external-data mechanisms covered elsewhere in this series, rather than relying on the model’s frozen training knowledge alone for anything time-sensitive or fact-critical.

None of this makes an LLM “just autocomplete,” a dismissal that undersells what emerges from this mechanism at scale — coherent multi-step reasoning, working code, translation between dozens of languages. But it also isn’t a database, and it isn’t a mind with settled interiority either. It’s a specific, well-understood mechanism — next-token prediction over a fixed context, using knowledge compressed into billions of numerical weights — that happens to produce results good enough to build products, and companies, on top of.

Frequently asked questions

Is an LLM the same thing as ChatGPT?

No. An LLM is the underlying model — GPT-4, Claude, Gemini, Llama are all LLMs. ChatGPT is a product: a chat interface, plus safety tuning (RLHF), plus product features (memory, file uploads, tool access) built around an LLM. The LLM is the engine; the chat app is the car built around it.

Does an LLM understand what it's saying?

This is genuinely disputed among researchers, not a settled question with a clean answer. What's not disputed: an LLM is trained on one objective — predict the next token — and everything it can do emerges from that single training goal applied at enormous scale. Whether the internal representations that make good predictions possible constitute 'understanding' in any meaningful sense is an open philosophical and empirical question, and treating it as obviously settled in either direction overstates the evidence.

Why do LLMs sometimes make things up confidently?

Because the training objective is 'predict a plausible next token,' not 'only say things you can verify.' A model trained this way has no built-in mechanism to distinguish a fact it saw thousands of times in training from a plausible-sounding pattern it's extrapolating — both get produced by the same process, with the same apparent confidence. This is the direct, mechanical explanation for what's commonly called 'hallucination.'