A language model is a frozen array of numbers that can do exactly one thing: score every possible next word fragment. The memory, the tools, the apparent deliberation are all machinery built around that single act.

A machine with no memory of the person asking, no plan, and no ability to change itself is producing sentences that look like thought. That is the thing worth explaining. The system cannot learn from the conversation it is having, holds nothing over from yesterday, and does not decide what to say before it starts saying it. Yet the output arrives in fluent order and often gets the answer right, and it costs real money that scales in ways that seem arbitrary until the mechanism is visible.
Two questions organise everything below. What is physically happening in the gap between the question and the answer? And why does that gap cost what it costs, in time and in money?
The answers separate cleanly into what is not in dispute and what is. The arithmetic is not in dispute: the operations are published, the architectures of several strong models are open, the hardware datasheets are public, and anyone can reproduce the numbers. What remains genuinely contested is the interpretation, whether the arithmetic amounts to understanding, why certain abilities appeared when models got bigger, and how much further the approach goes. What follows stays almost entirely in the first category, and flags the boundary when it arrives at it.
The worked example throughout is Llama 3 70B, an open-weight model from Meta released in 2024. It is used for one reason: every number in its design is published, so nothing here has to be guessed. Frontier commercial models are larger and differ in details, but they are the same kind of object doing the same kind of thing.
A trained language model is a file. Inside that file is a long list of numbers called parameters, or weights: individual values like 0.0412 or negative 1.87, arranged into rectangular grids called matrices. Llama 3 70B contains about 70.6 billion of them. Stored at 16 bits each, which is two bytes, the file weighs roughly 141 gigabytes.
Nothing in that file is a fact, a rule, or a sentence. There is no lookup table of capitals of countries, no stored copy of the training text, no if-then logic anywhere. There are only numbers, and one fixed procedure for multiplying incoming numbers against them.
That procedure is the same every time and does exactly one job. Given a sequence of text, it produces a score for every possible next fragment of text. Not a sentence, not an answer, not a plan. One score per candidate, 128,256 of them, because that is how many entries this model's vocabulary contains. Everything a language model appears to do is that one act, repeated.
Three properties of the file matter more than anything else here. It is frozen: the numbers do not change while it is answering, and nothing said to it is written back. It is large: 141 gigabytes must be moved from memory into the processor to produce a single word fragment, which turns out to set the speed limit for the entire industry. And it is expensive to have made: the arithmetic that fixed those numbers in place ran once, cost a sum Meta has never published, independently estimated between tens and hundreds of millions of dollars, and never runs again.
The word for using the file is inference. The word for making it is training. They are two different activities on the same object, and the report tracks both, because half the confusion about these systems comes from attributing to one what belongs to the other.
Text has to become numbers before any multiplication can happen. The component that does this is the tokenizer, a separate piece of software, trained once and then frozen, that the model itself never touches.
The tokenizer holds a fixed dictionary of text fragments called tokens. Common whole words are single tokens. Rare words split into pieces. Any text at all can be represented, because the dictionary bottoms out in individual bytes, the raw units that all digital text is made of. The dictionary is built by an algorithm called byte pair encoding, which starts with the 256 possible single bytes and then repeatedly finds the most frequent adjacent pair in a large corpus and merges it into a new entry, stopping when the target size is reached. Frequency decides what becomes a unit, so the alphabet is a compression artifact rather than a linguistic one.
Sizes have grown as models have. GPT-2 used 50,257 entries, GPT-4 used about 100,000, and the GPT-4o generation moved to roughly 200,000. Llama 3 uses 128,256. In English text, a token averages about four characters, so a thousand words is roughly 1,300 tokens.
Two consequences follow immediately, and both explain behaviour people find baffling. First, the model is billed and measured in tokens, not words or characters, and the same paragraph in a language poorly served by the dictionary costs more because it fragments into more pieces. Second, the model cannot see inside a token. If the fragment for a word arrives as one indivisible number, counting the letters in that word requires information the model was never given. Spelling failures are not reasoning failures; they are a consequence of the alphabet the system was handed before it saw anything.
The model does not read the word. It receives an integer that stands for the word, and integers have no spelling.
A token identity number is useless for arithmetic. Token 4,721 is not larger than token 300 in any meaningful sense. So the first thing the model does is a lookup: it holds a table with one row per vocabulary entry, and each row is a list of 8,192 numbers. That list is the token's embedding, and 8,192 is this model's hidden dimension, the width of everything that flows through it.
Treating a list of 8,192 numbers as a point in an 8,192-dimensional space is the standard intuition, and it is more than a metaphor. Direction in that space carries meaning, because the training process pushed the numbers into arrangements where tokens used in similar ways ended up pointing in similar directions. Nobody designed the arrangement, and nobody can fully read it.
This is the moment the object stops being text. From here to the end, the model is manipulating a stack of vectors, one per position in the sequence. A prompt of five tokens becomes a grid of 5 by 8,192 numbers, about forty thousand values. That grid is the only thing the model has. It contains no metadata about who is asking, no timestamp, no separate channel for instructions. Everything the system will act on has to be inside those numbers.
This grid is often called the residual stream, and the metaphor practitioners use is a conveyor belt. Each layer reads from the belt, computes something, and adds its result back onto the belt rather than replacing what was there. Information written early can survive to the end, and layers can specialise without having to preserve everything themselves.
Modern language models are built from one repeated block, described in a 2017 paper titled Attention Is All You Need. Llama 3 70B stacks 80 of these blocks. Each block has two halves.
The first half is attention, and it is the only part where positions in the sequence talk to each other. Every other operation in the model treats each position independently. The mechanism, stated as directly as it can be:
From each position's vector, three new vectors are produced by multiplying against three of the model's weight matrices. They are called the query, the key, and the value, written Q, K and V. The names come from database retrieval and are worth taking loosely: the query is what this position is looking for, the key is what each position offers, and the value is what gets passed along if the offer is taken up.
To decide how much position A should draw from position B, take the dot product of A's query with B's key. A dot product multiplies two lists of numbers pairwise and sums the results, producing a single number that is large when the two vectors point in similar directions. That number is a raw relevance score. Do this for every pair and a grid of scores results.
Those raw scores are then divided by the square root of the vector length, which keeps them in a range where the next step behaves, and pushed through the softmax function. Softmax takes any list of numbers and converts it to a list of positive numbers summing to one, exaggerating the gaps as it goes: it exponentiates each value and divides by the total. Now the scores are proportions. Multiply each position's value vector by its proportion, add them up, and the result is a weighted average of the whole sequence, mixed according to relevance the model computed on the spot. In compressed form, the entire operation is one line:
The transpose mark on K means the key matrix is flipped so the multiplication pairs every query with every key, and d is the length of a single head's vectors, 128 in this model. That is the whole of it. There is no search, no retrieval from a store, no symbolic step. There is a similarity score, a normalisation, and a weighted sum.
One structural constraint makes text generation possible. A position may only attend to itself and to positions before it, never after. This is the causal mask, applied by setting the forbidden scores to negative infinity so softmax assigns them zero weight. Training can then run on every position of a sentence at once, while the model stays usable to predict text it has not seen.
The block does not run one attention operation but 64 of them in parallel, called heads, each with its own smaller slice of the vectors and its own learned matrices, so different heads can track different kinds of relationship. And in this model the 64 query heads share only 8 sets of keys and values, an arrangement called grouped query attention. That choice is not about quality. It exists to shrink a memory cost that section VI will show is decisive.
The block's second half is a feed-forward network: each position's vector, independently, gets expanded from 8,192 numbers to 28,672, passed through a simple non-linear function that lets the model represent things a chain of pure multiplications cannot, and squeezed back down to 8,192. About 80 percent of the model's parameters live in these expansions, because this design uses three such matrices per block rather than the original two. If attention is where positions exchange information, the feed-forward half is where whatever was gathered gets processed.
Then it happens again. And 78 more times.
After the eightieth block, the vector at the final position gets multiplied by one last matrix, of shape 8,192 by 128,256. Only that position matters, because only it predicts what comes next. Out comes one number per vocabulary entry. These raw scores are called logits.
Softmax converts the logits to probabilities: 128,256 positive numbers summing to one. This is the model's complete and only output. Not a word. A distribution over every possible word fragment.
Something outside the model then has to choose one, and that choice is a separate, adjustable, and non-neural step. The options are few and worth knowing by name:
Greedy decoding takes the highest-probability token every time. It sounds obviously correct and produces noticeably worse text: research on decoding strategies established that maximising probability at each step yields repetitive, degenerate output, because human language does not follow a pattern of highest-probability next words. Temperature divides all the logits by a constant before softmax, which sharpens the distribution below 1 and flattens it above 1, changing how often unlikely candidates get a chance. Top-p sampling, also called nucleus sampling, restricts the candidates to the smallest set whose probabilities add to at least p, then samples from that set. Its advantage over a fixed-size shortlist is that the set grows when the model is unsure and shrinks when it is confident: most of the probability mass concentrates in a small subset that ranges between one and a thousand candidates.
This is where the randomness in a chatbot lives. It is not in the network. The network is a deterministic function: identical numbers in, identical numbers out. The variety comes from a dice roll applied afterwards, on a distribution the model computed.
One caveat, because it surprises people who set temperature to zero and still see variation. Even greedy decoding is not bit-for-bit reproducible in production, because floating-point addition is not associative, meaning the order in which numbers are summed can change the last digits of the result, and that order depends on how the serving system happened to group requests together. If it helps, the engineers find this annoying too. The distribution is stable. The exact tie-breaking at the bottom of the decimal is not.
The model has now produced one token. To produce the second, the chosen token is appended to the input and the entire process runs again from the top. That repetition is autoregression: every token the answer will ever contain costs one more full pass through the entire network.
A 500-token answer is 500 full passes through 80 blocks and 141 gigabytes of weights. Nothing was planned in advance. At the moment the first word was emitted, the last word did not exist as an intention anywhere in the system, which is also how many of us deliver wedding toasts, except here it is the design. Coherence over a long answer is a property of each step conditioning on everything already written, not of a stored plan.
Done naively this repeats work, because each new pass recomputes the keys and values for every earlier position, which have not changed. So they are stored. That store is the KV cache, the only state a language model has.
The cache splits generation into two phases with completely different characters, and the split explains most of what a person perceives as speed.
Prefill processes the prompt. Every token of it can be handled simultaneously, because they are all already known, so the hardware gets a large parallel job. This is the pause before the first word appears, and it grows with prompt length. Decode produces the answer, one token per pass, each pass depending on the previous one. That steady stream of words barely changes speed with the prompt.
Long prompt, slow start, then normal speed: that is not an artifact of network latency but the shape of the arithmetic.
Producing one token requires roughly two arithmetic operations per parameter, so about 141 billion operations for this model. It also requires reading all 141 gigabytes of weights out of memory. The ratio between those two quantities decides which part of the hardware is the constraint.
A graphics processor has two separate capacities. An NVIDIA H100 in its SXM form factor, the workhorse chip of this era, performs about 989 trillion 16-bit operations per second and moves data from its memory at 3,350 gigabytes per second. Divide one by the other and a threshold appears: to keep the arithmetic units busy, a workload must perform about 295 operations for every byte it reads. Below that, the chip finishes its sums and waits for data. This ratio is called arithmetic intensity, and the threshold is the ridge point.
Generating one token for one user performs about one operation per byte of weights read. That is roughly 300 times below the threshold. The expensive silicon is almost entirely idle, waiting on memory.
| Phase | Operations per byte | Limited by | Chip busy |
|---|---|---|---|
| Prefill reading the prompt, all at once |
200 to 400 | arithmetic | 90 to 95% |
| Decode writing the answer, one token at a time, batched across users |
60 to 80 | memory bandwidth | 20 to 40% |
| H100 threshold where the chip stops waiting |
295 | by definition | 100% |
The escape from that idleness is batching, and it is the economic foundation of every commercial model. If the server processes 64 users' requests together, it reads the 141 gigabytes once and produces 64 tokens from that single read. Memory traffic is unchanged, output multiplies by 64. A single token is cheap because the person is renting a slice of one read that dozens of others are sharing.
Modern serving systems push this further with continuous batching, which reconsiders the group after every single token rather than waiting for the slowest request in a fixed batch to finish, and with paged attention, which allocates KV cache memory in small blocks the way an operating system pages memory, instead of reserving each request's worst-case need in advance. The published result was two to four times the throughput of the previous best systems at the same latency, with vendor benchmarks against unoptimised baselines reporting considerably more.
The cache itself is why long context gets expensive. For this model, holding one token in the KV cache takes 327,680 bytes, or about 328 kilobytes, a figure that follows directly from its published shape: two tensors, 80 layers, 8 key-value heads, 128 numbers each, two bytes per number. Multiply by context length and the growth is unforgiving.
The cache is real enough that it has a price list. Because a conversation's opening tokens are identical across turns, providers store their computed keys and values and reuse them, a technique called prefix caching. Anthropic charges cache reads at one tenth of the base input price, and cache writes at 25% more than base for the default five-minute tier, on the condition that the prefix is byte-identical from one request to the next. A single changing timestamp near the top of a system prompt invalidates everything after it. That is not a billing quirk but the mechanism showing through the invoice: what is being sold is a saved intermediate computation, and a saved computation is only valid for the exact input that produced it.
Four further techniques round out how the physics gets bent, each attacking the same bottleneck from a different angle. Quantization stores weights at 8 or 4 bits instead of 16, which cuts bytes read roughly in half or better; Meta shipped its 405-billion-parameter model in 8-bit form so that it would fit on a single server node. Grouped query attention, met earlier, shrinks the cache by having many query heads share few key-value sets. Mixture of experts splits the feed-forward half into many parallel sub-networks and routes each token to a few of them, so total size and per-token cost decouple: DeepSeek-V3 holds 671 billion parameters but activates 37 billion for each token. And speculative decoding exploits the idle arithmetic directly: a small fast model drafts several tokens, the large model checks them all in one pass, and a rejection rule guarantees the accepted output follows exactly the distribution the large model would have produced alone. Measured on a 70-billion-parameter model, this delivered a two to two and a half times speedup with no quality cost.
Every serving optimisation of the last three years is a variation on one sentence: stop reading the same weights to produce one token.
The weights are frozen at inference, but they were set by a process that used the identical machinery in a different mode. Understanding it removes the last black box.
Pretraining works like this. Take a passage of real text. Run it through the model exactly as described above, but score every position at once rather than only the last. At each position, compare the probability distribution the model produced against the token that actually came next in the real text. The gap is the loss. Then run the arithmetic backwards through all 80 blocks to compute, for every one of the 70 billion parameters, whether nudging it up or down would have shrunk that gap. Nudge them all, slightly. Repeat on the next passage. Repeat trillions of times.
That is the entire objective. Not truthfulness, not helpfulness, not reasoning: predict the next token in human-written text. Everything the model appears to know is a side effect of getting better at that one game, because predicting text well enough eventually requires internal machinery that behaves like knowledge of what the text is about.
The cost follows a rule of thumb that is accurate enough to be used for planning budgets. A forward pass costs about two operations per parameter per token. Adding the backward pass makes it about six. So total training compute is approximately 6 × parameters × tokens. Check it against a real published run: Meta's 405-billion-parameter model was trained on 15.6 trillion tokens using over 16,000 H100 GPUs, at a total of 3.8 × 10²⁵ operations. Six times 405 billion times 15.6 trillion is 3.79 × 10²⁵. The rule lands within one percent of the reported figure.
How the budget is split between size and data is itself a research finding rather than a preference. A 2022 study trained over 400 models and concluded that a compute-optimal run uses about 20 training tokens per parameter, demonstrating it with a 70-billion-parameter model on 1.4 trillion tokens that beat a 280-billion-parameter model trained on 300 billion tokens for the same training compute, and at a quarter of the inference cost. Earlier models had been badly out of balance: GPT-3 used under 2 tokens per parameter.
Practice has since moved deliberately past that optimum, and the reason is a direct consequence of section VII. Compute-optimal minimises the cost of training. But a model is trained once and served billions of times, and a smaller model reads fewer bytes per token forever. So it pays to train a smaller model far longer than the optimum, spending more up front to buy permanently cheaper inference. Llama 3 sits far past compute-optimal at every size, from about 38 tokens per parameter for the 405-billion-parameter model to 221 for the 70-billion and 1,875 for the 8-billion. The smallest model is the clearest case: it was trained far longer than its size warranted precisely because it would be served the most.
Pretraining produces something that continues text but does not answer questions or refuse anything. A second, far cheaper stage called post-training shapes behaviour. Its published recipe is consistent across labs: supervised fine-tuning on demonstrations of good responses, then a reinforcement stage that optimises against human or model preferences between candidate answers. Meta describes iterative rounds of supervised fine-tuning, rejection sampling, which means generating many candidate answers and keeping only those that score well, and direct preference optimization, which adjusts the weights directly from pairs of better and worse answers; DeepSeek describes supervised fine-tuning followed by reinforcement learning. This stage is where helpfulness, format, tone and refusal come from. It moves the weights a little. It does not change the mechanism at all.
| Training | Inference | |
|---|---|---|
| Do the weights change | Yes, that is the point | Never |
| Arithmetic per token | ≈ 6 × parameters | ≈ 2 × parameters |
| Direction | Forward, then backward | Forward only |
| Worked figure | 3.8 × 10²⁵ ops, once | 8.1 × 10¹¹ ops, per token |
| Who pays | The lab, up front | Everyone, forever |
| Bottleneck | Arithmetic and cluster reliability | Memory bandwidth |
| Typical failure | Loss spikes, hardware faults | Cache exhaustion, queueing |
One implication deserves stating flatly, because it is the most common misconception about these systems. A model does not learn from conversations. The weights that answer a question in the afternoon are byte-identical to the ones that answered in the morning. Anything that looks like memory is text being fed back in by the surrounding software, which is the subject of the last section.
What has been described so far is a function: text in, one probability distribution out, no state, no side effects, no knowledge of anything that is not in the text it was handed. The software that surrounds that function has acquired a name, the harness, and a definition traced to a March 2026 post and now in common use: an agent is a model plus a harness, where the harness is every piece of code, configuration and execution logic that is not the model itself, covering the execution of tool calls, parsing what comes back, retrying on failure, routing work to sub-agents and enforcing limits. Practitioners estimate the harness at 80 to 90% of the codebase in a production system, though that figure is an impression rather than a measurement.
Its central job follows from statelessness. Models see only what is assembled for them each turn, so on every single call the harness reconstructs the world: the instructions, the tool definitions, the conversation so far, whatever documents were retrieved, the results of the last action. That assembly is the discipline called context engineering, which is mechanically a curation step that runs before every model call, not a system prompt written once and forgotten.
Around that sits the loop. The model emits text; the harness detects that some of it is a structured request to call a tool; it executes the tool as ordinary code; it appends the result to the transcript; it calls the model again. That cycle, reason then act then observe, is the standard agent pattern, and the interface for declaring tools has been standardised so that capabilities can be attached across vendors: the Model Context Protocol allows an agent's action space to be extended with external capabilities through a uniform interface.
Reading that diagram against the cost model in section VII explains a great deal of agent economics. Turn 40 of a conversation is not a short message. It is one enormous prompt containing all 39 previous turns, every tool result, and all the instructions, sent again from scratch. Cost per turn grows with the square of conversation length if nothing is done, which is why prefix caching moved from an optimisation to a necessity, and why harnesses summarise and prune aggressively.
The last piece connects the harness back to accuracy. Because more tokens can be spent before answering, quality became purchasable at inference time rather than only at training time. A model can be made to generate an extended internal working-out before its visible answer, or to produce several candidate answers and select among them. The benchmark movement has been large, and the trade is explicit: reasoning runs consume far more tokens per query, which is why total inference compute keeps climbing even as the price per token falls. The trade is not always favourable. Generating 8,000 tokens costs sixteen times the tokens of generating 500, published studies of the practice document diminishing and sometimes negative returns, and separate work documented models burning thousands of tokens on trivial arithmetic that a plain model answers correctly and instantly.
This also marks the honest boundary of the report. That extra generation improves accuracy on hard multi-step problems is measured and not seriously disputed. Why it does, whether the intermediate text is genuine intermediate computation or a prompt to itself that happens to land in better regions of the distribution, is not settled. Anyone asserting confidently in either direction is ahead of the evidence.
A mental model is worth what it can predict. These three can be checked in an afternoon against any commercial interface or a model running on a laptop, and each fails loudly if the picture above is wrong.
One. Separate the two latencies. Send a 50-token prompt asking for a long answer, then a 20,000-token prompt asking for an equally long answer. Time the gap before the first word separately from the words-per-second afterwards. The prediction: the first gap grows substantially, the streaming rate barely moves. If the pause scaled but the stream also collapsed, prefill and decode would not be the separate phases described in section VI.
Two. Make the cache miss on purpose. Run a long stable system prompt twice and note the reported cached-token counts and the time to first word. Then insert a current timestamp at the very top and run it twice more. The prediction: the cache hit disappears entirely, not partially, and both cost and initial latency jump. A partial hit would mean the reuse is semantic rather than a byte-exact prefix match, and section VII would be wrong about what is stored.
Three. Find the alphabet. Ask for the number of times a specific letter appears in an unusual long word, then ask the same question with the word written out with spaces between every letter. The prediction: accuracy improves sharply in the second case, because spacing forces the tokenizer to hand the model the letters as separate units. If performance were identical, the failure would be about reasoning rather than about what section II says the model receives.
The bottom line the three tests are checking is a single sentence. Nothing in the system knows anything; a fixed array of numbers scores the next fragment of text, and every other capability, the memory, the tools, the deliberation, the persona, is text that some surrounding program decided to put in front of it. That is not a deflation. A weighted average computed 80 times over is genuinely enough to do most of what these systems do, and the constraints in it are the ones that determine what they cost, where they fail, and which of those failures a better harness can fix and which it cannot.