The Model and the Machine Around It
A language model does one small thing: read some text, write the next word, do it again. Everything else you notice about an AI product comes from the software wrapped around it. This builds that software up one part at a time, so you can see where the model ends and the product begins.

The one thing the model does
Hand the model the text so far. It gives back the next token. A token is a chunk of text: a whole word, part of a word, or a mark of punctuation. Add that token to the text and ask again, and again, until a stop signal arrives. That loop is the entire model.
# the model: text in, one token out next = core_lm_op(text_so_far) text_so_far = text_so_far + next # append it # repeat until the stop signal
The model does not always return the same token for the same input. It weighs the odds across many possible next tokens, and one is picked. That is why two runs of the same prompt can read differently.
At this point the model knows nothing about users, conversations, files, memory, tools, or the internet. It reads tokens and writes one more. Everything that follows is the harness.
Choosing what the model sees
A real product rarely hands the model the raw conversation. Before each call, the harness assembles what the model will read: the standing instructions, earlier turns, saved notes, files, search results, today's date. It can add, cut, reorder, or shorten any of it. This is the first real power of a harness. It controls the model's input.
to_read = context_function(text_so_far, everything_else) next = core_lm_op(to_read) # the model sees only what was assembled
Context is not memory
Storing something is not the same as the model remembering it. When a product seems to recall last week, it kept that text somewhere and placed it back in front of the model this time.
The model itself remembers nothing between calls. Context is what it sees right now. Memory is what sits in storage until the harness decides to show it again.
Letting the output do something
So far information only flows toward the model. Now let it flow back out. The harness watches what the model produces, and when the model asks to read a file, run a command, search the web, or send mail, the harness actually does it and feeds the result back in.
The model still only writes tokens. It never touches the disk. The software around it reads those tokens and makes something happen in the world. A tool is exactly this: one wired-up reaction.
Deciding on another turn
The model stopping is not the same as the task finishing. After a stop, the harness can look at what just happened, decide there is more to do, and call the model again. That repeated cycle, model then action then model then action until the harness is satisfied, is what people mean by an agent loop.
So an agent is not a different kind of model. It is a model, plus some stored state, plus repeated calls, plus the ability to act between them.
Running more than one
Nothing says there is only one conversation. The harness can keep several going at once, each with its own text and its own state, all using the same model. The new power is that it can carry information from one to another. One thread investigates, the harness passes its findings to a second, and the second critiques them.
A subagent is just another separately-kept thread. This is where researcher then writer, or coder then reviewer then coder, or planner then several workers then a synthesiser all come from. The model has not changed. The harness is running several conversations and passing notes between them.
Choosing who runs next
Once there are several models, threads, or paths, something has to pick which one goes next. That picker is routing. It can be plain code (if the task is coding, use the coder), or another model call, or both together.
Routing is a different question from multiplicity. Multiplicity is what threads exist. Routing is which one gets control now.
Where this lands
Start with one operation: text in, one token out. Everything else is machinery around it. Two products can run the exact same model and behave nothing alike, because they differ in these five parts.
The five parts, in one place:
So the honest way to name what you are using is not the model but the model and its harness. Behaviour you credit to the model may be coming from either side of that line.