Inside the AI Kitchen: How LLMs Actually Run on Your Machine

Your GPU performs trillions of calculations per second — yet when you run an AI model at home, it writes... one... word... at... a... time. That gap between raw power and felt speed is the whole story of this article. To tell it, we'll use one metaphor: the AI restaurant.
| Character | Real thing | Role in the kitchen |
|---|---|---|
| The Recipe Book | The model file (weights) | Everything the kitchen knows, frozen on paper |
| The Manager | CPU | Smart, versatile, works step-by-step — but can't do the heavy cooking |
| The Chef | GPU | Thousands of short-order cooks doing simple math simultaneously |
| Counter & Fridge | RAM / VRAM | Fast workspace; the Chef can't cook from the grocery store (hard drive) |
| The Expeditor | Inference Engine | Runs the pass: decides what happens, when |
The promise: by the end, you'll know why prompts start fast but answers trickle, why chats slow down as they grow, and why a small GPU can run a surprisingly large model.
1. The Ingredients: What a Model Actually Is
A model is not a brain. It's a giant static list of numbers — billions of weights, like a warehouse-sized amplifier with billions of knob positions. A static file can do nothing on its own: no kitchen, no service.
Two facts to carry forward:
Tokens, not words. The book stores word-pieces. "The quick brown wolf jumps" is five words but perhaps six or seven tokens. (We'll treat them as interchangeable below.)
Editions matter. The full-precision book is enormous; quantized editions (8-bit, 4-bit, "Q4") record the knob positions less precisely at a quarter of the size — with nearly identical output. Remember this. Book size is destiny.
2. The Kitchen Crew
| Manager (CPU) | Chef (GPU) | |
|---|---|---|
| Strength | Any job, good judgment | One simple calculation, done thousands of times at once |
| Weakness | Sequential; no brute force | No judgment; can't manage anything |
| AI fit | Orchestration, housekeeping | 99% of an LLM's math is exactly its specialty |
The workspace rule governs everything: the Chef cannot cook from the grocery store. The model file sits on your hard drive (a warehouse across town); it must be hauled into fast memory — RAM, or the GPU's private counter, VRAM — before service. That "loading model..." bar is the haul. It happens once.
The reveal: the bottleneck in AI is rarely the Chef's speed. It's the walk to the fridge — memory bandwidth, how fast pages move between storage and the Chef's hands. Nearly every performance mystery in this article is that walk.
3. Phase One — Prefill: Reading the Ticket
You hit Enter on:
"The quick brown wolf jumps"
The Chef reads the entire ticket in one pass — all tokens simultaneously. This is the GPU's native gift: parallel processing. Long prompt? Slightly longer pause before the first word. Prefill scales with prompt size; it's a reading problem, and the Chef reads fast.
While reading, the Chef takes shorthand notes on a scratchpad — not sentences, but symbols recording what each word is and relates to. This scratchpad is the KV Cache: the AI's short-term memory (full treatment in Section 5).
At the end of the pass, the Chef predicts the most likely next word and writes the first word of the answer:
"over"
4. Phase Two — Decode: Plating Course by Course
Now the kitchen enters a loop, because each word depends on all the words before it:
| Loop | Consults | Writes |
|---|---|---|
| 1 | Ticket + scratchpad | "the" |
| 2 | + "the" added to scratchpad | "lazy" |
| 3 | + "lazy" added to scratchpad | "dog" |
| ... | ... | ... until the dish is complete |
Why so slow? Every loop requires consulting the entire recipe book — billions of knob-positions — and the book lives in the fridge. Writing one word means hauling the whole book to the counter; the next word, again. The Chef could flip pages in microseconds; the walk is the bottleneck. The GPU isn't idle during decode — it's starved, on a job that barely uses its parallel power.
Prefill is a reading problem (compute-bound). Decode is a hauling problem (bandwidth-bound).
And the condensed editions snap into focus: a 4-bit book is a quarter of the walking per word. That's how an 8GB GPU runs a model whose original wouldn't fit on a shelf. (Aside: the Chef actually computes probabilities for every possible next word and draws one; "temperature" dials in how adventurous the draw is. Everything else works the same.)
5. The Scratchpad: KV Cache — and Why Chats Crash
Without the scratchpad, every new word would require re-reading the entire conversation from the beginning. The KV Cache is what the Chef has already noted down — one entry per token — so each loop only consults the notes rather than reconstructing them.
The catch, in hardware terms:
| Property | Consequence |
|---|---|
| Scratchpad lives on the counter (VRAM) — consulted every loop | It competes for the same fast memory as the recipe book |
| It grows — one entry per token, prompt and answer | Long conversations eat counter space |
| Every book has a max scratchpad size — the context window | Past it, the kitchen literally can't take more notes |
Three familiar failure modes fall straight out:
Chats slow down — each loop wades through a longer scratchpad.
Chats crash — scratchpad + book overflow VRAM. That's "out of memory."
AI "forgets" your early messages — you hit the context ceiling. It's a design limit, not a bug.
"It worked fine for twenty messages, then died" is a memory story.
6. The Expeditor: What the Inference Engine Actually Does
Hardware is the staff; the engine is the discipline of running the pass. (It's a role the Manager steps into — software on your CPU directing the GPU — not a sixth employee.) Its job, by phase:
| Job | What it means |
|---|---|
| Translates the book | Compiles the generic model file into your hardware's dialect (CUDA, Metal...), choosing the exact math routines (kernels) for every step |
| Stages the kitchen | Decides what gets hauled to the counter, when, and how much space to reserve — memory doesn't organize itself |
| Runs the pass | Tokenizes your prompt, launches Prefill, drives the Decode loop, watches for the "dish complete" token, serves the result |
| Seats multiple tables | Batches several requests so the Chef cooks in parallel — boosts throughput (customers/hour), not any single meal's speed |
Notice what it never does: cook. Every dish is still the Chef's, from the same book, on the same counter. The engine determines whether the Chef works in rhythm or in a traffic jam.
Hardware sets what the kitchen can do. The inference engine determines what it actually does.
7. Closing Time
A model is a recipe book — inert until service. The CPU manages, the GPU cooks in parallel, and the walk to the fridge sets the pace of everything. Work splits into Prefill (fast parallel read) and Decode (slow sequential loop hauling the whole book per word). The KV Cache spares the kitchen from re-reading everything — at the cost of your counter space. Above it all, the inference engine runs the pass.
The diagnostic payoff:
| Symptom | What's happening | What helps |
|---|---|---|
| Pause before the first word | Prefill still reading a long ticket | Nothing's broken; shorter prompt |
| Slow word-by-word output | Decode hauling the book per word | A smaller edition (quantized model) |
| Crash mid-conversation | Scratchpad + book overflowed VRAM | Shorter history, smaller book |
| Long "loading model..." | The haul from the warehouse | One-time cost; smaller books load faster |
| AI ignores early messages | Context window ceiling | Trim the conversation |
You don't need a supercomputer — you need to know that the speed limit is the walk to the fridge, and that the expeditor, the scratchpad, and the edition you choose decide how smoothly the kitchen runs.
Now you do.





