Skip to main content

Command Palette

Search for a command to run...

How a Laptop Learned to Run a Giant

Updated
•5 min read•View as Markdown
H
Father of two, tech lover. Building systems by day, raising curious minds by night.

New to the jargon? Start with the glossary →


A 70-billion-parameter model is roughly 140 GB of data. A mid-range gaming laptop has between 8 and 16 GB of video memory.

That's like trying to pour a swimming pool into a teacup. It shouldn't fit. So how is it running?

The answer is that we stopped trying to force the whole model into memory at once. Instead, we shrink it, we share the load, and we only use the parts we actually need. Here are the three tricks that make it possible.

Trick 1: Shrinking the Model

The first problem is pure size. You can't fit 140 GB into an 8 GB space. So we make the model smaller.

Quantization is the biggest lever. Think of it like a high-resolution photo. The original model uses 16-bit precision (called FP16), where every single parameter takes up 2 bytes of memory. That's the "high-res" version. Quantization drops that down to 8-bit (INT8), then 4-bit (INT4). At INT4, every parameter takes just 0.5 bytes.

That 140 GB model suddenly drops to around 35 GB. You lose a little bit of quality—the photo gets slightly blurrier—but you save an enormous amount of memory.

Pruning is the second lever. Researchers look at the model's billions of parameters and find the ones that aren't pulling their weight. They zero them out, effectively deleting them. It's like trimming dead branches off a tree so the rest can breathe.

Knowledge distillation is the third. Instead of compressing a big model, you train a small model to copy a big one. The small model learns from the big model's answers and reasoning, like an apprentice learning from a master. You end up with a model a fraction of the size, carrying a surprising amount of the original's capability.

Trick 2: Sharing the Load

Shrinking helps, but a 35 GB compressed model still doesn't fit comfortably into 16 GB of video memory. So we use everything the computer has.

Unified memory changed the game on Apple Silicon. In a traditional PC, the CPU has its own RAM, and the GPU has its own VRAM. Moving data between them is slow. Apple's M-series chips use a single pool of memory shared by both the CPU and the GPU. A MacBook Pro with 128 GB of unified memory can hold a ~65B model comfortably—no copying required.

CPU offloading is the PC equivalent. Instead of demanding that the entire model live on the graphics card, tools like llama.cpp split the model up. The "hot" layers that are needed right now stay on the GPU. The rest sit in system RAM, and the CPU handles them.

Is it slower than pure GPU inference? Yes. But it's a thousand times better than not running at all.

Trick 3: Doing Less Work

The final trick is architectural. Modern models are simply smarter about how they use themselves.

Mixture of Experts (MoE) is the biggest shift. Instead of a monolithic model that uses all 70 billion parameters for every single question, an MoE model is made of many smaller specialist teams. When a question comes in, a router sends it to just a few of those experts—maybe 8 to 10 billion parameters' worth. The model is huge on paper, but computationally cheap in practice. Examples like Mixtral, Qwen2.5-MoE, and Grok use this to deliver massive capability without massive compute costs.

Sparse attention reduces the memory and compute required for long contexts. Instead of forcing the model to look at every single word it has ever seen, sparse attention focuses only on the relevant parts. It cuts the math down drastically.

And then there's just better design. Models like Llama 3 and Qwen 2.5 are simply more parameter-efficient than their predecessors. A smaller model today beats a bigger model from two years ago.

The Reality Check

So what does this actually look like on your desk?

If you have a Mac Mini (M-series) with 32 GB of RAM, you can run Qwen2.5-14B or Llama-3-8B comfortably at INT4. You might even squeeze in a 32B model if you're aggressive with quantization and don't have much else open.

If you have a MacBook Pro with 64 GB or 128 GB of unified memory, the world opens up. You can run a 70B model at INT4. It's slow—roughly 5-10 tokens per second—but it runs, completely offline, no API calls needed. We're talking batch processing and careful prompting, not instant chat.

There is a tradeoff. Larger models mean slower inference, more heat, and louder fans. But here's the part that surprises most people.

Training an AI model requires massive, sustained compute. It's like running a marathon at sprint speed for weeks. But inference—just using the model to get an answer—doesn't need that same intensity. You're reading a huge matrix multiplication table, not calculating it from scratch.

That's why a laptop can do this. It's not performing miracles. It's just reading the answer, one compressed piece at a time.


Next in the series: The New Shape of AI Models →

New to the jargon? Every term in this series has a plain-English explanation in the glossary.

A

The 35 GB INT4 example is a useful starting point, but the usable context length deserves its own memory budget. The compressed weights can fit while a longer conversation still runs out of memory because the attention cache and runtime buffers also occupy the same pool. That makes 'the model loads' a weaker test than 'the intended conversation runs.'

For the laptop reality check, I would record memory and response speed at several prompt lengths with the same offload settings. Leaving headroom for the operating system also matters on unified-memory machines. It would make the hardware examples easier to translate into an actual local workload.

Big Models, Small Machines

Part 3 of 3

*How local hardware learned to run giant AI — and why size stopped measuring intelligence.* A five-part series on why parameter count stopped measuring intelligence, and how ordinary machines started running extraordinary models. --- ## 📚 Table of Contents | # | Essay | Status | |---|---|---| | 0 | [Plain English Glossary](https://nightthoughts.hashnode.dev/plain-english-glossary-ai-terms-for-the-rest-of-us) | ✅ Published | | 1 | [Bigger Isn't Always Smarter](https://nightthoughts.hashnode.dev/bigger-isn-t-always-smarter) | ✅ Published | | 2 | [How a Laptop Learned to Run a Giant](https://nightthoughts.hashnode.dev/how-a-laptop-learned-to-run-a-giant) | ✅ Published | | 3 | [The New Shape of AI Models](#) | 🚧 Coming soon | | 4 | [Who Actually Saves Money?](#) | 🚧 Coming soon | | 5 | [The Future Is Both](#) | 🚧 Coming soon |

Start from the beginning

Plain English Glossary: AI Terms for the Rest of Us

The Basics Term What it actually means Parameter A dial the model tuned while learning. Billions of them. Think of them as the model's settings. Model size (e.g., 70B) How many dials the model