AI Frontiers, part 17: Quantization — running 70B on a laptop
Part 17from the AI Frontiers series · 65 parts in all
The claim in the title deserves to be audited before it is celebrated. A 70-billion parameter model at 16-bit precision needs about 140 gigabytes of memory just to hold its weights, plus the key-value cache and activations, which is why serving one has historically required a rack of accelerators. At 4-bit, the weights need roughly 35 to 40 gigabytes. Whether that fits on your desk depends entirely on which machine you have: a laptop with 64 gigabytes of unified memory runs it comfortably, and a laptop with 16 gigabytes does not run it at all. The genuinely surprising part of the last two years is not that quantization makes big models smaller — that has been standard practice since the first CNN on a phone — but that it now costs almost nothing in quality, and that the enabling research was published freely enough for the tooling to converge on a small set of formats.
What quantization actually changes
A neural network stores numbers. Most of those numbers are the weights of matrix multiplications, and they have been stored in 16- or 32-bit floating point since forever because that is what training produces. Quantization replaces them with lower-precision representations: 8 bits, 4 bits, and in the research literature lower still. The naive version of this is trivial to describe and disastrous in practice. Take a weight distribution, find its minimum and maximum, map the interval linearly onto the available integer range, and store a scale factor. Do that per tensor and a handful of outliers stretch the range so badly that the majority of weights collapse into a fraction of the available buckets.
Every serious method of the last several years is an answer to that outlier problem. LLM.int8() (Dettmers et al., "LLM.int8()") noticed that a small number of activation dimensions carry enormous magnitudes and split them out into a separate high-precision computation, quantizing everything else. SmoothQuant (Xiao et al.) migrated the difficulty from activations to weights by scaling channels, because weights quantize more gracefully than activations do. GPTQ (Frantar et al.) approached it as an optimization problem: quantize the weights column by column, and after each column, adjust the remaining unquantized weights to compensate for the error already introduced, using second-order information from a calibration set. AWQ (Lin et al.) observed that protecting a small fraction of important weight channels — identified by activation magnitude, not weight magnitude — preserves quality better than protecting channels chosen by weight statistics, and turned that into a scaling search.
The pattern across all of them is worth naming because it explains why quantization went from a 1–2% quality loss at 8 bits to near-lossless at 4: the field stopped quantizing uniformly and started quantizing selectively, spending bits where they change the output and economizing where they do not. Modern schemes mix precisions freely — group sizes of 32 or 64 weights sharing a scale, some layers in 8 bits with others in 4, embedding and output layers handled separately, and outliers kept at higher precision. That is also why there is no single "4-bit" number to compare; the actual bit budget depends on how many groups need extra bits, and the honest unit is bytes of file size.
The formats that won, and why
Two of the most consequential things to happen in this area were not papers. The first was llama.cpp adopting a family of block-quantization formats — the Q4_K_M and Q5_K_M names that appear in every model download today — which blend different bit widths within and between blocks and use a small amount of extra precision for the most sensitive tensors (Gerganov et al.). The second was the emergence of a shared file format, GGUF, that bundles weights, tokenizer, template and metadata into one artifact a dozen runtimes can load. Together they produced the thing that actually determined adoption: a model file you could download, verify against a checksum, and run in Ollama, llama.cpp, LM Studio or a Python binding without conversion.
On the GPU-serving side, vLLM and TensorRT-LLM support the GPTQ and AWQ formats (Kwon et al.), and they matter because a quantized model is not only smaller but faster at inference. That is the point people miss. Decoding a token is memory-bandwidth bound — the accelerator spends its time reading weights, not multiplying them — so halving the bytes per weight roughly halves the time per token for the same batch size. Quantization buys you memory and throughput, and the only thing it costs is some quality and a calibration step.
The research frontier has moved further. QuIP# and AQLM push toward 2-bit with lattice-based and codebook approaches; BitNet b1.58 (Ma et al.) asks a different question entirely, training a model from scratch whose weights are constrained to ternary values −1, 0 and +1, and reporting competitive quality at a fraction of the energy. That last result matters because it relocates the problem: if you train for low precision instead of post-processing into it, you can go much lower than post-training quantization allows. The practical catch, as always, is that it requires training the model, so it is a technique for people building models rather than people running them.
The arithmetic you should do before believing a demo
Memory, not parameters, is the constraint, and the accounting is worth doing explicitly. Take a 70B model at four bits per weight with a group size of 32 and an eight-bit scale per group: that is 4.25 bits per weight on average, or roughly 37 gigabytes for the weights. Now add the KV cache, which scales with context length, number of layers, number of KV heads and precision. For a 70B model with grouped-query attention and a few thousand tokens of context, that is a couple of gigabytes at 16-bit and half that at 8-bit. Add the runtime, the tokenizer, the framework's own overhead and the operating system, and you land around 40–44 gigabytes for a comfortable session.
It is also worth remembering that memory bandwidth, not capacity, decides whether the result is usable. A 4-bit 70B model that fits in 40 gigabytes but streams those weights from a slow pool of shared memory will produce tokens at a rate that makes the exercise pointless. The laptop grade is a memory-bandwidth grade as much as a memory-size one.
Which is why the answer to "can I run this on a laptop?" is almost always a question about your laptop's unified memory, and why the Macs with 64 or 128 gigabytes became the default demonstration platform for local inference. It is also why mixture-of-experts models look so attractive here: if only a fraction of the parameters are active per token, you need the full weight set in memory but you do not need the compute of a dense model of that size — which makes a 4-bit MoE a much better fit for a single machine than a 4-bit dense model of equivalent quality.
Two more honest caveats. Quantization shifts quality unevenly across tasks; perplexity moves very little while long-form coherence, exact copying or arithmetic can degrade noticeably, so a benchmark number is a poor guide and your own evaluation set is a good one. And quantized weights interact badly with some fine-tuning workflows, which is the reason the QLoRA result from part 3 was so important: it showed you could attach trainable low-rank adapters to a frozen 4-bit backbone and recover the quality of full fine-tuning on a single GPU. The two technologies — quantized inference and quantized training — compound.
Choosing a setting, without guesswork
The practical procedure is short and mostly consists of not trusting a leaderboard.
Start from the memory budget, not the model. Measure what your machine has available after the operating system and the runtime — on many laptops that is several gigabytes less than the marketing number — then divide by bytes per weight to get the largest model that fits, leaving a few gigabytes for the KV cache and activations. Pick the quantization level from that budget rather than the other way around. A model that does not fit is not a model.
Quantize for the task, not for the file size. Perplexity is nearly unmoved by 4-bit quantization, which is why 4-bit is safe in general, but the tasks people care about are not general. Extraction, classification and summarization tolerate aggressive quantization well; arithmetic, exact quoting and long-range coherence degrade earlier. Build an evaluation set of fifty representative prompts with expected outcomes, run it at each level, and pick the smallest one that passes. This takes an afternoon and is the only calibration step that matters.
Do not forget the KV cache. Weight quantization gets all the attention, but the cache grows with conversation length and can dominate at long contexts. Quantizing it to 8 bits is usually free in quality terms and halves that memory, and 2-bit cache schemes (Liu et al.) are viable with some care about which layers get protected. A configuration that works at 4K context can fall over at 32K purely because of the cache.
Consider the mixture-of-experts shortcut. If a 4-bit dense model does not fit, a 4-bit mixture-of-experts model with a similar quality profile often will run at an acceptable speed on the same machine, because active parameters drive compute while total parameters drive memory — and memory is usually the binding constraint.
Training in low precision, which is the other half
Everything above is post-training quantization: the model was trained at high precision and squeezed afterwards. The alternative is to train in low precision from the start, and the research keeps suggesting that the ceiling is much lower than post-training methods allow. BitNet b1.58 is the sharpest example — weights restricted to three values, trained that way, reporting quality competitive with comparable full-precision models (Ma et al.). Mixed-precision training standards have also shifted: FP8 training is now routine at the frontier, which is part of how the training-cost story of part 13 became possible in the first place.
The split between the two approaches mirrors a split that runs through the whole field. If you are consuming a model, post-training quantization is the only option and it is good enough. If you are producing one, training-aware low precision is where the real energy savings live, because you can make the architecture itself cheap rather than paying to fix it afterwards. The reason both exist is that the majority of practitioners are in the first category, and the tooling reflects that: one command-line flag versus one training run.
What this does to the shape of the industry
The strategic consequence is not that models got smaller. It is that a capable model can now exist in places a network call cannot reach: on a phone, on a factory floor, inside a hospital, in a car, on an air-gapped machine. That is a genuine change in what is buildable, and it directly contradicts the assumption embedded in most 2023-era architectures that inference is an API call with a per-token price.
It also changes the economics in a way that only becomes visible at scale. If a task requires a hundred million tokens a month, the difference between an API price and a fixed capital cost is the difference between an operating expense that grows with success and one that does not. Quantization is the technical precondition for the second option, and it is why every serious platform team I know has one person quietly running local models against a sample of their traffic to see where the crossover point is.
What quantization does not do is make a small model smart. Squeezing a 7B model onto a phone does not give you a 70B model's judgment; it gives you a 7B model of judgment at a much better cost. Which is why the systems that work in 2025 are hybrids — a small local model for the high-volume, low-difficulty majority of requests, a frontier model for the minority that need it, and a router in front that decides which is which. The whole stack has quietly converged on that architecture, and every part of it exists because someone solved a memory problem.
Works Cited
Dettmers, Tim, et al. "8-bit Optimizers via Block-wise Quantization." arXiv, 2021, arxiv.org/abs/2110.02861. Accessed 18 Sept. 2025.
---. "The Case for 4-bit Precision: k-bit Inference Scaling Laws." arXiv, 2022, arxiv.org/abs/2212.09720. Accessed 18 Sept. 2025.
---. "LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale." arXiv, 2022, arxiv.org/abs/2208.07339. Accessed 18 Sept. 2025.
Dettmers, Tim, et al. "QLoRA: Efficient Finetuning of Quantized LLMs." arXiv, 2023, arxiv.org/abs/2305.14314. Accessed 18 Sept. 2025.
Egiazarian, Vage, et al. "AQLM: Extreme Compression of Large Language Models via Additive Quantization." arXiv, 2024, arxiv.org/abs/2401.06118. Accessed 18 Sept. 2025.
Frantar, Elias, et al. "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers." arXiv, 2022, arxiv.org/abs/2210.17323. Accessed 18 Sept. 2025.
Gerganov, Georgi, et al. "llama.cpp: LLM Inference in C/C++." GitHub, github.com/ggerganov/llama.cpp. Accessed 18 Sept. 2025.
Liu, Zechun, et al. "KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache." arXiv, 2024, arxiv.org/abs/2402.02750. Accessed 18 Sept. 2025.
Lin, Ji, et al. "AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration." arXiv, 2023, arxiv.org/abs/2306.00978. Accessed 18 Sept. 2025.
Kwon, Woosuk, et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." arXiv, 2023, arxiv.org/abs/2309.06180. Accessed 18 Sept. 2025.
Ma, Shuming, et al. "The Era of 1-bit LLMs: All Large Language Models Are in 1.58 Bits." arXiv, 2024, arxiv.org/abs/2402.17764. Accessed 18 Sept. 2025.
Tseng, Albert, et al. "QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks." arXiv, 2024, arxiv.org/abs/2402.04396. Accessed 18 Sept. 2025.
Xiao, Guangxuan, et al. "SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models." arXiv, 2022, arxiv.org/abs/2211.10438. Accessed 18 Sept. 2025.