Weights
Q4 uses much less memory than FP16. Metadata and quantization format mean that “4 bit” is not exactly “0.5 byte per parameter.”
Model memory made understandable
The short answer is enough fast memory for model weights, context cache and headroom. The detailed answer maps 8B, 14B, 32B and 70B models to practical GPU and memory tiers without confusing marketing metrics with usable model memory.
The simple formula
Parameter count alone is not a memory figure. What matters is the number of bits per weight and everything that must remain in memory during use.
Q4 uses much less memory than FP16. Metadata and quantization format mean that “4 bit” is not exactly “0.5 byte per parameter.”
It stores information from the current context. Long prompts, concurrent sessions and some architectures make it grow rapidly.
Ollama, llama.cpp, LM Studio and other backends need working buffers. Implementation and GPU offload affect the total.
A pool that only just fits on paper is a poor buying plan. The OS, display and companion services need breathing room.
Memory needs with headroom
Ranges roughly assume Q4-like quantization plus ordinary runtime headroom. Long context, MoE architectures, vision encoders or concurrent users can require more.
| Model class | Approx. weights | Sensible fast-memory class | Typical use |
|---|---|---|---|
| 7B–8B | about 5–6 GB | 12–16 GB | chat, summarization, simple local helpers |
| 12B–14B | about 8–10 GB | 16 GB | better generalists, entry-level coding |
| 20B–32B | about 14–22 GB | 24–32 GB | coding, RAG, capable assistants |
| 65B–70B | about 40–45 GB | 64 GB or more | higher model quality, complex tasks |
| 100B–120B | about 65–80 GB | 96–128 GB or more | professional capacity class |
“Fits in memory” does not mean “runs fast.” Bandwidth, compute, backend and offload distribution determine speed.
Choose GPU by model
The best GPU for an LLM is not automatically the most expensive model. It is the hardware that keeps model and context in fast memory and is reliably supported by your runtime.
A good fit for many 7B to 14B models in suitable quantization. Headroom becomes tight for 20B+, long context, vision or concurrent sessions.
Typical use: chat and first coding helpersA strong tier for fast 20B to 32B inference. An RTX 5090 provides 32GB of dedicated VRAM; 70B usually needs system-memory offload.
Typical use: coding, RAG and agentsInspect RTX 5090 PCsUnified-memory mini PCs and professional GPUs make room for quantized 70B models. Bandwidth, reserved system memory and backend determine speed.
Typical use: large models and long contextCompare AI mini PCsCheck more than the GPU name: desktop and mobile versions can carry different memory capacities. With unified memory, the usable UMA or VGM allocation also matters.
Two memory paths
Memory directly attached to a separate GPU. High bandwidth makes it the fastest route while the entire model fits.
CPU and integrated GPU share a large pool. Bigger models can fit, but the OS and applications use that same memory.
Common questions
It can be enough for small 7B to 8B models with short to medium context in a memory-efficient quantization. Headroom is limited, though; larger models, vision input and long context can trigger offloading quickly.
12GB is a workable entry point for many 7B to 8B models and selected 12B to 14B models in suitable quantization. Check the actual model file together with KV cache and runtime headroom before buying.
Yes, for many 7B to 14B models in suitable quantization. Long context, large vision models or 32B models can exceed it.
The physical VRAM on a discrete graphics card is normally fixed. System-RAM offloading may still launch an oversized model, but it is usually much slower. A suitable unified-memory system can instead assign a larger share of its common pool to the GPU.
No. Dedicated VRAM sits next to a discrete GPU and offers high bandwidth. System RAM is the PC's main memory. Unified memory is a third design in which the CPU and integrated GPU use the same pool.
Often yes by offloading part of it to system RAM or using stronger quantization. It will not reside entirely in VRAM, and performance can drop substantially.
No. Only suitable unified-memory architectures can share a large portion. The OS and runtime need headroom, and the maximum configurable UMA allocation must be checked.
For large local LLMs, an NPU alone is rarely the main buying criterion. Runtime support, memory capacity, bandwidth and GPU performance matter more.
From memory needs to a system
The PC Finder applies this memory logic and explains its recommendation.