Model memory made understandable

Which GPU and how much VRAM does a local LLM need?

The short answer is enough fast memory for model weights, context cache and headroom. The detailed answer maps 8B, 14B, 32B and 70B models to practical GPU and memory tiers without confusing marketing metrics with usable model memory.

The simple formula

Weights + KV cache + runtime + headroom

Parameter count alone is not a memory figure. What matters is the number of bits per weight and everything that must remain in memory during use.

Model weightsparameters × quantizationKV cachecontext × architectureRuntimebuffers and backendHeadroomOS and other tools

Weights

Q4 uses much less memory than FP16. Metadata and quantization format mean that “4 bit” is not exactly “0.5 byte per parameter.”

KV cache

It stores information from the current context. Long prompts, concurrent sessions and some architectures make it grow rapidly.

Runtime

Ollama, llama.cpp, LM Studio and other backends need working buffers. Implementation and GPU offload affect the total.

Headroom

A pool that only just fits on paper is a poor buying plan. The OS, display and companion services need breathing room.

Memory needs with headroom

Typical needs for quantized text models

Ranges roughly assume Q4-like quantization plus ordinary runtime headroom. Long context, MoE architectures, vision encoders or concurrent users can require more.

Model classApprox. weightsSensible fast-memory classTypical use
7B–8Babout 5–6 GB12–16 GBchat, summarization, simple local helpers
12B–14Babout 8–10 GB16 GBbetter generalists, entry-level coding
20B–32Babout 14–22 GB24–32 GBcoding, RAG, capable assistants
65B–70Babout 40–45 GB64 GB or morehigher model quality, complex tasks
100B–120Babout 65–80 GB96–128 GB or moreprofessional capacity class

“Fits in memory” does not mean “runs fast.” Bandwidth, compute, backend and offload distribution determine speed.

Choose GPU by model

What is the best GPU for a local LLM?

The best GPU for an LLM is not automatically the most expensive model. It is the hardware that keeps model and context in fast memory and is reliably supported by your runtime.

12–16GB: practical entry

A good fit for many 7B to 14B models in suitable quantization. Headroom becomes tight for 20B+, long context, vision or concurrent sessions.

Typical use: chat and first coding helpers

24–32GB: speed sweet spot

A strong tier for fast 20B to 32B inference. An RTX 5090 provides 32GB of dedicated VRAM; 70B usually needs system-memory offload.

Typical use: coding, RAG and agentsInspect RTX 5090 PCs

64–128GB: capacity tier

Unified-memory mini PCs and professional GPUs make room for quantized 70B models. Bandwidth, reserved system memory and backend determine speed.

Typical use: large models and long contextCompare AI mini PCs

Check more than the GPU name: desktop and mobile versions can carry different memory capacities. With unified memory, the usable UMA or VGM allocation also matters.

Two memory paths

VRAM and unified memory are not the same

Dedicated VRAM

Memory directly attached to a separate GPU. High bandwidth makes it the fastest route while the entire model fits.

  • excellent inference speed
  • clear, fixed capacity limit
  • offloading costs performance

Unified memory

CPU and integrated GPU share a large pool. Bigger models can fit, but the OS and applications use that same memory.

  • 64 to 128 GB in compact systems
  • configurable UMA allocation matters
  • less bandwidth than high-end VRAM

Common questions

What buyers often confuse

Is 8GB of VRAM enough for local AI?

It can be enough for small 7B to 8B models with short to medium context in a memory-efficient quantization. Headroom is limited, though; larger models, vision input and long context can trigger offloading quickly.

Is 12GB of VRAM enough for a local LLM?

12GB is a workable entry point for many 7B to 8B models and selected 12B to 14B models in suitable quantization. Check the actual model file together with KV cache and runtime headroom before buying.

Is 16 GB of VRAM enough for local AI?

Yes, for many 7B to 14B models in suitable quantization. Long context, large vision models or 32B models can exceed it.

Can you increase the VRAM on a graphics card?

The physical VRAM on a discrete graphics card is normally fixed. System-RAM offloading may still launch an oversized model, but it is usually much slower. A suitable unified-memory system can instead assign a larger share of its common pool to the GPU.

Is VRAM the same as regular RAM?

No. Dedicated VRAM sits next to a discrete GPU and offers high bandwidth. System RAM is the PC's main memory. Unified memory is a third design in which the CPU and integrated GPU use the same pool.

Can I launch a 70B model with 32 GB of VRAM?

Often yes by offloading part of it to system RAM or using stronger quantization. It will not reside entirely in VRAM, and performance can drop substantially.

Does 128 GB RAM mean 128 GB for the GPU?

No. Only suitable unified-memory architectures can share a large portion. The OS and runtime need headroom, and the maximum configurable UMA allocation must be checked.

Do I need an NPU?

For large local LLMs, an NPU alone is rarely the main buying criterion. Runtime support, memory capacity, bandwidth and GPU performance matter more.

From memory needs to a system

Calculate the right memory class from model size and context

The PC Finder applies this memory logic and explains its recommendation.

Start the PC Finder