Local AI. Your hardware. Your data.

Find the right PC to run your own local LLM

Work out how much model memory you really need, then compare complete systems for your model size, speed, space and budget.

  • no cloud required
  • memory-first comparisons
  • transparent evaluation

Three steps to the right class

Not every “AI PC” is a good LLM PC

For local language models, the crucial question is whether model weights, context cache and runtime headroom fit into fast memory together. An impressive NPU TOPS figure alone does not answer it.

01

Choose model size

A quantized 8B model is modest. 32B and 70B models need considerably more memory but can cover more demanding tasks.

02

Speed or capacity?

A discrete high-end GPU delivers excellent speed. Large unified memory can hold bigger models, with a different performance profile.

03

Check the exact system

Cooling, memory allocation, storage, upgradeability and the exact GPU configuration matter as much as the processor name.

Memory tiers at a glance

Which memory tier fits?

Understand model memory
16 GB

Compact models

A practical starting point for 7B to 14B models in suitable quantization and normal context lengths.

Chat · summarization · first coding workflows
64–128 GB

Large models

The capacity class for 70B models and long contexts, depending on runtime and quantization.

70B · long context · agent stacks

These ranges include practical headroom but are not a guarantee. Architecture, quantization, context, KV cache and runtime all change memory use.

Build a private coding assistant

Choose hardware by autocomplete, code chat or repository agent

A coding agent needs more than model weights: tool calls, 64K context and parallel sessions can change the memory target. Our new guide connects the workflow to a suitable complete PC.

Set up OpenCode with Ollama
Open the local coding guide

Specific complete systems

Choose by model size and software needs: two Ryzen AI Max+ PCs with different networking, an NVIDIA AI appliance, or a high-speed GPU tower.

All systems and criteria
GMKtec EVO-X2 128 GB

Capacity focus · compact

GMKtec EVO-X2

Ryzen AI Max+ 395 with 128 GB of unified memory and a 2 TB configuration: compelling when fitting a large quantized model matters more than maximum discrete-GPU speed. Read the model guide.

  • 128 GBunified memory
  • 2 TBSSD configuration
  • Compactform factor
Check configurations*
Beelink GTR9 Pro 128 GB / 2 TB

Capacity and networking · compact

Beelink GTR9 Pro

A 128 GB Ryzen AI Max+ 395 system with a 2 TB SSD and dual 10GbE for a fast home-network model server. Read the model guide.

  • 128 GBunified memory
  • 2 TBSSD configuration
  • 2× 10GbEwired networking
Check configurations*
HP OMEN 45L RTX 5090

Speed focus · desktop

HP OMEN 45L GT22-3090

A roomy tower configuration combining a GeForce RTX 5090 32 GB with 64 GB of system memory. Suited to fast local inference where the active workload fits in GPU memory.

  • 32 GBdiscrete VRAM
  • 64 GBsystem memory
  • Desktopcooling
Check configurations*
ASUS Ascent GX10 128 GB / 4 TB

NVIDIA stack · AI appliance

ASUS Ascent GX10

A GB10 Grace Blackwell system with 128 GB of coherent unified memory and an Ubuntu-based NVIDIA environment. It suits CUDA-native local AI work, but is not a general Windows desktop. Read the buying guide.

  • 128 GBcoherent unified memory
  • 4 TBSSD configuration
  • Arm/Linuxdevelopment platform
Check configurations*

* Paid link. We may earn a commission if you buy; your price is unchanged. As an Amazon Associate we earn from qualifying purchases. Please verify the exact configuration and current availability with the retailer.

Your workload matters

From model to suitable PC in about a minute

The finder considers model class, context length, intended use and your main priority. It returns an explainable hardware class, not false precision.

Calculate my recommendation

Choose by your main goal

Which kind of system actually helps?

Fast coding and RAG

For mostly 8B to 32B models, an RTX 5090 with 32 GB of VRAM is the speed-focused choice. Pair it with at least 64 GB of system RAM and 2 TB of NVMe storage.

RTX 5090 buying check

Run 70B at home

128 GB of unified memory can hold many quantized 70B models with useful reserve, though response speed remains below large professional GPUs.

Compare 70B hardware options

Plan a large model library

A 2 TB SSD is a useful starting point, but several model families, quantizations and project data can fill it quickly. Check for another M.2 slot or fast external storage before buying.

Compare storage and upgrades