From the box to your first prompt

Run a local LLM with Ollama or LM Studio

A safe setup sequence for Windows and Linux: check hardware requirements, choose a runtime, load a suitable model, verify GPU offload, then optimize.

Before installation

Ollama and LM Studio hardware requirements

LM Studio recommends at least 16GB of RAM and 4GB of dedicated VRAM on Windows. That is enough only for small models. A new local AI PC should be sized around the model class you intend to run.

ComponentTechnical entry pointPractical local LLM target
System RAM16GB for small models and short context32GB for 7B–14B; 64GB or more for larger models and offload
GPU memory4GB is LM Studio's recommendation for starting out12–16GB for 7B–14B; 24–32GB for fast 20B–32B inference
CPULM Studio requires AVX2 on x64; Arm64 is available on supported systemsA modern multi-core CPU, especially when model layers reside in system memory
SSDspace for the app and one small modelplan at least 1TB of usable space, preferably 2TB for several model families
GPU backendmust match hardware and operating systemNVIDIA CUDA, supported AMD ROCm/Vulkan or an officially listed GB10/Ryzen AI path

Check current support before buying: Ollama lists supported NVIDIA, AMD and GB10 hardware, while LM Studio documents OS, CPU, RAM and GPU requirements. Use our VRAM guide to size the actual model.

From model to complete system

What PC is right for Ollama?

Ollama does not have one useful RAM or VRAM requirement for every model. Model size, quantization, context and memory used by the operating system and other services set the real target. The current hardware list includes RTX 50-series GPUs, GB10 and Ryzen AI Max+ 395, but operating-system and backend support still need checking.

Planned model classPractical memory tierSuitable complete systemsMost important buying check
7B to 14B12 to 16GB dedicated VRAM; 32GB system RAMChoose an entry GPU PCThe exact model and context must fit inside fast memory.
20B to 32B, speed first24 to 32GB dedicated VRAM; 64GB system RAMHP OMEN 45L RTX 5090Confirm the desktop GPU has 32GB and check chassis cooling and power.
65B to 70B, compact capacity128GB unified memory with system headroomGMKtec EVO-X2, Beelink GTR9 Pro or HP Z2 Mini G1aVerify the 128GB configuration, usable UMA allocation and Radeon runtime path.
70B with NVIDIA tooling128GB coherent unified memoryASUS Ascent GX10Verify Arm64, DGX OS and every required extension before purchase.

Six-step setup

Six steps to a local LLM

The interface is personal preference. Drivers, memory headroom and a compatible model format are the technical foundations.

  1. 01

    Update firmware and drivers

    Install operating-system updates plus current graphics and chipset drivers from official sources. On unified-memory systems, also check BIOS/UEFI and available UMA allocation.

  2. 02

    Check memory and free SSD space

    Keep useful capacity free beyond the model itself. Several quantizations of a large model can quickly consume hundreds of gigabytes.

  3. 03

    Choose a runtime

    Ollama is convenient for a local service and command-line workflows. LM Studio provides a graphical interface. llama.cpp gives experienced users extensive control.

  4. 04

    Start small

    Begin with a well-documented 7B or 8B instruct model from your runtime's catalog. Verify drivers, GPU offload and operation before starting very large downloads.

  5. 05

    Watch utilization

    Monitor GPU/UMA memory, system RAM, temperatures, fan behavior and time to first token. If heavy offloading occurs, a smaller or more strongly quantized model is often better.

  6. 06

    Secure access

    Initially bind a local API to 127.0.0.1 only. Do not expose it to your LAN or the internet without review. Sensitive documents stay local only if plugins, web search and telemetry are configured accordingly.

Choose your tool

Which runtime fits you?

Ollama

A simple local model service and useful foundation for apps and automation. Models are managed through the runtime.

Good for:beginners, developers, local APIsOfficial documentation

LM Studio

Graphical model discovery, chat interface and local server in a desktop application.

Good for:visual setup, comparing modelsOfficial documentation

llama.cpp

A lightweight open-source backend with extensive controls for offloading, quantization and platforms.

Good for:control, experiments, custom integrationsProject page

Ollama hardware, answered

Common questions before buying a PC

Does Ollama require a GPU?

No. Ollama can run models on a CPU, but a supported GPU or suitable unified-memory system can accelerate inference substantially when much of the model resides in fast memory.

How much RAM does Ollama need?

Small 7B to 8B models can run with 16GB, but 32GB is a more practical floor for a new PC. Larger models, long context and CPU offloading benefit from 64GB or 128GB. The actual model file plus headroom remains the deciding figure.

Can Ollama run a 70B model locally?

Yes, with suitable quantization and enough fast memory. For many Q4-like 70B models, a 64GB tier is a tight lower bound; 96GB to 128GB leaves more room for context, the runtime and the operating system.

Is NVIDIA or AMD better for Ollama?

NVIDIA provides a widely used CUDA path, while currently supported AMD hardware can accelerate through ROCm or Vulkan. The exact GPU, operating system, driver, Ollama version and available memory matter more than the vendor name alone.

Windows

Start native, add WSL when needed

Native installation is usually enough for Ollama and LM Studio. WSL2 helps when your development stack expects Linux. Avoid parallel services consuming the same GPU memory.

  • current GPU driver
  • check performance mode under sustained load
  • configure local-server autostart deliberately

Linux

Flexible for services and development

Linux works well for headless operation and reproducible environments. Check which kernel, GPU driver and backend are officially supported for your hardware.

  • install runtimes only from trusted sources
  • do not bind the API publicly
  • keep models on a sufficiently large partition

Private does not automatically mean secure

Four security rules

Check offline behavior

Verify whether your chosen application uses online features or telemetry.

Limit the API

Bind locally, or protect it with authentication and a firewall.

Read the license

Model licenses can impose different rules on use and redistribution.

Monitor sustained load

Watch temperatures, dust, power consumption and noise.

Hardware still undecided?

Choose the right memory class first

The finder translates your desired model into a realistic system class. If several devices will use it, also read the local AI server hardware guide.

Open the PC Finder