Ollama
A simple local model service and useful foundation for apps and automation. Models are managed through the runtime.
Good for:beginners, developers, local APIsOfficial documentationFrom the box to your first prompt
A safe setup sequence for Windows and Linux: check hardware requirements, choose a runtime, load a suitable model, verify GPU offload, then optimize.
Before installation
LM Studio recommends at least 16GB of RAM and 4GB of dedicated VRAM on Windows. That is enough only for small models. A new local AI PC should be sized around the model class you intend to run.
| Component | Technical entry point | Practical local LLM target |
|---|---|---|
| System RAM | 16GB for small models and short context | 32GB for 7B–14B; 64GB or more for larger models and offload |
| GPU memory | 4GB is LM Studio's recommendation for starting out | 12–16GB for 7B–14B; 24–32GB for fast 20B–32B inference |
| CPU | LM Studio requires AVX2 on x64; Arm64 is available on supported systems | A modern multi-core CPU, especially when model layers reside in system memory |
| SSD | space for the app and one small model | plan at least 1TB of usable space, preferably 2TB for several model families |
| GPU backend | must match hardware and operating system | NVIDIA CUDA, supported AMD ROCm/Vulkan or an officially listed GB10/Ryzen AI path |
Check current support before buying: Ollama lists supported NVIDIA, AMD and GB10 hardware, while LM Studio documents OS, CPU, RAM and GPU requirements. Use our VRAM guide to size the actual model.
From model to complete system
Ollama does not have one useful RAM or VRAM requirement for every model. Model size, quantization, context and memory used by the operating system and other services set the real target. The current hardware list includes RTX 50-series GPUs, GB10 and Ryzen AI Max+ 395, but operating-system and backend support still need checking.
| Planned model class | Practical memory tier | Suitable complete systems | Most important buying check |
|---|---|---|---|
| 7B to 14B | 12 to 16GB dedicated VRAM; 32GB system RAM | Choose an entry GPU PC | The exact model and context must fit inside fast memory. |
| 20B to 32B, speed first | 24 to 32GB dedicated VRAM; 64GB system RAM | HP OMEN 45L RTX 5090 | Confirm the desktop GPU has 32GB and check chassis cooling and power. |
| 65B to 70B, compact capacity | 128GB unified memory with system headroom | GMKtec EVO-X2, Beelink GTR9 Pro or HP Z2 Mini G1a | Verify the 128GB configuration, usable UMA allocation and Radeon runtime path. |
| 70B with NVIDIA tooling | 128GB coherent unified memory | ASUS Ascent GX10 | Verify Arm64, DGX OS and every required extension before purchase. |
Six-step setup
The interface is personal preference. Drivers, memory headroom and a compatible model format are the technical foundations.
Install operating-system updates plus current graphics and chipset drivers from official sources. On unified-memory systems, also check BIOS/UEFI and available UMA allocation.
Keep useful capacity free beyond the model itself. Several quantizations of a large model can quickly consume hundreds of gigabytes.
Begin with a well-documented 7B or 8B instruct model from your runtime's catalog. Verify drivers, GPU offload and operation before starting very large downloads.
Monitor GPU/UMA memory, system RAM, temperatures, fan behavior and time to first token. If heavy offloading occurs, a smaller or more strongly quantized model is often better.
Initially bind a local API to 127.0.0.1 only. Do not expose it to your LAN or the internet without review. Sensitive documents stay local only if plugins, web search and telemetry are configured accordingly.
Choose your tool
A simple local model service and useful foundation for apps and automation. Models are managed through the runtime.
Good for:beginners, developers, local APIsOfficial documentationGraphical model discovery, chat interface and local server in a desktop application.
Good for:visual setup, comparing modelsOfficial documentationA lightweight open-source backend with extensive controls for offloading, quantization and platforms.
Good for:control, experiments, custom integrationsProject pageOllama hardware, answered
No. Ollama can run models on a CPU, but a supported GPU or suitable unified-memory system can accelerate inference substantially when much of the model resides in fast memory.
Small 7B to 8B models can run with 16GB, but 32GB is a more practical floor for a new PC. Larger models, long context and CPU offloading benefit from 64GB or 128GB. The actual model file plus headroom remains the deciding figure.
Yes, with suitable quantization and enough fast memory. For many Q4-like 70B models, a 64GB tier is a tight lower bound; 96GB to 128GB leaves more room for context, the runtime and the operating system.
NVIDIA provides a widely used CUDA path, while currently supported AMD hardware can accelerate through ROCm or Vulkan. The exact GPU, operating system, driver, Ollama version and available memory matter more than the vendor name alone.
Windows
Native installation is usually enough for Ollama and LM Studio. WSL2 helps when your development stack expects Linux. Avoid parallel services consuming the same GPU memory.
Linux
Linux works well for headless operation and reproducible environments. Check which kernel, GPU driver and backend are officially supported for your hardware.
Private does not automatically mean secure
Verify whether your chosen application uses online features or telemetry.
Bind locally, or protect it with authentication and a firewall.
Model licenses can impose different rules on use and redistribution.
Watch temperatures, dust, power consumption and noise.
Hardware still undecided?
The finder translates your desired model into a realistic system class. If several devices will use it, also read the local AI server hardware guide.