Autocomplete
Short suggestions while you type. Low latency matters more than a very large model or repository-wide context.
Plan: 8–16 GB fast memoryLocal coding · private repository · your hardware
Start with the job: fast autocomplete needs different hardware from a repository agent that reads files, runs tests and keeps 64K or more context.
The short answer
Model quality matters, but a useful result also needs enough fast memory, usable context, reliable tool calls and acceptable loop time. Pick the row closest to what you actually do.
Short suggestions while you type. Low latency matters more than a very large model or repository-wide context.
Plan: 8–16 GB fast memoryExplain code, draft tests and inspect selected files. A capable 14B–32B class model benefits from 16–32 GB.
Plan: 16–32 GBReads files, edits, searches and runs commands. Tool reliability and at least 64K usable context become central.
Plan: 24–32 GB or moreSeveral sessions multiply KV cache and runtime buffers. Large shared pools help capacity; discrete GPUs remain stronger for speed.
Plan: 64–128 GB capacityMemory before marketing labels
These are planning bands, not model guarantees. The exact quantization, context cache, runtime, operating system and concurrent sessions all share the available pool.
| Hardware class | Good fit | Practical model path | Main limit |
|---|---|---|---|
| 12–16 GB VRAM | Autocomplete, focused code chat | Smaller 7B–14B coding models in a suitable quantization | Large repositories and long agent context consume headroom quickly. |
| 24 GB VRAM | Capable single coding agent | Many 20B–30B quantized models; verify the actual file and context | 64K context can make an otherwise fitting model tight. |
| 32 GB VRAM | Fast repo agents and heavier code chat | Strong 20B–32B tier with more room for cache and tools | Large 70B weights usually require offload. |
| 64–128 GB unified memory | Large models, long context, several services | Capacity-focused 32B to quantized 70B paths | Fitting a model does not equal discrete-GPU throughput. |
| 128 GB coherent NVIDIA memory | CUDA-oriented model appliance | Large models and server-style development workloads | Arm64/Linux compatibility and software support must fit. |
A useful reference model
The official model card lists 30.5 billion total parameters, 3.3 billion activated parameters and a native 262,144-token context. It is a mixture-of-experts model: the active figure describes compute per token, while all expert weights still need storage and memory when loaded.
Do not buy by parameter count alone
A model that loads at 8K may fail or slow down at 64K. Repository agents also keep tool schemas, file excerpts, command output and conversation history in the prompt.
Agent and model are separate choices
Terminal coding agent that can inspect a repository, edit files and run commands. Ollama provides a direct launch path.
Best for:repo work and terminal workflowsFollow our local setupIDE extension with separate roles for chat, edit and autocomplete. Ollama models can be detected or configured explicitly.
Best for:VS Code and JetBrains workflowsOfficial Ollama guideIDE-based agent with local model support. Its own guide recommends planning hardware and compact prompts together.
Best for:guided agent work inside the editorOfficial local-model guideGit-aware terminal pair programmer. Its Ollama guide stresses using the chat endpoint and checking configured context.
Best for:focused edits with Git disciplineOfficial Ollama guideAn RTX 5090 has 32 GB of dedicated VRAM and is the strong path when the chosen model and context fit. Faster prompt processing and generation shorten edit–test–repair loops.
A 128 GB Ryzen AI Max+ 395 system can hold larger quantized models and more context in a compact chassis. It is a capacity decision, not a claim of RTX 5090-class throughput.
A local model can still run dangerous tools
Local inference limits where prompts are processed, but an agent may still read secrets, execute commands or use network tools if you allow it.
Commit first, then work in a branch or worktree so every change is reviewable and reversible.
Keep credentials, production data and private keys outside the accessible workspace.
Do not run the agent as administrator. Approve destructive commands and external network access deliberately.
Review the diff and run tests, linters and security checks before merging generated changes.
Complete-system shortlist
| Priority | System path | Why it fits | Verify before buying |
|---|---|---|---|
| Fast single agent | RTX 5090 tower | 32 GB dedicated VRAM and strong CUDA throughput for models that fit. | Exact desktop GPU, 64 GB+ system RAM, cooling, PSU and SSD. |
| Large-model capacity | Beelink GTR9 Pro or MINISFORUM MS-S1 Max | 128 GB unified memory in a compact Ryzen AI Max+ 395 system. | Exact 128 GB configuration, usable graphics allocation, OS and backend. |
| NVIDIA model appliance | ASUS Ascent GX10 | 128 GB coherent memory and NVIDIA software environment. | Arm64/Linux compatibility for every IDE extension, container and dependency. |
Buying questions
There is no single winner for every PC and workflow. For autocomplete, prefer a small low-latency model. For repository agents, prioritize tool calling, patch quality and enough usable context. Qwen3-Coder 30B-A3B is a useful current reference for a capable local agent, but test it against your languages, repository and runtime.
Yes for many 7B–14B quantized models and focused code chat. It is less comfortable for a 30B-class model, 64K context or parallel agents. Do not confuse system RAM with dedicated VRAM unless the platform explicitly provides a supported shared-memory path.
Choose 32 GB of high-bandwidth discrete VRAM when speed matters and the model fits. Choose a 128 GB shared pool when model size, long context or several services would exceed 32 GB. The larger number does not by itself mean faster inference.
Inference can remain local, but check the agent, extensions, telemetry, web tools, package managers and commands it invokes. An offline claim is only true after those network paths are disabled or controlled.
Primary documentation
Total and active parameters, native context, tool calling and steps for out-of-memory errors. Official Qwen model card.
Direct launch, local model context and manual provider configuration. Official integration guide.
Local Ollama endpoint, model IDs and tool-call troubleshooting. Official provider documentation.
From decision to working agent
Use a small reversible issue, measure whether the patch passes tests, and only then decide whether the bottleneck is model quality, context or hardware speed.