Local terminal agent · Ollama endpoint · repository tools

OpenCode with Ollama: set up a local coding model correctly

The fast path is one command. A reliable setup also needs a tool-capable model, at least 64K configured context for repository work, enough memory for that context and a safe Git workspace.

Fastest supported route

Launch OpenCode through Ollama

Install both applications from their official sources, make sure Ollama is running, then use the integration command:

ollama launch opencode

Ollama asks you to choose an available model and starts OpenCode with an inline local configuration. The integration does not overwrite your persistent OpenCode configuration.

The check that prevents most false starts

Plan for 64K context before choosing the model

Ollama's OpenCode documentation calls for a context window of 64K or higher. That context contains repository excerpts, tools, command output and history. A model can load successfully at a smaller window yet fail during real agent work.

  • select a model with reliable tool calling
  • reserve memory for KV cache and runtime buffers
  • begin with one repository and one agent
  • test edits and commands on a disposable branch

Reproducible setup

Six steps from installation to a verified patch

Use the direct launch first. Add manual configuration only when you need a fixed model, another host or explicit context controls.

  1. 01

    Install from official sources

    Download Ollama and follow the OpenCode installation guide for your operating system. Confirm both commands are available in a new terminal.

  2. 02

    Choose a coding model with tools

    Autocomplete-only models are not automatically good agents. Check the model card for instruction following, tool or function calling and context support. Qwen3-Coder 30B-A3B is one current reference for agentic coding.

  3. 03

    Confirm the exact model ID

    Run ollama list and use exactly the displayed name and tag. A different tag may point to a different parameter count or quantization.

  4. 04

    Launch the integration

    Run ollama launch opencode, choose the local model and open a small test repository. Use ollama launch opencode --config if you want Ollama to prepare the configuration without starting an interactive session.

  5. 05

    Verify tools, not just chat

    Ask the agent to read one file, make a small reversible edit and run a harmless test command. A fluent answer does not prove that structured tool calls work.

  6. 06

    Measure the real bottleneck

    Watch memory use, prompt-processing delay, generation speed and test success. Reduce context or choose a smaller quantization if memory is exhausted; choose a stronger model if patches are fast but wrong.

Manual provider configuration

Connect OpenCode to Ollama explicitly

Create or update opencode.json when you want a fixed local provider. Replace both placeholders with the exact ID shown by ollama list.

{   "$schema": "https://opencode.ai/config.json",   "provider": {     "ollama": {       "npm": "@ai-sdk/openai-compatible",       "name": "Ollama (local)",       "options": {         "baseURL": "http://localhost:11434/v1"       },       "models": {         "your-model-id": {           "name": "Your local coding model"         }       }     }   } }

What each line changes

Keep the local endpoint local

  • baseURL: Ollama's OpenAI-compatible endpoint on this PC.
  • models: only IDs actually installed in Ollama should be listed.
  • npm: the OpenAI-compatible adapter specified by OpenCode's provider guide.
  • localhost: avoids sending requests to another machine by accident.

Do not replace localhost with a public bind address merely to make another device connect. Remote access needs authentication, firewall rules and deliberate network design.

Model + 64K context + headroom

How much memory does OpenCode with Ollama need?

OpenCode itself is not the main memory consumer. The loaded model, quantization, KV cache and parallel sessions set the hardware target.

Local model classPractical starting hardwareSuitable workloadBefore you commit
7B–14B quantized12–16 GB VRAM; 32 GB system RAMFocused chat, small edits, first agent testsConfirm 64K context fits and tools work reliably.
20B–24B quantized24 GB VRAM; 64 GB system RAMMore capable single-agent workLeave reserve for context, display use and build tools.
30B–32B quantized32 GB VRAM for speed, or 64 GB+ unified memory for capacityRepository agents and stronger code generationThe exact quantization may fit at 32K but become tight at 64K.
Large model or several agents96–128 GB unified/coherent memoryCapacity-first workflows and multiple servicesValidate runtime support and expect a different speed profile from a discrete GPU.

These ranges deliberately include working headroom but remain estimates. Use the actual model file, configured context and measured peak use. The PC Finder turns those inputs into a hardware class.

Troubleshooting by symptom

Fix the layer that is actually failing

SymptomLikely causeCheckUseful next move
Model not listedOllama is stopped, wrong tag or launch picker limitationollama list and the Ollama servicePull the exact tag or define the installed ID in opencode.json.
Chat works, edits do notWeak or incompatible tool callingModel card and a one-file tool testUse a tool-capable coding model; do not force capabilities the model lacks.
Out of memory after a whileContext cache or concurrent sessionsConfigured context and peak memoryClose other sessions, reduce context deliberately or use a smaller quantization.
Very slow before first tokenLong prompt ingest, CPU offload or cold model loadGPU offload, prompt length and storage activityKeep the model in fast memory, reduce irrelevant context or choose a faster hardware path.
Agent repeats or loses filesContext truncation or weak repository retrievalActual context limit and OpenCode logsIncrease context only if memory allows; split the task and keep instructions concise.
Endpoint unreachableWrong host/port or service not runninghttp://localhost:11434/v1 and local firewallRestore the default local endpoint before attempting remote access.

Local inference is only one privacy layer

Give the agent a safe workspace

OpenCode can read files, edit them and run terminal commands. Treat model output like an untrusted junior contributor: useful, fast and always reviewed.

Commit first

Use a clean branch or worktree and keep an easy rollback point.

Narrow the folder

Open only the project needed for the task; exclude credentials and production dumps.

Review commands

Do not approve destructive, privileged or networked commands without understanding them.

Verify the patch

Inspect the diff and run the project's tests, lint and security checks before merging.

OpenCode and Ollama FAQ

Answers before you change hardware

Can OpenCode use an Ollama model without an API key?

Yes. The local Ollama endpoint on localhost does not need a cloud-provider API key. Keep that endpoint local unless you deliberately add authentication and network controls.

Why does the same model work in chat but fail in OpenCode?

An agent needs structured tool calls, more prompt context and enough output discipline to edit files and run commands. A model that writes convincing chat text may not support those operations reliably.

Does OpenCode really need 64K context?

Ollama's official OpenCode integration recommends 64K or higher for local models. Small focused tasks may use less, but repository work can quickly fill the prompt with tools, files, output and history.

Should I choose an RTX 5090 or a 128 GB mini PC?

Choose the RTX 5090 when the model plus context fit in 32 GB and loop speed is the priority. Choose 128 GB unified memory when larger models, longer context or several services need capacity beyond 32 GB.

Keep commands current

Official setup references

Still choosing the PC?

Size the model and context before the complete system

Enter your model class, context and workload. The result separates fast dedicated VRAM from larger shared-memory capacity.

Open the PC Finder