Running LLMs Locally with Ollama: Quantization Tags, num_ctx, GPU Offload and the API
Key takeaways
Ollama makes 'ollama run llama3.1' work in one command, but the defaults hide a few traps: a context window that truncates long prompts without an error, models that spill onto the CPU when VRAM runs out, and an API with no authentication once you bind it to 0.0.0.0. This guide covers the commands, the API, and those trade-offs.
What Ollama does (and what it doesn’t)
Ollama is a local model runner. It downloads open-weight models in GGUF format, loads them onto your GPU (or CPU), and serves them over an HTTP API on port 11434. The inference engine underneath comes from the llama.cpp lineage. What Ollama adds is model management with Docker-like names and tags, automatic GPU detection and layer offloading, loading and unloading models on demand, and a stable API.
It’s a good fit for development, private data you don’t want to send to a cloud API, offline use, and small internal tools. It’s not a replacement for a frontier cloud model on hard reasoning tasks, and it’s not a high-throughput serving stack for many concurrent users. For that, dedicated servers like vLLM are designed differently, with batching and paged KV caches as the central concern.
“Free” also needs a qualifier. Ollama itself is open source (MIT), but each model comes with its own license. Some are Apache 2.0 or MIT. Others, like Meta’s Llama models, use custom licenses with their own conditions. Check the license on the model’s library page before you ship a product on it.
Install and first run
# macOS: download the app from ollama.com, or
brew install ollama
# Linux: installs a binary and a systemd service
curl -fsSL https://ollama.com/install.sh | sh
# Windows: installer from ollama.com (runs in the tray)
The desktop apps and the Linux service start the server automatically. Run ollama serve yourself only if nothing is already listening on 11434. Otherwise you get Error: listen tcp 127.0.0.1:11434: bind: address already in use, which means the server is already running, not that something is broken.
ollama run llama3.1 # pulls on first use, then opens an interactive chat (/bye to exit)
ollama pull qwen2.5-coder:7b # download without running
ollama list # installed models and their sizes
ollama ps # models currently loaded, their memory and CPU/GPU split
ollama show llama3.1 # architecture, parameters, quantization, context length, license
ollama rm llama3.1:70b # free disk space
Names, tags and quantization
A model reference is name:tag. Without a tag you get latest, which for most library models means a mid-size instruct model at 4-bit quantization. Tags encode size and quantization, and they’re worth reading:
| Tag fragment | Meaning |
|---|---|
8b, 70b | Parameter count |
instruct / text | Chat-tuned vs base model |
q4_K_M | ~4.5 bits per weight, “K-quant, medium”. The usual default: good quality-to-size ratio |
q5_K_M, q6_K | More bits: slightly better quality, more memory |
q8_0 | 8-bit: close to original quality, about twice the size of q4 |
fp16 | Unquantized half precision: largest, mostly useful for comparison |
A rough sizing rule: memory ≈ parameters × bits-per-weight / 8, plus the KV cache for the context. That’s why an 8B model at q4 is a download of around 5 GB, and a 70B model at q4 needs over 40 GB before any context. Going from q4 to q8 on the same model is usually a smaller quality jump than going to a bigger model at q4, so if you have memory to spare, a larger model is often the better use of it.
Downloaded models live under ~/.ollama/models (on Windows, %USERPROFILE%\.ollama\models; the Linux service install uses the ollama user’s home, typically /usr/share/ollama/.ollama/models). Layers are content-addressed blobs, so tags that share weights don’t use extra space. To store models on another disk, set OLLAMA_MODELS to the new path for the server process. A variable set in your shell doesn’t affect a server started by systemd or the desktop app.
The context window trap (num_ctx)
This is the problem that costs people the most time. Ollama loads each model with a context window, num_ctx, that is often much smaller than what the model supports. For a long time the default was 2048 tokens. It has since been raised, and newer versions may choose it based on available memory, but it’s still frequently smaller than a model’s advertised 128K. When a prompt doesn’t fit, Ollama doesn’t return an error. It truncates. The server log shows a line like:
level=WARN msg="truncating input prompt" limit=4096 prompt=11873 ...
The model then answers on the part that survived. With chat templates, the system prompt or the beginning of your pasted document is often what gets cut. The symptoms look like model stupidity: it “ignores” instructions, summarizes only the end of a file, or a RAG pipeline “forgets” the retrieved context.
I’ve been bitten by exactly this: a summarization script that worked on short test files and produced confident but partial summaries on real documents. Nothing failed, and the answers looked plausible. Only the server log showed the truncation. Since then, the first thing I do with any Ollama-backed pipeline is set num_ctx explicitly and check the log once with a realistic input.
There are three ways to set it:
# 1. Per request (native API only)
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1",
"messages": [{"role": "user", "content": "Summarize the following ..."}],
"options": { "num_ctx": 16384 },
"stream": false
}'
# 2. Server-wide default for every model (set it where the server runs)
OLLAMA_CONTEXT_LENGTH=16384 ollama serve
# 3. Baked into a derived model
FROM llama3.1
PARAMETER num_ctx 16384
ollama create llama3.1-16k -f Modelfile
The trade-off is memory. The KV cache grows linearly with context length, so a 32K context can need several extra gigabytes on top of the weights. That can push a model that used to fit in VRAM partly onto the CPU (next section). Changing num_ctx also forces the model to reload. Clients that send different num_ctx values for the same model cause repeated reloads, which look like random multi-second stalls.
GPU, CPU and partial offload
Ollama detects NVIDIA (CUDA), AMD (ROCm on supported cards) and Apple Silicon (Metal) automatically. When a model plus its context doesn’t fit in VRAM, Ollama puts as many layers on the GPU as fit and runs the rest on the CPU. ollama ps shows where things landed:
NAME ID SIZE PROCESSOR UNTIL
llama3.1:70b ... 47 GB 52%/48% CPU/GPU 4 minutes from now
qwen2.5:7b ... 6.0 GB 100% GPU 4 minutes from now
Anything short of 100% GPU usually means a big slowdown, because the CPU layers become the bottleneck for every token. Options, roughly in order of effectiveness:
- use a smaller model or a lower quantization,
- lower
num_ctx, - reduce
OLLAMA_NUM_PARALLEL: each parallel request slot needs its own share of context memory, - close other GPU consumers (browsers and games take VRAM too; check
nvidia-smi).
On Apple Silicon, the GPU shares unified memory with the system, so “VRAM” is a portion of your RAM. A 16 GB Mac runs 7B-8B models comfortably and struggles with anything much larger.
If the GPU isn’t used at all (100% CPU on a machine with an NVIDIA card), check the server log at startup. It reports which GPUs it discovered and why it skipped any, such as a driver that’s too old or an unsupported compute capability. In Docker you need the NVIDIA Container Toolkit and --gpus=all. Without them, the container silently runs on the CPU.
Keep-alive and concurrency
Loaded models are unloaded after 5 minutes of inactivity by default, and the next request pays the load time again. Control that with keep_alive per request ("10m", -1 to keep it loaded indefinitely, 0 to unload immediately) or with OLLAMA_KEEP_ALIVE on the server. OLLAMA_MAX_LOADED_MODELS limits how many models stay resident at once. On a shared GPU box, an embedding model and a chat model evicting each other on alternating requests is a common cause of unexplained latency.
The API
Native endpoints
# Chat (streams NDJSON by default; "stream": false for one JSON response)
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1",
"messages": [
{"role": "system", "content": "You are a senior backend engineer. Be concise."},
{"role": "user", "content": "When should I pick gRPC over REST?"}
],
"stream": false
}'
# Structured output: pass a JSON schema in "format"
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1",
"messages": [{"role": "user", "content": "Extract: Widget Pro costs $29.99 and is in stock"}],
"format": {
"type": "object",
"properties": {"name": {"type": "string"}, "price": {"type": "number"}, "in_stock": {"type": "boolean"}},
"required": ["name", "price", "in_stock"]
},
"stream": false
}'
# Embeddings (current endpoint is /api/embed; /api/embeddings is the older one)
curl http://localhost:11434/api/embed -d '{"model": "nomic-embed-text", "input": ["first text", "second text"]}'
Structured output constrains the response format, not its correctness. A small model can return perfectly valid JSON with the wrong price. Still validate the values you care about.
OpenAI-compatible endpoint
Ollama also serves /v1/chat/completions, /v1/embeddings and /v1/models, so existing OpenAI SDK code can target it:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama") # key required by the SDK, ignored by Ollama
resp = client.chat.completions.create(
model="llama3.1",
messages=[{"role": "user", "content": "Explain idempotency keys in two sentences."}],
)
print(resp.choices[0].message.content)
“Drop-in” works for basic chat, streaming and embeddings. Two catches:
- No
num_ctxin the request. The OpenAI format has no field for it, so requests through/v1run with the server default, and the truncation trap above applies. Use a Modelfile-derived model orOLLAMA_CONTEXT_LENGTH. - Behavior still depends on the model. Tool calling and JSON mode go through the same endpoint, but a 7B local model follows tool schemas far less reliably than a large hosted model. Code that worked against a cloud API needs re-testing, not just a new
base_url.
Python and JavaScript clients
The official ollama packages wrap the native API, including options and keep_alive:
import ollama # pip install ollama
resp = ollama.chat(
model="llama3.1",
messages=[{"role": "user", "content": "How do I reverse a list in Python?"}],
options={"num_ctx": 8192, "temperature": 0.2},
)
print(resp["message"]["content"])
for chunk in ollama.chat(model="llama3.1", messages=[{"role": "user", "content": "Tell me a story"}], stream=True):
print(chunk["message"]["content"], end="", flush=True)
import ollama from 'ollama'; // npm install ollama
const stream = await ollama.chat({
model: 'llama3.1',
messages: [{ role: 'user', content: 'Explain async/await in JavaScript.' }],
options: { num_ctx: 8192 },
stream: true,
});
for await (const part of stream) process.stdout.write(part.message.content);
For LangChain, the langchain-ollama package provides ChatOllama and OllamaEmbeddings, which take num_ctx as a constructor argument. See LangChain 1.x in Practice for building chains and RAG on top.
Modelfiles: pinning behavior
A Modelfile creates a named variant with its own system prompt and parameters. That’s the cleanest way to make settings stick for every client, including /v1 clients:
FROM qwen2.5-coder:7b
SYSTEM """You are a senior TypeScript reviewer. Point out type-safety issues first. Be concise."""
PARAMETER temperature 0.2
PARAMETER num_ctx 16384
ollama create ts-reviewer -f Modelfile
ollama run ts-reviewer
ollama show --modelfile llama3.1 prints the Modelfile of an existing model, including its chat TEMPLATE. Don’t override the template unless you know the model’s expected prompt format. A wrong template produces rambling output or the model talking to itself.
Exposing Ollama on the network
By default the server binds to 127.0.0.1:11434. To reach it from another machine or a container, people set OLLAMA_HOST=0.0.0.0. Understand what that does: the Ollama API has no authentication. Anyone who can reach the port can generate text on your GPU, list your models, pull new ones, and delete existing ones through the same API your clients use. Internet scans for open 11434 ports are a known phenomenon, and a publicly reachable Ollama is an open, free compute endpoint.
Where you set the variable depends on how the server runs:
# Linux (systemd service)
sudo systemctl edit ollama.service
# [Service]
# Environment="OLLAMA_HOST=0.0.0.0:11434"
sudo systemctl daemon-reload && sudo systemctl restart ollama
# macOS app
launchctl setenv OLLAMA_HOST "0.0.0.0:11434" # then restart the Ollama app
Safer patterns, from simplest:
- SSH tunnel for personal remote use:
ssh -L 11434:localhost:11434 gpu-box, and the server stays on localhost. - Private network only: bind to the LAN or VPN interface address instead of
0.0.0.0, plus a firewall rule. - Reverse proxy with auth for shared use:
server {
listen 443 ssl;
server_name llm.internal.example.com;
location / {
auth_basic "LLM";
auth_basic_user_file /etc/nginx/.htpasswd;
proxy_pass http://127.0.0.1:11434;
proxy_set_header Host localhost:11434;
proxy_buffering off; # stream tokens as they're generated
proxy_read_timeout 300s; # long generations
}
}
proxy_buffering off matters for streaming. Without it, nginx collects the whole response and the UI gets everything at once at the end. See Nginx Reverse Proxy Configuration for the TLS side.
Browser apps calling Ollama directly are a separate concern: cross-origin requests are restricted by OLLAMA_ORIGINS. Add your app’s origin there instead of * if the page is served from somewhere else.
Open WebUI
For a ChatGPT-style interface on top of your models:
docker run -d -p 3000:8080 \
--add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data \
--name open-webui \
ghcr.io/open-webui/open-webui:main
The container reaches Ollama on the host through host.docker.internal, so Ollama can stay bound to localhost. Open WebUI has its own user accounts. The first account created becomes the admin, so create it immediately if the port is reachable by others.
Choosing a model
Model names change quickly, so treat this as a starting point and check the Ollama library for current versions:
| Use case | Starting point | Notes |
|---|---|---|
| General chat on 8-16 GB | llama3.1:8b, qwen2.5:7b, gemma3:4b | Fit fully on most modern GPUs at q4 |
| Code | qwen2.5-coder:7b (or larger sizes) | Keep num_ctx high for whole-file context |
| Reasoning | deepseek-r1:8b and larger | Emits long thinking output before the answer: slower and uses more tokens |
| Vision | llama3.2-vision, gemma3 (4b and larger) | Images go in the images field of a message |
| Embeddings for RAG | nomic-embed-text, mxbai-embed-large | Vectors from different embedding models aren’t compatible |
Evaluate on your own tasks. Leaderboard positions say little about how a 7B model handles your prompts, and quality differences between quantizations of the same model are best judged on real inputs too.