
I run local vision-language models on an AMD Strix Halo mini PC for a camera project — feed it a frame, ask it what’s going on. The text models scream on this thing. But the moment I pointed a vision model (qwen3-vl:4b) at a 4K camera frame, inference fell off a cliff: ~40 seconds per image. Worse, while it churned, one CPU core group sat pegged near 1000% (about 10 cores) and the GPU was basically idle for the image step.
A 128GB unified-memory APU with a perfectly capable iGPU, and it’s grinding images on the CPU. Something was off. Here’s what’s happening and how to get it back.
The symptom
- Vision inference is slow — tens of seconds per frame.
- During the image step, CPU is maxed (a big multi-core spike) and the GPU is idle.
- Text-only generation on the same model/GPU is fine.
That last point is the tell. If text is fast but images are slow, the language model is on the GPU — it’s the vision encoder that isn’t.
How to diagnose it
Watch the Ollama logs while you run a vision request. The smoking gun is this line:
disabling multimodal projector offload reason=shared-memory-gpu
Right after that, Ollama launches its internal llama-server with the flag --no-mmproj-offload. Translation: the CLIP/vision projector — the image encoder, the "mmproj" — is being forced onto the CPU, while the language half stays on the GPU. That split is exactly why you see a massive CPU spike for the image and an idle GPU.
You can confirm the CPU side with htop or top during a request: you'll see one process eating ~1000% CPU for the duration of the encode.
Why Ollama does this
It’s not a bug, it’s a deliberate, hardcoded heuristic. Ollama disables GPU projector offload for any shared-memory GPU — that’s every iGPU and APU, where the GPU borrows system RAM instead of having dedicated VRAM. The reasoning is OOM protection: on some shared-memory setups, loading CLIP onto the GPU can blow up memory or produce garbage, so Ollama plays it safe.
The frustrating part: it’s a blanket rule. It doesn’t matter that you have 128GB of unified memory sitting mostly free — if the GPU is shared-memory, the projector goes to CPU. And as of Ollama 0.30.11, there is no environment variable to override it. (See Ollama GitHub issues #10889 and #13742 — this has bitten a lot of APU users.)
The fix: run llama-server directly with the projector on the GPU
Here’s the good news. Ollama is built on llama.cpp, and it bundles the llama-server binary and has already downloaded your model. You can run that same binary yourself, on a different port, with the projector offload enabled — and point your app at it.
First, find the bundled binary and the model files. The binary typically lives here:
# The llama-server Ollama ships with: ls /usr/local/lib/ollama/llama-server
# Find the model files Ollama already pulled. ls ~/.ollama/models/manifests/registry.ollama.ai/library/qwen3-vl/ ls ~/.ollama/models/blobs/
The model is stored as content-addressed blobs (files named sha256-...), and the manifest JSON maps the model name to those blobs. Pull the model-layer digest from the manifest:
python3 -c "import json; m=json.load(open('$HOME/.ollama/models/manifests/registry.ollama.ai/library/qwen3-vl/4b')); print(next(l['digest'].replace(':','-') for l in m['layers'] if l['mediaType']=='application/vnd.ollama.image.model'))"
# -> sha256-... (the file lives in ~/.ollama/models/blobs/)Important — qwen3-vl ships as one combined GGUF. The vision projector is packed into the same file as the language weights, so you pass the same blob to both --model and --mmproj — there is no separate mmproj file to hunt for. (Some other vision models do split the projector into its own layer; if yours does, your manifest will show a second projector/mmproj layer and you'd point --mmproj at that one. For qwen3-vl:4b, it's one file.)
Then launch your own vision server with the projector on the GPU. I reused the llama-server binary and ROCm backend Ollama already bundles, so the three env vars below point at them — if you built llama.cpp yourself, drop the env lines and use your own binary:
BLOB=~/.ollama/models/blobs/sha256-<model-layer-digest>
GGML_BACKEND_PATH=/usr/local/lib/ollama/rocm_v7_2/libggml-hip.so \ LD_LIBRARY_PATH=/usr/local/lib/ollama:/usr/local/lib/ollama/rocm_v7_2 \ ROCR_VISIBLE_DEVICES=0 \ /usr/local/lib/ollama/llama-server \ --model "$BLOB" \ --mmproj "$BLOB" \ --mmproj-offload \ --host 127.0.0.1 --port 8081 \ -c 8192 -np 1 \ --flash-attn on \ --image-min-tokens 1024 \ -b 1024 -ub 1024 --context-shift --keep 4
The flag that does the work is --mmproj-offload — the one Ollama refuses to set for you. Note --model and --mmproj point at the same blob (combined GGUF, see above). If you find the language layers aren't all landing on the GPU, add --n-gpu-layers 999. Then point your app's OpenAI-compatible calls at http://127.0.0.1:8081.
Gotcha: if your vision model is a “thinking” model, you’ll get empty answers
This one cost me an hour, so let me save you the trouble. qwen3-vl:4b is a thinking variant — left to the default template it opens every reply with a long <think>…</think> reasoning monologue before the actual answer. That's harmless until you add a token cap to keep responses short (which you'll want to). Then the cap gets eaten by the hidden thinking and you get a truncated, empty answer. (Ollama's think: false doesn't save you — it just hides the monologue in a separate field; the tokens are still spent.)
The fix is a tiny chat template that prefills a closed, empty think block on the assistant turn, so the model skips the monologue and answers directly. Save this as nothink.jinja:
{%- for message in messages -%}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>\n' -}}
{%- endfor -%}
{%- if add_generation_prompt -%}
{{- '<|im_start|>assistant\n<think>\n\n</think>\n\n' -}}
{%- endif -%}Then add --jinja --chat-template-file /path/to/nothink.jinja to the launch command above. Clean, direct, single-digit-second descriptions, no output-stripping needed. If your vision model isn't a thinking model, skip this entirely.
Make it a persistent “vision sidecar”
You don’t want to babysit a terminal. Wrap that command in a small systemd service so it stays running and the model stays warm in memory. Run it on its own port and leave Ollama alone for everything else — text generation, embeddings, model management. Ollama handles those untouched; your sidecar just owns vision. Best of both worlds, and you skip the cold model-load on every call.
A minimal unit is just ExecStart= pointing at the command above, plus Restart=always. Enable it and forget it.
The benchmark
I ran this apples-to-apples: two isolated llama-server instances, same 4K frame, identical except for the offload flag.
Run Projector Cold (first) Warm (cached) --no-mmproj-offload (Ollama default) CPU 39.8 s 1.0 s --mmproj-offload (the fix) GPU 10.0 s 1.0 s
About 4x faster on a cold image encode. The output was coherent and identical in quality with GPU offload — so on this gfx1151 box, GPU CLIP is perfectly stable. (Warm/cached calls are ~1s either way, because at that point you’re not re-encoding.)
Two bonus tips that stack
1. Downscale the image before you send it. Image-encode time scales hard with resolution, and a 4K frame is overkill for most vision tasks. Resize to roughly 720p / ≤1280px on the long edge before the request. You’ll often shave a big chunk off the encode with no meaningful loss in answer quality — the model isn’t reading license plates, it’s telling you there’s a person on the porch.
2. Keep the model resident. That’s the sidecar point again: a persistent server means the weights stay loaded, so you don’t pay a cold model-load on top of the encode for every single call. These two stack with the GPU-offload fix.
The honest caveat — test your own hardware
Ollama disables GPU projector offload on shared-memory GPUs for a reason. Some APUs really may OOM or spit out garbage with CLIP on the GPU. I’m not telling you Ollama is wrong to be cautious — I’m telling you it’s too cautious for this box.
So:
- Test on your own hardware. Run a few real frames through the sidecar.
- Verify the output is coherent — not just fast, but correct. If it’s hallucinating nonsense, GPU CLIP isn’t stable on your chip; back off.
- Keep Ollama as a fallback. Since the sidecar is a separate service on a separate port, Ollama is still right there if you need it.
Tested on Strix Halo gfx1151 — your mileage may vary. But if you’ve got a capable iGPU sitting idle while one CPU sweats through every image, this is very likely your 4x.