← All writing
Strix Halo

I tuned llama.cpp on a Strix Halo mini-PC and it beats Ollama on the same weights — here’s the recipe (and why “Vulkan beats ROCm” is a myth)

Originally published on Medium. This archived article reflects the projects, opinions, and versions at the time of publication. View the Medium original ↗

In this article
The finding that changes everythingThe setupThe numbersThe head-to-head that proves the pointThe honest part (because you’ll ask)No sidecar
llama.cpp direct no sidecar or wrapper
llama.cpp direct no sidecar or wrapper

I have an AMD Strix Halo box — a Ryzen AI Max+ 395, Radeon 8060S iGPU (gfx1151), 128 GB of unified memory. On paper it’s a monster for local LLMs: that unified pool means the GPU can address ~100 GB of model weights with no discrete-VRAM ceiling. In practice, most people running one are leaving half the performance on the floor, and repeating two pieces of folklore that cost them the other half:

  1. “Just use Ollama / LM Studio / Lemonade.” Wrappers on wrappers on llama.cpp, each one adding a layer between you and the engine.
  2. “Vulkan is faster than ROCm on this APU.” You’ll read this in every Strix Halo thread and in more than one benchmark writeup.

I got tired of both, built llama.cpp directly against ROCm/HIP with the gfx1151 tuning that actually matters, and measured it. Here’s what I found.

The finding that changes everything

Stock HIP really is slow on gfx1151. If you cmake -DGGML_HIP=ON and call it a day, your prompt processing is bad, and that's exactly where "Vulkan wins" comes from — those comparisons are testing untuned HIP.

Tuned HIP wins. Two levers do most of the work:

That’s the whole secret. It’s not exotic. It’s a build flag and an env var that the ecosystem’s wrappers hide from you.

The setup

The build, in full:

export ROCM_PATH=/opt/rocm
export HIPCXX="$(hipconfig -l)/clang"
cmake -S llama.cpp -B llama.cpp/build \
  -DGGML_HIP=ON \
  -DGPU_TARGETS=gfx1151 \
  -DGGML_HIP_ROCWMMA_FATTN=ON \
  -DGGML_HIP_NO_VMM=ON \
  -DCMAKE_BUILD_TYPE=Release
cmake --build llama.cpp/build -j"$(nproc)"

-DGGML_HIP_NO_VMM=ON matters: HIP's virtual-memory manager is buggy on gfx1151 and will corrupt allocations if you leave it on. And on a 128 GB unified box, set your BIOS dedicated-VRAM to the minimum and let the driver use GTT — there's no performance difference between "VRAM" and system memory on this APU, and GTT is reclaimable.

The numbers

All measured on the box, ROCBLAS_USE_HIPBLASLT=1 llama-bench -m <model> -ngl 999 -fa 1. pp512 = prompt processing, tg128 = token generation, both tokens/sec, higher is better.

Local model throughput (tokens per second)
ModelParametersQuantizationpp512tg128
Qwen3–4B4 B denseQ4_K138869.7
Qwen3–30B-A3B30.5 B / 3 B active (MoE)Q4_K_M84373.7
gpt-oss-20B20.9 B (MoE)MXFP496173.0

Read the middle row again: a 30-billion-parameter model generating at ~74 tokens/sec on a mini-PC. It’s a MoE — only ~3 B params fire per token — so it actually out-generates the 4 B dense model despite carrying 7× the weights. On this hardware, MoE is the play.

And it holds at depth. Same 30B, with 4096 tokens already in the KV cache:

MetricDepth 0Depth 4096
pp512843209
tg12873.747.5

Still ~47 t/s of generation with real context loaded. That graceful long-context falloff is exactly where the tuned rocWMMA-FA + hipBLASLt path pulls ahead of Vulkan.

The head-to-head that proves the point

Everyone defaults to Ollama. So I ran the same GGUF blob — llama3.2:3b, the exact file Ollama had already pulled — through both runtimes, both GPU-offloaded on the same box:

RuntimeGeneration
Tuned llama.cpp, direct85.1 t/s
Ollama82.1 t/s

~4% faster on generation for a tiny 3B where both are already saturated — and the real gap is in prompt processing (1918 t/s direct on that blob). Strip the wrapper, tune the flags, and you beat the tool everyone reaches for, on identical weights. That’s not a claim. That’s the same file, two runtimes, one box.

The honest part (because you’ll ask)

No sidecar

The whole thing is a repo you can reproduce:

git clone <repo>
cd halo-llamacpp
source env.sh          # ROCm env + ROCBLAS_USE_HIPBLASLT=1
./build.sh             # clone/pull llama.cpp @ the pinned commit, configure, build
./serve.sh model.gguf  # tuned llama-server: -ngl 999 -fa on -b 2048 -ub 512 --no-mmap

No Ollama, no container, no Python server in front. The binary is the engine. There’s a bench.sh so you can reproduce every number above on your own box.

If you’ve got a Strix Halo, clone it, run bench.sh, and drop your numbers. I want the real cross-box table — not folklore.

git https://github.com/bkpaine1/halo-llamacpp · measured 2026 on Strix Halo gfx1151 / ROCm 7.2.4 / llama.cpp @ 1425386.

Built at MindSpark Business Studio — msbs.com. We ship real tools for people who do real work: MultiScan (multi-vendor network config capture, one exe, no install) and StrixScreener (AI resume screening that reads under the fluff). If this writeup was useful, come see what else we make.

More from the studio

Explore all 13 articles →See the current work →