← All writing
Strix Halo

My $1,999 AMD Mini PC Outscored a $1,400 NVIDIA GPU on ML Training

Originally published on Medium. This archived article reflects the projects, opinions, and versions at the time of publication. View the Medium original ↗

In this article
93 autonomous experiments. $0.54 to benchmark the competition. The data speaks for itself.Before You Say “That’s Not Fair”The Recipe Was Optimized On-DeviceThe Real Problem Isn’t AMD’s HardwareThe Karpathy ConnectionWhat AMD Should DoThe Data
Illustration accompanying My $1,999 AMD Mini PC Outscored a $1,400 NVIDIA GPU on ML Training

93 autonomous experiments. $0.54 to benchmark the competition. The data speaks for itself.

I let an AI coding agent loose on my AMD Radeon 8060S — an integrated GPU inside a $1,999 mini PC sitting on my desk. No discrete GPU. No cloud subscription. 54 watts total system power. The kind of machine you’d use for spreadsheets.

It ran 93 experiments autonomously. Modify the neural network code, train for 5 minutes, keep improvements, discard regressions, repeat. Zero human intervention. I slept through most of it.

The result: a training quality score (val_bpb) of 1.227.

Then I rented an NVIDIA RTX 4090 on RunPod for $0.54 — fifty-four cents — and ran the exact same recipe.

The 4090 scored 1.844.

The $1,999 mini PC beat the $1,400 GPU by 33%.

Before You Say “That’s Not Fair”

You’re right. Let me be honest about the numbers.

The RTX 4090 is 6x faster on raw throughput — 320,000 tokens per second vs 51,000. If you need to train a production model, NVIDIA wins on speed. Full stop.

But speed isn’t the only metric that matters. Here’s the full picture:

The AMD system cost $1,999 complete — mini PC, ready to go. A 4090 build needs the GPU ($1,400), a CPU ($350), 64GB RAM ($180), motherboard ($180), power supply ($130), case and storage ($150). That’s $2,390+ for a desktop tower that still needs a monitor.

The AMD system draws 54 watts. The 4090 system draws 550+. Over a year of 8-hour days, that’s $19 in electricity vs $193.

The AMD system has 128GB of unified memory — up to 96GB allocatable as VRAM. The 4090 has 24GB. Hard limit. When you’re experimenting with model architectures overnight, memory headroom matters more than throughput.

And the AMD system achieved 25% utilization of its theoretical compute. The 4090 achieved 7.7%. The most expensive consumer GPU on the market is leaving 92% of its capability on the table.

The Recipe Was Optimized On-Device

This is the key insight: hardware-specific tuning matters more than raw hardware power.

Over 33 ablation experiments, the AI agent discovered two hyperparameters that are critical specifically on AMD:

HEAD_DIM 64 — reverting to the default 128 causes a 1.11% regression.
WARMDOWN_RATIO 0.7 — reverting to the default 0.3 causes a 1.08% regression.

The 4090 ran the same recipe without hardware-specific tuning. It’s like comparing a stock car to one with a custom tune — the engine matters, but so does the map.

The Real Problem Isn’t AMD’s Hardware

It’s the software.

NVIDIA has 15 years of ecosystem investment. pip install works. Flash Attention auto-selects the right kernel. PyTorch releases stable builds. Everything just works.

On AMD, I had to:

Use nightly PyTorch builds because stable ROCm doesn’t ship kernels for my GPU.

Set an undocumented environment variable that gives a 19x attention speedup — the difference between “AMD is unusable for ML” and “AMD is competitive.”

Manually unset a default shell variable that crashes PyTorch. Yes, AMD’s own default config crashes AMD’s own ML framework.

I also found 5 critical bugs where bf16 math produces NaN — small batch sizes, small attention head dimensions, deep networks, wide aspect ratios, and a sharp precision cliff at specific learning rates. All reproducible. All documented with exact steps.

These bugs cascade through the entire AMD ML ecosystem. ComfyUI breaks randomly. Users blame numpy. Numpy blames ML frameworks. ML frameworks blame numpy. Everyone points fingers while the actual bug sits in AMD’s bf16 accumulation layer.

The Karpathy Connection

This project is a fork of Andrej Karpathy’s autoresearch — an autonomous ML experimentation loop. We submitted AMD support upstream as a pull request. It was rejected. They preferred CUDA-only.

Karpathy now consults for NVIDIA.

So we forked his tool, ran it on AMD hardware, used it to benchmark his employer’s flagship GPU, and published the results. Open source is a beautiful thing.

What AMD Should Do

The hardware works. The silicon is capable. The unified memory architecture is a genuine advantage for ML workloads. Here’s what needs to happen:

One — make AOTriton flash attention the default, not hidden behind an undocumented environment variable. This single change transforms AMD from “unusable” to “competitive.”

Two — ship stable PyTorch wheels for consumer GPUs. Asking users to install from nightlies is asking them to leave.

Three — fix the bf16 accumulation bugs. We filed them with exact reproduction steps at the ROCm GitHub.

Four — stop pretending consumer GPUs don’t exist for ML. The Strix Halo with 128GB unified memory is the most accessible ML research platform on the market. Market it that way.

The Data

Everything is open source. Every experiment. Every bug. Every benchmark. Every script to reproduce it.

GitHub: github.com/bkpaine1/amdsense

ROCm bug report: github.com/ROCm/ROCm/issues/6034

93 experiments. $0.54 to benchmark the competition. The data speaks for itself.

Bkpaine builds AI tools and occasionally picks fights with GPU companies. The autonomous research agent was powered by Claude Code (Opus 4.6). The RTX 4090 was rented for less than the cost of a gas station coffee. The AMD mini PC is still on his desk, still drawing 54 watts, still outscoring hardware that costs more than his monthly rent.

More from the studio

Explore all 13 articles →See the current work →