THE RIG

Five GPUs in a gaming PC.
My local LLM rig

Zero server parts. An ordinary home-built PC that also plays games, with five GPUs crammed in using every slot plus riser cables, for 80GB of VRAM. Every local LLM in AutoDev LG runs on this one machine.

Inside the PC with the side panel off: three GPUs lying flat on the left, two on a vertical bracket on the right, a big air cooler in the middle
Side panel off. Three cards lying flat on the left, two on the vertical bracket on the right.
As for the cable management — please pretend you didn't see it.
5
GPUs (RTX 5070 Ti ×1 + 5060 Ti ×4)
80 GB
total VRAM
128 GB
main memory
~220 GB
of models in use

01Hardware

PartWhat's in itNote
CPUIntel Core i9-14900KF (24 cores / 32 threads)Bought for gaming. All 32 threads used from WSL
Memory128 GB (DDR4, 32 GB × 4)About half (62 GB) is given to WSL
MotherboardASRock B760 Pro RS/D4 WiFiA regular gaming board — server boards were out of reach
CaseLIAN LI O11D EVO XLSwapped in to fit five GPUs. Shockingly huge when it arrived
GPU bracketLIAN LI O11DEXL-1X (vertical GPU bracket, an option made for the O11D EVO XL)Used to mount the GPUs connected by riser cables
Power supplyThermaltake TOUGHPOWER GT 1200W (ATX 3.1, PS-TPT-1200FNFAGJ-3)The five GPUs' power limits add up to 1,020 W, but in practice the whole PC stays within 750 W even at full load
StorageNVMe SSD 2 TB × 2 + USB HDD 6 TBModels live on NVMe (loading from USB is much slower)
NetworkI-O DATA ET10G-US3C (USB Type-C 10-gigabit LAN adapter, Realtek chip)With the PCIe slots taken by GPUs, there was nowhere left for a 10G NIC, so it runs over USB — plugged into the front USB 3.2 Gen2 Type-C port, the only 10Gb/s USB port on this board. Link speed: 10 Gbps

02The five GPUs

Six empty GPU boxes in a row: an RTX 5070 Ti, an RTX 5070 and four RTX 5060 Ti
A parade of empty boxes. From the left: RTX 5070 Ti (PNY), the RTX 5070 (MSI) retired on 8/27, then four RTX 5060 Ti.
That many boxes in just over a month. I've lost it.
...Oh. There's a spare 5070 lying around...

Three cards plug straight into PCIe slots. The two that didn't fit go through riser cables (PCIE2, and M2_2 — an M.2 slot). As a result, every card has a different link width.

GPUSlotVRAMPower limitLink (measured)
RTX 5070 TiPCIE116 GB300 Wx16 direct
RTX 5060 Ti ①PCIE216 GB180 Wx1 riser
RTX 5060 Ti ②PCIE416 GB180 Wx1 direct
RTX 5060 Ti ③PCIE316 GB180 Wx4 direct
RTX 5060 Ti ④M2_216 GB180 Wx2 riser (M.2 adapter)

The four RTX 5060 Ti cards come from MSI (VENTUS 2X), Palit (INFINITY) × 2 and GAINWARD (PYTHON). Zero sense of uniformity.
"Link" is the current lane count reported by the NVIDIA driver. The five power limits add up to 1,020 W.

Which GPU sits in which slot?

Using the motherboard layout and block diagram, we matched the PCI locations Windows reports against each link's measured width. The bar length is the lane count (x1–x16).

CPUCore i9-14900KF
PCIE1Gen4 x16
RTX 5070 Tidirect
M2_1Gen4 x4
NVMe SSD
ChipsetIntel B760
PCIE3Gen4 x4
RTX 5060 Ti ③direct
M2_3Gen4 x4
NVMe SSD
M2_2Gen4 x2
RTX 5060 Ti ④riser (M.2-to-PCIe adapter)
PCIE2Gen3 x1
RTX 5060 Ti ①riser
PCIE4Gen3 x1
RTX 5060 Ti ②direct
PCIe 4.0 lanesPCIe 3.0 lanesunused

Slot names follow the ASRock B760 Pro RS/D4 WiFi manual.

Does x1 even work?
For MoE models, it does (dense models are covered in section 04 below). AutoDev LG splits one model across all five cards by layers (llama.cpp's "layer split"). That way, during generation the GPUs only pass small data at the layer boundaries, so even thin links run reasonably well. Where the thin links really show is when a model is being loaded.
Watch out for tensor parallelism, though. Splitting each layer across several GPUs makes them exchange data constantly, and on this setup performance collapses completely. Few people will build something this eccentric, but consider yourself warned.

03Software

Windows 11 Pro→WSL2 (Ubuntu 24.04)→CUDA 13.3→llama.cpp (llama-server)→AutoDev LG (LangGraph + OpenHands)
LayerWhat runsNote
OSWindows 11 Pro + WSL2 (Ubuntu 24.04.5 LTS)A Windows-native version was tried too, but it was slow on long contexts, so it went back to WSL
GPU stackNVIDIA driver 617.14 / CUDA 13.3
Inference enginellama.cpp official (build 11151) + Unsloth build (build 11160)The Unsloth build is for Qwen3.8-Flash-Next's MTP; the other models use the official build
LanguagePython 3.12AutoDev LG itself is written in Python
WorkflowLangGraph 1.2Runs the steps in order and decides where to go back when something fails
AI agentOpenHands SDK 1.46An open-source platform for autonomous AI software-engineering agents. The worker's local LLM edits files and runs commands through it

04The models

One model at a time is loaded across all five GPUs. Every time the role changes the model is swapped, which means reading tens of gigabytes again.

ModelTypeRoleQuantization · sizeSpeed-up (MTP etc.)
Qwen3.8-Flash-NextMoELeader / workerAP-Q4_K_XL · 95 GBMTP head (separate file · Q8_0 · 3.9 GB)
Qwen3.8-27BDenseWriting test codeUD-Q4_K_M · 16 GBMTP (built into the model)
Gemma-4-26B-A4B-itMoE (about 4B of 26B active)ReviewerBF16 · 48 GBMTP helper model (Q8_0 · 0.4 GB)
gpt-oss-120bMoEBackupMXFP4 · 60 GBEAGLE3 (Q8_0 · 0.8 GB) *
Dense models? Not really. This is basically an MoE-only machine.
Dense models (which use every weight for every token) don't perform when split across several GPUs on this setup. The bottleneck is PCIe bandwidth: the thin x1 and x2 links, and the DMI Gen4 x4 that four of the five cards share through the chipset, get choked.
With an MoE (Mixture of Experts) model, the total parameter count can be large, but only a few "experts" work on each token. Gemma-4-26B-A4B-it, for example, uses about 4B of its 26B parameters each time (the "A4B" in its name).
So what runs here is almost all MoE: Qwen3.8-Flash-Next, Gemma-4-26B-A4B-it and gpt-oss-120b. Dense models are used only when they fit in a single GPU's VRAM (16 GB). The one dense model here, Qwen3.8-27B, is sized to fit that rule (and gets a further boost from MTP). If "80 GB of VRAM" made you think "big dense models, then?" — sorry, this is an MoE-only machine.

Dense: a model that uses every one of its weights for every token. The classic, most common design.
MoE (Mixture of Experts): a model with many "experts" inside that uses only a few of them for each token. The total parameter count can be large while the amount of work per token stays small.
MTP (Multi-Token Prediction): predicts several tokens ahead instead of just the next one, and the main model checks them all at once. Because the main model checks every token, the output doesn't change.
* gpt-oss-120b has no MTP layers, so it gets an add-on helper model using a different method built on the same idea, called EAGLE3.
Context length is 65,536 tokens, the KV cache is compressed to 4 bits (q4_0), and Flash Attention is on. What the speed-ups bought us is in the tuning log.

05Also

The PC interior from an angle, with the GPUs and memory glowing in rainbow colors
It glows. For the record, no amount of glowing makes the LLMs a single token faster.
An M6 Mac mini and its box
Unboxed it, took a photo, and... that's it.

There's an M6 Mac mini (base model) too. "I'll need it to make iPhone apps someday," I told myself. ...So far it hasn't been used once (lol).

How much of the GPU cost has paid for itself (what if it all ran on the Claude API?) is updated daily on the payback meter in the dev log.

See the dev log →How it works →