Zero server parts. An ordinary home-built PC that also plays games, with five GPUs crammed in using every slot plus riser cables, for 80GB of VRAM. Every local LLM in AutoDev LG runs on this one machine.

| Part | What's in it | Note |
|---|---|---|
| CPU | Intel Core i9-14900KF (24 cores / 32 threads) | Bought for gaming. All 32 threads used from WSL |
| Memory | 128 GB (DDR4, 32 GB × 4) | About half (62 GB) is given to WSL |
| Motherboard | ASRock B760 Pro RS/D4 WiFi | A regular gaming board — server boards were out of reach |
| Case | LIAN LI O11D EVO XL | Swapped in to fit five GPUs. Shockingly huge when it arrived |
| GPU bracket | LIAN LI O11DEXL-1X (vertical GPU bracket, an option made for the O11D EVO XL) | Used to mount the GPUs connected by riser cables |
| Power supply | Thermaltake TOUGHPOWER GT 1200W (ATX 3.1, PS-TPT-1200FNFAGJ-3) | The five GPUs' power limits add up to 1,020 W, but in practice the whole PC stays within 750 W even at full load |
| Storage | NVMe SSD 2 TB × 2 + USB HDD 6 TB | Models live on NVMe (loading from USB is much slower) |
| Network | I-O DATA ET10G-US3C (USB Type-C 10-gigabit LAN adapter, Realtek chip) | With the PCIe slots taken by GPUs, there was nowhere left for a 10G NIC, so it runs over USB — plugged into the front USB 3.2 Gen2 Type-C port, the only 10Gb/s USB port on this board. Link speed: 10 Gbps |

Three cards plug straight into PCIe slots. The two that didn't fit go through riser cables (PCIE2, and M2_2 — an M.2 slot). As a result, every card has a different link width.
| GPU | Slot | VRAM | Power limit | Link (measured) |
|---|---|---|---|---|
| RTX 5070 Ti | PCIE1 | 16 GB | 300 W | x16 direct |
| RTX 5060 Ti ① | PCIE2 | 16 GB | 180 W | x1 riser |
| RTX 5060 Ti ② | PCIE4 | 16 GB | 180 W | x1 direct |
| RTX 5060 Ti ③ | PCIE3 | 16 GB | 180 W | x4 direct |
| RTX 5060 Ti ④ | M2_2 | 16 GB | 180 W | x2 riser (M.2 adapter) |
The four RTX 5060 Ti cards come from MSI (VENTUS 2X), Palit (INFINITY) × 2 and GAINWARD (PYTHON). Zero sense of uniformity.
"Link" is the current lane count reported by the NVIDIA driver. The five power limits add up to 1,020 W.
Which GPU sits in which slot?
Using the motherboard layout and block diagram, we matched the PCI locations Windows reports against each link's measured width. The bar length is the lane count (x1–x16).
Slot names follow the ASRock B760 Pro RS/D4 WiFi manual.
| Layer | What runs | Note |
|---|---|---|
| OS | Windows 11 Pro + WSL2 (Ubuntu 24.04.5 LTS) | A Windows-native version was tried too, but it was slow on long contexts, so it went back to WSL |
| GPU stack | NVIDIA driver 617.14 / CUDA 13.3 | |
| Inference engine | llama.cpp official (build 11151) + Unsloth build (build 11160) | The Unsloth build is for Qwen3.8-Flash-Next's MTP; the other models use the official build |
| Language | Python 3.12 | AutoDev LG itself is written in Python |
| Workflow | LangGraph 1.2 | Runs the steps in order and decides where to go back when something fails |
| AI agent | OpenHands SDK 1.46 | An open-source platform for autonomous AI software-engineering agents. The worker's local LLM edits files and runs commands through it |
One model at a time is loaded across all five GPUs. Every time the role changes the model is swapped, which means reading tens of gigabytes again.
| Model | Type | Role | Quantization · size | Speed-up (MTP etc.) |
|---|---|---|---|---|
| Qwen3.8-Flash-Next | MoE | Leader / worker | AP-Q4_K_XL · 95 GB | MTP head (separate file · Q8_0 · 3.9 GB) |
| Qwen3.8-27B | Dense | Writing test code | UD-Q4_K_M · 16 GB | MTP (built into the model) |
| Gemma-4-26B-A4B-it | MoE (about 4B of 26B active) | Reviewer | BF16 · 48 GB | MTP helper model (Q8_0 · 0.4 GB) |
| gpt-oss-120b | MoE | Backup | MXFP4 · 60 GB | EAGLE3 (Q8_0 · 0.8 GB) * |
Dense: a model that uses every one of its weights for every token. The classic, most common design.
MoE (Mixture of Experts): a model with many "experts" inside that uses only a few of them for each token. The total parameter count can be large while the amount of work per token stays small.
MTP (Multi-Token Prediction): predicts several tokens ahead instead of just the next one, and the main model checks them all at once. Because the main model checks every token, the output doesn't change.
* gpt-oss-120b has no MTP layers, so it gets an add-on helper model using a different method built on the same idea, called EAGLE3.
Context length is 65,536 tokens, the KV cache is compressed to 4 bits (q4_0), and Flash Attention is on. What the speed-ups bought us is in the tuning log.


There's an M6 Mac mini (base model) too. "I'll need it to make iPhone apps someday," I told myself. ...So far it hasn't been used once (lol).
How much of the GPU cost has paid for itself (what if it all ran on the Claude API?) is updated daily on the payback meter in the dev log.
See the dev log →How it works →