We left AutoDev LG's app-building logic alone and tuned only for speed. Here is what we changed in each local LLM, and how much faster it got.
Every change here does the same thing: the answers stay the same, they just arrive sooner.
draft_n_max)ubatch)load_mode)--n-cpu-moe 2)Four models split the work in AutoDev LG. Here's what each one got, and what we tried and dropped.
| Model | Main job | Adopted | Dropped |
|---|---|---|---|
| Flash-Next | Leader / coding | MTP (draft 2 → 4), ubatch 256 → 512, partial CPU offload, switched to an MTP-capable llama.cpp build (Unsloth) | Thread count, load mode (no gain) |
| 27B | Writing test code | MTP (draft 3 → 6), load mode none | ubatch 512 (no effect) |
| gemma | Reviews (acceptance tests, spec audits) | MTP (draft 4, dedicated drafter model), ubatch 256 → 512, load mode none | Other draft lengths |
| gpt-oss | Backup | Unchanged (the model has no MTP layers) | — |
The same requests, sent with and without MTP (tok/s = tokens per second; higher is faster).
| Model | Test | No MTP | With MTP | Draft hit rate |
|---|---|---|---|---|
| Flash-Next | 40k-token context | 20.3 | 31.7 (+56%) | ~0.71 |
| Flash-Next | Short prompts | 28.5 | 40–45 | 0.76–0.93 |
| 27B | 2 short prompts | 29.4 / 29.0 | 56–71 | 0.74 / 0.92 |
| gemma | 2 short prompts | 33.6 / 30.6 | 52.2 / 62.2 | 0.56 / 0.75 |
With MTP on, every model still returned answers in the required format (JSON). MTP went into 27B and gemma with v141 on 9/28, and into Flash-Next with v142 on 9/29.
We recorded the actual requests from loop68, the run that passed on the MTP setup, then replayed them while changing one setting at a time (v143).
| Model | Change | Generation tok/s | Prompt reading tok/s | Total time |
|---|---|---|---|---|
| Flash-Next | draft 2→4, ubatch 256→512 | 41.0 → 49.5 | 247 → 330 | 599 → 464 s |
| 27B | draft 3→6 | 58.3 → 70.6 | ~1205 → ~1175 | 151 → 135 s |
| gemma | ubatch 256→512 | 63.9 → 52.2 | 1023 → 1361 | 33 → 29 s |
gemma actually generated more slowly, but it read its long prompts so much faster that its total time went down, so the change was kept.
The real runs (loop68 → loop69) showed almost the same gains.
| Model | Metric | loop68 (v142) | loop69 (v143) |
|---|---|---|---|
| Flash-Next | Generation | 40.9 tok/s | 49.8 tok/s |
| Flash-Next | Prompt reading | 281 tok/s | 366 tok/s |
| 27B | Generation | 57.0 tok/s | 68.3 tok/s |
AutoDev LG swaps models as roles change, and every swap means reading tens of gigabytes. In loop68, 187 seconds of the roughly 54 minutes went to loading models. The big models don't fit in the PC's memory cache, so disk read speed matters.
| Model | File size | Default (mmap) | none | Result |
|---|---|---|---|---|
| gemma | 50 GB | 58.9 / 35.0 s | 28.0 / 29.1 s | Adopted |
| 27B | — | 7.1 / 22.1 s | 6.1 / 6.1 s | Adopted |
| Flash-Next | 95 GB | 54.0 / 49.3 s | 52.9 / 62.1 s | Dropped (no difference) |
| gpt-oss | 60 GB | Unchanged | ||
Each setting was measured twice, taking turns. In the real loop69 run, gemma went from 33 → 27–31 s and 27B from 8 → 5–7 s.
On the first MTP setup (v142), taskboard PASSED in 54 minutes instead of 79. The v143 tuning hit its speed targets too.
loop69 on v143 and the next run, loop70, were still NOT PASS, though. Both tripped over how tests were written, which had nothing to do with speed. After v144 and v145 fixed that, loop71 ran on the same v143 speed settings and PASSED in about 38 minutes (2261 s), the fastest taskboard run yet! That run also needed fewer model requests (52, vs 81 in loop68) and spent 180 s loading models.
Faster runs mean more tries per day, and that's the real win here.
Read the dev log → How it works →