LLM TUNING LOG · 2026-09-29

Same GPUs,
AutoDev LG now builds taskboard
25 minutes faster.

We left AutoDev LG's app-building logic alone and tuned only for speed. Here is what we changed in each local LLM, and how much faster it got.

79→54 min
taskboard run time (loop67 → loop68, −32%)
+56%
Flash-Next generation speed on a 40k-token context
~2.8×
27B generation speed (24.4 → 68.3 tok/s)
33→~29 s
gemma model load time

01The toolbox

Every change here does the same thing: the answers stay the same, they just arrive sooner.

MTP (speculative decoding)
A small "drafter" guesses the next few tokens, and the main model checks them all at once. Because the main model verifies every token, the output doesn't change. When the guesses are right, it jumps ahead.
Draft length (draft_n_max)
How many tokens ahead the drafter guesses. Longer drafts pay off more when they hit, and miss more often.
Batch size (ubatch)
How big a chunk of a long prompt is read in at a time.
Load mode (load_mode)
How the model file is read onto the GPUs. We tried reading it straight in ("none") instead of the default "mmap".
Partial CPU offload (--n-cpu-moe 2)
Part of Flash-Next's first 2 layers moves to main memory, freeing a GPU for its drafter.

02Settings per model

Four models split the work in AutoDev LG. Here's what each one got, and what we tried and dropped.

ModelMain jobAdoptedDropped
Flash-NextLeader / codingMTP (draft 2 → 4), ubatch 256 → 512, partial CPU offload, switched to an MTP-capable llama.cpp build (Unsloth)Thread count, load mode (no gain)
27BWriting test codeMTP (draft 3 → 6), load mode noneubatch 512 (no effect)
gemmaReviews (acceptance tests, spec audits)MTP (draft 4, dedicated drafter model), ubatch 256 → 512, load mode noneOther draft lengths
gpt-ossBackupUnchanged (the model has no MTP layers)—

03Generation speed: adding MTP

The same requests, sent with and without MTP (tok/s = tokens per second; higher is faster).

ModelTestNo MTPWith MTPDraft hit rate
Flash-Next40k-token context20.331.7 (+56%)~0.71
Flash-NextShort prompts28.540–450.76–0.93
27B2 short prompts29.4 / 29.056–710.74 / 0.92
gemma2 short prompts33.6 / 30.652.2 / 62.20.56 / 0.75

With MTP on, every model still returned answers in the required format (JSON). MTP went into 27B and gemma with v141 on 9/28, and into Flash-Next with v142 on 9/29.

04Squeezing harder: replaying real requests

We recorded the actual requests from loop68, the run that passed on the MTP setup, then replayed them while changing one setting at a time (v143).

ModelChangeGeneration tok/sPrompt reading tok/sTotal time
Flash-Nextdraft 2→4, ubatch 256→51241.0 → 49.5247 → 330599 → 464 s
27Bdraft 3→658.3 → 70.6~1205 → ~1175151 → 135 s
gemmaubatch 256→51263.9 → 52.21023 → 136133 → 29 s

gemma actually generated more slowly, but it read its long prompts so much faster that its total time went down, so the change was kept.

The real runs (loop68 → loop69) showed almost the same gains.

ModelMetricloop68 (v142)loop69 (v143)
Flash-NextGeneration40.9 tok/s49.8 tok/s
Flash-NextPrompt reading281 tok/s366 tok/s
27BGeneration57.0 tok/s68.3 tok/s

05Model load time

AutoDev LG swaps models as roles change, and every swap means reading tens of gigabytes. In loop68, 187 seconds of the roughly 54 minutes went to loading models. The big models don't fit in the PC's memory cache, so disk read speed matters.

ModelFile sizeDefault (mmap)noneResult
gemma50 GB58.9 / 35.0 s28.0 / 29.1 sAdopted
27B—7.1 / 22.1 s6.1 / 6.1 sAdopted
Flash-Next95 GB54.0 / 49.3 s52.9 / 62.1 sDropped (no difference)
gpt-oss60 GBUnchanged

Each setting was measured twice, taking turns. In the real loop69 run, gemma went from 33 → 27–31 s and 27B from 8 → 5–7 s.

06Wrap-up, and what came next

On the first MTP setup (v142), taskboard PASSED in 54 minutes instead of 79. The v143 tuning hit its speed targets too.

loop69 on v143 and the next run, loop70, were still NOT PASS, though. Both tripped over how tests were written, which had nothing to do with speed. After v144 and v145 fixed that, loop71 ran on the same v143 speed settings and PASSED in about 38 minutes (2261 s), the fastest taskboard run yet! That run also needed fewer model requests (52, vs 81 in loop68) and spent 180 s loading models.

Faster runs mean more tries per day, and that's the real win here.

Read the dev log → How it works →