In AutoDev LG, several local LLMs split the work. Who works the most, what are they good at, what do they struggle with? Counted from the records of every run.
Turns: times a model was loaded to handle one job. Test-writing pass rate: share of test-writing turns that didn't end in "couldn't write the tests" (retries count separately). PASSes closed: times it was the coder when the run PASSed.
The hard-working captain. Shows up almost everywhere, from requirements and design to coding. Of 18 PASSes, it closed 17. 141.7 hours of work — nobody else comes close.
The review machine. 231 checks done, and it appeared in more runs than any other model. Ask it to write tests, though, and it passes only 39% of the time.
The test-writing craftsman: 81% pass rate. Short, sharp turns — only 3.3 hours of work in total.
The last line of defense — the backup called in when the others get stuck. 73% pass rate at writing tests.
The retired previous test-writer. Handed over to Qwen3.8-27B after 9/24. 68% pass rate.
Worked as the coder just 4 times early on, then retired. But it's the model that landed the very first PASS on 9/14.
Counted from every run's per-turn records (escalation-chain-run.json). Hours worked include model loading time.