AutoDev LG's version number is past 150. Here are ten fixes whose effect clearly showed up in the results that followed, oldest first.
01v699/20
The "bump the version on every fix" rule is born
What was going wrong
Until then, fixes went straight into a development line (LG1.10.68) whose version number never changed.
The fix
New rule: every time AutoDev LG is fixed, the version number goes up by one and the code is saved as a ZIP. v69 was the first.
What changed
Every run can always be traced to its version. This rule is why the version number is past 150.
"The start of version-number inflation."
02v769/21
The first PASS in a week
What was going wrong
In loop16 the designer forgot the task-board's title input box, so the tests had to guess at it and failed. A test also had a stand-in function that kept calling itself forever.
The fix
Put an input box in the design example, and added a check that catches tests whose stand-ins call themselves.
What changed
The next run, loop17, PASSed on 9/21 — the first success since 9/14.
"A task board with no input box isn't much of a task board."
03v869/23
The endless re-reading bug
What was going wrong
After writing the correct file, the worker kept re-opening the same three files about 50 times, until it hit its step limit.
The fix
The cause: the conversation was auto-summarized after about 12 steps, and the summary wrote in a mistaken "still to do" list that the worker then followed. The summary now starts after 32 steps instead of 12.
What changed
The worker finishes once the file is written, and the pointless re-reading dropped.
"Took so many notes it forgot what it was doing."
04v1019/24
Admitting the tests can be wrong
What was going wrong
Only 3 of the last 20 runs PASSed (15%). Of the 17 failures, 13 were caused by wrong tests that no correct app could pass — and once written, the tests couldn't be changed.
The fix
A policy change approved by the human: under guards, only the broken tests are rewritten by a different model, checked by yet another model, and then swapped in.
What changed
The next day, 9/25, taskboard PASSed four times (v101, v103, v105, v108).
"If the answer key is wrong, fix the answer key."
05v1149/26
Same answer twice? Change tack right away
What was going wrong
In loop59 the diagnosis repeated word-for-word identical answers; eight repairs over about 90 minutes even broke tests that had been passing. A human stopped it after 2 h 56 min.
The fix
At the human's request: if the same result comes back twice, switch to a fallback immediately.
What changed
Less time wasted going around in circles.
"If you've told the same story twice, change the subject."
06v117・v1189/26
Three PASSes in a row
What was going wrong
The prompt for repairing tests pasted the same tests over and over, swelling to about 69k tokens — more than the model could read (65,536). And the "partial design fix" added in v114 had a bug of its own and had never once succeeded.
The fix
Shrank that prompt from 85 KB to 6.4 KB and fixed the partial-fix bug. v118 also added a way to fill small gaps in a reviewer's answer.
What changed
loop63 PASSed on v117, then loop64 and loop65 on v118: three PASSes in a row on taskboard.
"The fixing mechanism itself needed fixing."
07v1309/27
gpt-oss-120b joins as the last line of defense
What was going wrong
Rewrites and reviews need a model other than the one that wrote the code, and in INVADERS run inv10 there were no such models left.
The fix
At the human's request, gpt-oss-120b (63.4 GB, spread over five GPUs) was added as a third model family. It outputs about twice as fast as Flash-Next.
What changed
73% pass rate on test writing (see the model scorecard). In loop69 and loop70 it got test writing through after the other models could not.
"Keep a heavyweight on the bench."
08v134 → v1409/28
WSL vanished — and came back
What was going wrong
The whole WSL (Linux on Windows) development environment disappeared.
The fix
It was first ported to run natively on Windows (v134), but llama-server was slow on long contexts, so that was dropped. WSL was rebuilt and everything moved back (v140).
What changed
loop67 PASSed on the rebuilt WSL (about 79 minutes). The rebuild worked.
"Moved out, then moved right back in."
09v141〜v1439/28〜9/29
MTP: 79 minutes → 54
What was going wrong
Building taskboard once took well over an hour, which capped how many tries fit in a day.
The fix
Added MTP (multi-token prediction) to each model (v141, v142), then replayed recorded real requests to try every speed setting (v143).
What changed
Building taskboard went from 79 minutes in loop67 to 54 minutes in loop68 (v142).
"Not a single GPU added."
10v144・v1459/29
The fastest PASS: 38 minutes
What was going wrong
loop69 and loop70 were NOT PASS. In loop69 the tests changed a drop-down's value but never ran what selecting it triggers. In loop70 the tests expected ValueError where the design said KeyError.
The fix
UI actions now need a step that runs the app's real handler, such as pressing a button (v144), and a test expecting a different error from the design is judged to be the test's mistake and fixed (v145).
What changed
loop71 PASSed in about 38 minutes — the fastest taskboard build yet.
"Changing the value isn't the same as choosing it."
+v1509/29
Bonus: fix one thing, break another
A check for gaps in the design was added (v150). As a side effect, in INVADERS run inv26 the design lost the part that draws the screen, and Qwen3.8-27B, Gemma-4-26B-A4B-it and gpt-oss-120b all stopped at test writing. The next fix waits until the approach is settled.
"Whack-a-mole, still in progress."
Based on each version's CHANGELOG, the handoff notes and the run records.