How AutoDev LG Works

Who does what, and how a single automatic development attempt (a “run”) unfolds — in two diagrams.

01The cast and their roles

AutoDev LG runs on four layers. The local LLMs at the bottom build the apps. Claude Code grows the system that makes them (the framework), and a human sets the direction.

👤 Human the owner

  • Says what app to build (e.g. “Make me an Invaders game”)
  • Judges only the finished app (the final output), and gives a quick instruction if a feature is missing. Honestly, never even looked at the app spec
  • Decides AutoDev LG's goals, policy and spec changes (when Claude Code asks; e.g. “3 PASSes in a row on INVADERS”, “no app-specific logic in the framework”)

🛠 Claude Code develops the framework (ChatGPT did too, earlier)

  • ① Writes the spec for AutoDev LG from the human's request
  • ② Starts and monitors AutoDev LG, then analyzes the logs to find why it failed
  • ③ Fixes AutoDev LG itself: write the tests first → bump the version → release gate (1,200+ tests) → deploy AutoDev LG → back to ② until app generation succeeds reproducibly
  • ④ Writes handoff notes in case the Claude Code session gets cut off

⚙️ AutoDev LG an automatic development workflow built on LangGraph

It takes a spec and hands out work to the agents below, in order. When an agent gets stuck, it automatically swaps in a different model. At the end of a run, it also writes a user guide to go with the finished app.

Leaderrequirements, design, test plan
Qwen3.8 Flash-Next if stuck → Qwen3.8 27B → Gemma 4 → gpt-oss-120b
Reviewerchecks that the design and tests really mean what the spec says
Gemma 4 26B
Test writerwrites the test code that decides pass or fail
Qwen3.8 27B
Workerwrites and fixes the actual code through OpenHands
Qwen3.8 Flash-Next if stuck → Gemma 4

🖥 Local LLMs running on my home PC with llama.cpp (llama-server)

One model at a time is loaded across five GPUs and swapped whenever the role changes. No cloud AI is used.

RTX 5070 Ti 16GBRTX 5060 Ti 16GBRTX 5060 Ti 16GBRTX 5060 Ti 16GBRTX 5060 Ti 16GB= 80GB VRAM

02The flow of one run

It starts from a spec, and if every test passes, it's a PASS. If something fails along the way, it diagnoses the cause, fixes it according to its type, and verifies again.

workcheck (gate)test run
📄 SpecSPEC (e.g. taskboard / INVADERS)written by Claude Code → see the real ones
RequirementsRequirementsLeader
DesignArchitectLeader
Check the design for gapsContract Gate
freeze the design
Write testsTest GenerateTest writer + Reviewer
Check what the tests meanSemantic Audit
CodingCoding (OpenHands)Worker
Match code against designImplementation Contract Gate
Run the testsQuality: Public + Acceptance tests
all passed
🎉 PASS
failures
Diagnose the causeDiagnose (Public / Acceptance)
A bug in the testsfixture / test bug
Safely repair or regenerate the teststest repair / regenerate
A gap in the designcontract gap
Review the designcontract reconcile / review
A bug in the codeimplementation bug
Fix the codecode repair (coding)
Verify again → continue if it made progressre-validate → validated progress?

Repeats the same failure → switch to another model. Still no progress → the run ends (NOT PASS)

at the end of the run
📦 Output: the app files + 📘 user guideproject files + USAGE_GUIDE.md

The user guide is written by a program, not an LLM: it states only what can be verified from the spec, the design and the generated code (the launch command comes from the real code, so “I did what it says and it didn't start” is unlikely).

* The real workflow has more than 20 steps, including test-plan audits, harness recovery and scope checks. The diagram shows only the main path.

03After a run

If it's NOT PASS, Claude Code analyzes the logs. Then it fixes the framework in a general way — never with a fix tailored to one specific app — bumps the version, and starts the next run. Thanks to this loop, the version number has passed 140.

The daily results are updated every morning in the dev log.