REPLAY · loop71

Behind the 38 minutes
the fastest PASS, taken apart

The fastest run so far (loop71, v146) built taskboard in about 38 minutes. Straight from the run records, here's what the local LLMs were doing between receiving the spec and the PASS.

37.7 min
from start to PASS
52
calls to the LLMs
1.08M
tokens read (input)
~$3.2
if it ran on Claude Opus 5.5

Where the time went

1
3
6

Width is time. Color is the model (blue = Qwen3.8-Flash-Next, green = Gemma-4-26B-A4B-it, yellow = Qwen3.8-27B).

What happened

"+N min" is time since the start. The seconds on the right are how long the LLM took to answer.

+0.0 min
1. DesignQwen3.8-Flash-Next19.1 min
  • Turn the spec into requirements113 s
  • Design (1st try)212 s
  • Check the design against the spec → gaps found159 s
  • Re-check the gaps it found67 s
  • Try a small design patch15 s
  • Redo the design (2nd try)188 s
  • Check the new design → OK159 s
  • Write the test plan170 s
+19.1 min
2. Test-plan reviewGemma-4-26B-A4B-it0.7 min
  • Check the test plan against the spec6 s
+19.8 min
3. Test writingQwen3.8-27B2.3 min
  • Write the pass/fail tests (Acceptance)87 s
  • Repair the tests once39 s
+22.1 min
4. Test reviewGemma-4-26B-A4B-it0.8 min
  • Check what the tests mean12 s
+22.9 min
5. Wrap-up checkQwen3.8-27B0.2 min
  • Take the review result → all good
+23.1 min
6. CodingQwen3.8-Flash-Next14.6 min
  • OpenHands:Write its own tests (Public) (10 calls, 93% cache hits)209 s
  • OpenHands:Part 1: task_model.py (8 calls, 85% cache hits)97 s
  • OpenHands:Part 2: task_store.py (6 calls, 87% cache hits)89 s
  • OpenHands:Part 3: task_service.py (7 calls, 89% cache hits)70 s
  • OpenHands:Part 4: app.py (the UI) (8 calls, 90% cache hits)92 s
  • Final check: does the app match the spec?246 s
  • Ran 38 tests → 🎉 PASS

Highlights

All numbers come from loop71's run records (escalation-chain-run.json, llm_calls.jsonl, token-summary.json). The Claude API estimate uses Opus 5.5 list prices with the measured cache hits.