Assemble the line. Then pull a block out.
Six stations on the belt, five modules under the floor. Click a block to add or remove it. The gauges are a rule set calibrated on one real week — remove the inspection gate and watch risk, remove the planner and watch first-pass, remove all the agents and watch the line run at human speed.
A person reads the file-level plan and approves it — one card drag in the board. Reject unclear tickets here, not after they are built.
18 tickets a night for about $31, with 68 human minutes. A modeled result; the human minutes are assumptions.
Where the humans go · minutes per ticket
- Planning1.5approving a plan
- Verifying0agent + evidence
- Reviewing1.5reading evidence, merging
- Incidents0.83% escaped defects × 25 min
- Forensics0a folder per ticket
01 · Calibration
One real week, one real night
July 20–27, 2026: 50 tickets, 48 completed, 2 rejected by the factory, 390 engine calls, $145 ≈ $2.90 per ticket. One night, 01:12–05:47: 20 claimed, 18 done, $31, 42 tests green. Those are the anchors.
02 · Rules, not a simulation
How removing a block scores
No planner: −25 pts completion. No isolation: −15 (the fourth parallel ticket). No verifier: +18 pts escaped defects. No inspection gate: +25 risk. Humans without agents do the stage by hand: 20, 12 and 15 minutes. A night has 480 human minutes.
03 · What it doesn't say
Most of this is reliability engineering
The prompts were marginal. The queue, the budgets, the contracts and the artefact trail were the work — and none of it is a product you can buy. That is the part a team builds once, with someone who has done it before.
Same tools. Different factory.
Model A is what nearly everyone has: AI speeds up fragments of an unchanged process. Model B redesigns the process around agents and puts humans where judgement is needed. The difference is not the model you rent. It is who does the waiting.
The flow.
What the human does.
What breaks.
What you can measure.
What makes a factory boring enough to run at night
The six rules the real week was built on. Every one of them is a block in the composer above — try running the line without it.
Deterministic code does the infrastructure
Git, CI, budgets, routing, artefacts — all handled by plain code in the harness. Agents never touch a repository directly. The builder may edit only the files named in the approved plan.
Fail closed, everywhere
Every verdict is a strict JSON contract. If a planner, verifier or reviewer cannot say yes or no unambiguously, the workflow halts and a human looks. Nothing is ever guessed through.
Two human gates, zero auto-merge
Intake: a person approves the plan before anything is built. Inspection: a person merges, with the evidence trail open. Both are minutes per ticket. Neither is optional.
One ticket, one worktree
Isolation per ticket with file-level reservations, and a merge queue that re-verifies every waiting ticket when main moves. Added after semantic conflicts appeared at four parallel tickets.
Different models for different roles
A frontier model plans, a faster model builds, a different model reviews, and the verifier can be a local one. Routing is per role, per domain, per project — cost and independence in one setting.
Economics built in from day one
Per-ticket budgets, a circuit breaker on spend per hour, plan reuse, and full metrics per invocation. The week cost $145 because someone could see what every call cost while it ran.
One week, counted.
The composer's gauges are anchored to these numbers. The experiment was one person, two repositories, one week — run and written up by our AI R&D lead. The full write-up is in our case studies.

