The agent moves. The checks run. People make the calls.
This is the line, drawn as the loop it is. Pick a real starting point from the build, make the calls at the five human stops, then try Merge before the checks finish. In the middle: product definition, requirements, acceptance criteria, code, tests, automation and documentation did not come in phases. They grew in the same changes, all at once. An illustrative walk through the way we worked, not a replay of one recorded change.
One ticket. Your calls.
The agent moves. The checks run. At the human stops, you make the call.
Agent work: Read the ticket, locate the cause and prepare a fix.
Human contribution: Answer behaviour questions and review the final code.
Real task inputs from the team lead; they illustrate the pattern, not timed experiments. UAT means the acceptance environment.
Your walk: 0 human calls · 1 station · 0 returns · not merged yet. Your walk, not a measurement: the project did not count interventions per ticket.
- 1Changes needed → Implement
- 2Deploy or smoke failed → Investigate
- 3QA finding → new task → Investigate
- ✦Grew at once, in the same changes: Product definition · Requirements · Acceptance criteria · Code · Tests · Automation · Documentation.
Team workflow, not GitHub merge gates: at day 37 one rule enforced (every review thread resolved), the rest procedure; the 9 checks became required on day 41. AI review and CI run in parallel; smoke and QA + PO run in parallel. Select a station to read its detail.
Station detail · HUMAN
Define the task
PO, team lead and developer set expected behavior. They could open an existing reference product instead of specifying everything from scratch. Pick a starting point to define this task.
Your call
Define the task
Merge now becomes available as soon as the model has a PR, before developer acceptance or completed checks.
The decision was a developer's. The button press could be the agent's, once authorized.
We did not audit every merge against the procedure; the day-37 snapshot shows what was configured, not what happened on each of the 756.
Switching restarts this input. Not on day 37. From day 41 the repository asked GitHub for exactly this, minus the approval: the 9 checks became required, applied by hand. Whether it was safer is not measured. Developer Accept supplies the model's required approval; resolving a thread does not.
Your walk so far
- HUMAN Define the task
Rules, not measurement
The loop is a model of an illustrative ticket, not a recorded PR. The snapshot facts and counts are measured; the order of checks is ours. Human calls count your actions, including retries, Wait and resolving a thread. Let it proceed skips an answer and does not add a call. Stations count visits, including returns.
Agent and check stations advance automatically. Human stations wait. The failed-build input enters through the failure return; the model never invents a smoke failure. Invoice inspection ends in a report without a PR. Payment can also end with a report. A screenshot requires an explanation. QA follow-up ends the walk.
Wait first completes core checks, integration and load-test tooling; then the 2 high and 3 medium Playwright jobs; then the 1 low job and closes the CodeRabbit thread. Returning to implementation resets checks and developer acceptance. These are model rules, not elapsed times or recorded completion orders.
The day-37 snapshot, as the ship procedure recorded it: no required status checks, no minimum approvals, conversation resolution enforced. The procedure told the agent to wait; GitHub did not refuse the merge once every review thread was marked resolved.
Gate timeline, from our copy of the repository: day 37, 0 required checks; day 41, rulesets with 9 required checks on the integration branch and 2 release gates, applied by hand, approvals 0; day 44, a gate in front of the main branch. Applied by hand means the file is not proof a rule was live for every merge.
The launch total includes 13 merges that only pulled one branch into another, not features. The median means half merged faster, half slower. ClickUp counts use the timestamp a task was marked done, not an independent product-acceptance audit. Counts include the full milestone day.
The snapshot contained 4 applications, 731 test files (including 66 browser E2E and 36 integration), 107 documentation files and 10 repository skills. File counts, not executions, coverage or an AI productivity measure.
| Through | Merges | Median | Within 24 h | Done tasks |
|---|---|---|---|---|
| Day 37 | 664 | 50 min | 96.2% | 135 |
| Day 44 | 756 | 48 min | 95.9% | 181 |
GitHub Docs: required checks, reviews and resolved conversations are separate settings. The switch models the day-41 checks plus one approval the rulesets did not require, without an administrator bypass.
Three roles on the line, none of them optional.
From the team's notes and the repository as recorded on day 37. The agent does the work between decisions. The checks reject what they can. The decisions stay with people.
The agent
Given a one-line task, a bug report, a failed deployment or a screenshot, the agent read the ticket, the code and the environment, found the cause and wrote the change, with its tests and its documentation page in the same change. For tasks like these it needed no file-by-file plan. It also drafted test scenarios for the tester.
The people
Four experienced developers, a trainee, a tester, a product owner, one person on environments, one running the schedule. People wrote the one-line tasks, answered behaviour questions the same day, inspected every result and tested it on the published version. A person accepted each of the 756 changes. The product owner alone put in 313 h over 44 days: a full-time job, not a weekly check-in.
The checks
Lint, build, types, dead code, architecture rules, our own contracts (API spec freshness, error codes, currency units, test-case IDs), integration tests on a throwaway Postgres, six Playwright jobs in the browser, and a short real-provider run after every deploy to a test environment. Feedback for correction, not approval: no AI decided whether a change was acceptable.
Forty-four days, counted.
Counts from our copy of the code history, the task list and our time tracker, as of September 2026. Two windows, never mixed: day 37 was ready, day 44 was live.
Day 0 was not empty. Six weeks start earlier.
Day 0 was a Saturday. 52 h went into that weekend, and by Monday morning the architecture, the first connections to outside services and the documentation were in place, because most of it had been built before. Three kinds of things came in on day 0, none of them a tool you can buy.
Backend pieces
The product shipped as four applications: a web app, an admin console, an API and a background worker. Payments, subscriptions and a credits ledger followed shapes we had built before; so did the throwaway-database integration harness and the deployment that publishes every accepted change to a test version. What had worked stayed. What had hurt was redesigned.
Frontend pieces
The web app and the admin console started from shells and components we had used before, with component tests and the browser test setup carried in, with a stand-in for the paid media provider and the payment provider's test mode. The reference product supplied the rest: developers opened it to see how a screen should behave instead of reading a document about it.
Knowledge and technique
Shared instructions the agents read before the first task: architecture, conventions, how to verify. Ten written procedures for agents, from ship to review, including the review skill that asks the same five scale questions of every change. Scripts to set up, start and reset an environment, usable by people who do not write code. Two short meetings a day. All of it from earlier projects, including our own second brain.
More human involvement, not less.
Two things people assume when they hear six weeks and AI in one sentence, next to what the tracker and the team's notes say.
AI builds it for a fraction.
Six calm weeks.
Three things we will not say.
The numbers above are ours to defend. These three are the edges of them.
What AI saved
2,168 h of work went into 44 days. That number says nothing about what the same product would have cost without agents, because nobody built it that way. We publish the hours; the savings claim we leave to others.
That you can copy the calendar
You can copy the way of working: one-line tasks, two meetings a day, automatic publishing, agents taught each new chore, documentation in the same change. You cannot copy the existing product the developers opened to find expected behaviour, or a team that had built this shape before. Without a picture that clear, expect to write more, and to plan the time for it.
What $60k–$150k buys
A first working version of a SaaS or an internal tool, built by the same factory with the pieces above carried in: six to eight weeks, $60k–$150k depending on complexity. Picture almost any product and it fits in that range. This build was a specific engagement with its own scope and its own invoice, which stays between us and the client.
Custom software is no longer a luxury.
A year ago a first working version of a product was a year of work and a budget that only a funded company could sign. It is now six to eight weeks and $60k–$150k, depending on complexity, a budget a department can own, if the people who know the product give it their days and the factory brings its pieces. The tool you rent because building it was unthinkable is now a six-week question. So is the product you shelved.
What the factory is made of, with the humans at two hard gates, is brief 005. What launch usually starts, thirty weeks of keeping a product alive, is brief 004. This brief stops on day 44 on purpose.

