All Briefs
Mission Brief · 018—

A working product
in six weeks.

Picture any SaaS, an existing tool or a new one. A first working version used to be a year and a seven-figure budget. Our factory, agents on the line and people at the decisions, ships one in six to eight weeks for $60k–$150k, depending on complexity. The proof below is one real build: a consumer product with payments and media generation, nine people, 44 days from the first line of code to live. The client's name stays under NDA; the numbers are ours.

44 days
from the first line of code to live; ready on day 37; nine people, 2,168 h of work in that window
6–8
weeks to a first working version of a SaaS or an internal tool, at $60k–$150k depending on complexity: our standing offer
756
changes accepted into the product in 44 days, each by a person; half within 48 minutes of being proposed
Walk the loop
How we build · interactive

The agent moves. The checks run. People make the calls.

This is the line, drawn as the loop it is. Pick a real starting point from the build, make the calls at the five human stops, then try Merge before the checks finish. In the middle: product definition, requirements, acceptance criteria, code, tests, automation and documentation did not come in phases. They grew in the same changes, all at once. An illustrative walk through the way we worked, not a replay of one recorded change.

One ticket. Your calls.

The agent moves. The checks run. At the human stops, you make the call.

Pick a real starting point

Agent work: Read the ticket, locate the cause and prepare a fix.

Human contribution: Answer behaviour questions and review the final code.

Real task inputs from the team lead; they illustrate the pattern, not timed experiments. UAT means the acceptance environment.

Your walk: 0 human calls · 1 station · 0 returns · not merged yet. Your walk, not a measurement: the project did not count interventions per ticket.

The delivery loop123EVERY CHANGEall of itat onceProduct definitionRequirementsAcceptance criteriaCodeTestsAutomationDocumentation
  • 1Changes needed → Implement
  • 2Deploy or smoke failed → Investigate
  • 3QA finding → new task → Investigate
  • ✦Grew at once, in the same changes: Product definition · Requirements · Acceptance criteria · Code · Tests · Automation · Documentation.

Team workflow, not GitHub merge gates: at day 37 one rule enforced (every review thread resolved), the rest procedure; the 9 checks became required on day 41. AI review and CI run in parallel; smoke and QA + PO run in parallel. Select a station to read its detail.

Station detail · HUMAN

Define the task

PO, team lead and developer set expected behavior. They could open an existing reference product instead of specifying everything from scratch. Pick a starting point to define this task.

Your call

Define the task

Merge now becomes available as soon as the model has a PR, before developer acceptance or completed checks.

The decision was a developer's. The button press could be the agent's, once authorized.

We did not audit every merge against the procedure; the day-37 snapshot shows what was configured, not what happened on each of the 756.

Switching restarts this input. Not on day 37. From day 41 the repository asked GitHub for exactly this, minus the approval: the 9 checks became required, applied by hand. Whether it was safer is not measured. Developer Accept supplies the model's required approval; resolving a thread does not.

Your walk so far

  1. HUMAN Define the task
Rules, not measurement

The loop is a model of an illustrative ticket, not a recorded PR. The snapshot facts and counts are measured; the order of checks is ours. Human calls count your actions, including retries, Wait and resolving a thread. Let it proceed skips an answer and does not add a call. Stations count visits, including returns.

Agent and check stations advance automatically. Human stations wait. The failed-build input enters through the failure return; the model never invents a smoke failure. Invoice inspection ends in a report without a PR. Payment can also end with a report. A screenshot requires an explanation. QA follow-up ends the walk.

Wait first completes core checks, integration and load-test tooling; then the 2 high and 3 medium Playwright jobs; then the 1 low job and closes the CodeRabbit thread. Returning to implementation resets checks and developer acceptance. These are model rules, not elapsed times or recorded completion orders.

The day-37 snapshot, as the ship procedure recorded it: no required status checks, no minimum approvals, conversation resolution enforced. The procedure told the agent to wait; GitHub did not refuse the merge once every review thread was marked resolved.

Gate timeline, from our copy of the repository: day 37, 0 required checks; day 41, rulesets with 9 required checks on the integration branch and 2 release gates, applied by hand, approvals 0; day 44, a gate in front of the main branch. Applied by hand means the file is not proof a rule was live for every merge.

The launch total includes 13 merges that only pulled one branch into another, not features. The median means half merged faster, half slower. ClickUp counts use the timestamp a task was marked done, not an independent product-acceptance audit. Counts include the full milestone day.

The snapshot contained 4 applications, 731 test files (including 66 browser E2E and 36 integration), 107 documentation files and 10 repository skills. File counts, not executions, coverage or an AI productivity measure.

Measured activity · August and September 2026
ThroughMergesMedianWithin 24 hDone tasks
Day 3766450 min96.2%135
Day 4475648 min95.9%181

GitHub Docs: required checks, reviews and resolved conversations are separate settings. The switch models the day-41 checks plus one approval the rulesets did not require, without an administrator bypass.

Who does what

Three roles on the line, none of them optional.

From the team's notes and the repository as recorded on day 37. The agent does the work between decisions. The checks reject what they can. The decisions stay with people.

01
diagnosis · code · tests · documentation

The agent

Given a one-line task, a bug report, a failed deployment or a screenshot, the agent read the ticket, the code and the environment, found the cause and wrote the change, with its tests and its documentation page in the same change. For tasks like these it needed no file-by-file plan. It also drafted test scenarios for the tester.

— team's notes · five real task inputs, in the loop above
02
define · answer · inspect · accept

The people

Four experienced developers, a trainee, a tester, a product owner, one person on environments, one running the schedule. People wrote the one-line tasks, answered behaviour questions the same day, inspected every result and tested it on the published version. A person accepted each of the 756 changes. The product owner alone put in 313 h over 44 days: a full-time job, not a weekly check-in.

— team's notes · lead developer's report, approximate
03
9 CI jobs · smoke after deploy

The checks

Lint, build, types, dead code, architecture rules, our own contracts (API spec freshness, error codes, currency units, test-case IDs), integration tests on a throwaway Postgres, six Playwright jobs in the browser, and a short real-provider run after every deploy to a test environment. Feedback for correction, not approval: no AI decided whether a change was acceptable.

— repository as recorded on day 37 · CI workflow
What we carried in

Day 0 was not empty. Six weeks start earlier.

Day 0 was a Saturday. 52 h went into that weekend, and by Monday morning the architecture, the first connections to outside services and the documentation were in place, because most of it had been built before. Three kinds of things came in on day 0, none of them a tool you can buy.

01
API · worker · payments · credits

Backend pieces

The product shipped as four applications: a web app, an admin console, an API and a background worker. Payments, subscriptions and a credits ledger followed shapes we had built before; so did the throwaway-database integration harness and the deployment that publishes every accepted change to a test version. What had worked stayed. What had hurt was redesigned.

— team's notes · repository as recorded on day 37
02
app shell · admin console · browser tests

Frontend pieces

The web app and the admin console started from shells and components we had used before, with component tests and the browser test setup carried in, with a stand-in for the paid media provider and the payment provider's test mode. The reference product supplied the rest: developers opened it to see how a screen should behave instead of reading a document about it.

— team's notes
03
instructions · 10 skills · review rules

Knowledge and technique

Shared instructions the agents read before the first task: architecture, conventions, how to verify. Ten written procedures for agents, from ship to review, including the review skill that asks the same five scale questions of every change. Scripts to set up, start and reset an environment, usable by people who do not write code. Two short meetings a day. All of it from earlier projects, including our own second brain.

— team's notes · repository as recorded on day 37 · brief 015
Human in the loop · heavily

More human involvement, not less.

Two things people assume when they hear six weeks and AI in one sentence, next to what the tracker and the team's notes say.

01
Assumed vs happened

AI builds it for a fraction.

What people assumePoint the agents at the reference product, wait six weeks, pay a fraction of the old bill.
What happened2,168 h of work in 44 days. Our estimate is 2,511 human calls if every change followed the stated workflow: defining, answering, inspecting, accepting. The product owner put in 313 h. Nobody can say what AI saved; what changed is that a person's day went into decisions, not typing.
AI changed how, not whether.
02
Assumed vs happened

Six calm weeks.

What people assumeThe team keeps office hours. The machine does the overtime.
What happened163 h on Saturdays and Sundays, 8% of the total; 5 of the 8 people worked at least one weekend; one day of 99 h across the team. The busiest person averaged about 70 h a week. Decisions came the same day, in two short meetings and a call when something was stuck.
Either you have people who will do that, or you take a longer calendar.
What nobody can claim

Three things we will not say.

The numbers above are ours to defend. These three are the edges of them.

01
no reliable number

What AI saved

2,168 h of work went into 44 days. That number says nothing about what the same product would have cost without agents, because nobody built it that way. We publish the hours; the savings claim we leave to others.

— lead developer's report · our time tracker · our estimate
02
a reference product did a lot

That you can copy the calendar

You can copy the way of working: one-line tasks, two meetings a day, automatic publishing, agents taught each new chore, documentation in the same change. You cannot copy the existing product the developers opened to find expected behaviour, or a team that had built this shape before. Without a picture that clear, expect to write more, and to plan the time for it.

— team's notes
03
our offer, not this bill

What $60k–$150k buys

A first working version of a SaaS or an internal tool, built by the same factory with the pieces above carried in: six to eight weeks, $60k–$150k depending on complexity. Picture almost any product and it fits in that range. This build was a specific engagement with its own scope and its own invoice, which stays between us and the client.

— our pricing, October 2026
Where this leaves us

Custom software is no longer a luxury.

A year ago a first working version of a product was a year of work and a budget that only a funded company could sign. It is now six to eight weeks and $60k–$150k, depending on complexity, a budget a department can own, if the people who know the product give it their days and the factory brings its pieces. The tool you rent because building it was unthinkable is now a six-week question. So is the product you shelved.

What the factory is made of, with the humans at two hard gates, is brief 005. What launch usually starts, thirty weeks of keeping a product alive, is brief 004. This brief stops on day 44 on purpose.

Mission control standing by

Bring the product you have in mind. We will tell you the six weeks.

We write custom software. Some pieces are ready-made. We join your team and work alongside it. Bring the product you have in mind, or the tool you pay for and would rather own. We will tell you what the first working version takes: the people, the weeks, the budget, and what you need to have ready before day 0.

Pirxey · Aleja Grunwaldzka 472, 80-309 Gdańsk, Poland·130+ engineers · 100+ missions delivered