All Briefs
Mission Brief · 005—·updated

Agents build.
Humans decide.
Somebody builds the factory.

Copilots can speed up coding fragments while the process stays unchanged. DORA's 2025 report puts it plainly: 90% of engineers use AI, more than 80% feel more productive, throughput rises — and delivery stability still falls. A factory is the other model: agents on a line built from deterministic code, with a person approving every plan and merging every PR. We ran one for a week. Assemble the line yourself and see what each block is worth.

50
tickets in one week · 48 completed
$2.90
per ticket · API-equivalent cost
2
human decisions per ticket · plan, merge
Build the line
Composer · interactive

Assemble the line. Then pull a block out.

Six stations on the belt, five modules under the floor. Click a block to add or remove it. The gauges are a rule set calibrated on one real week — remove the inspection gate and watch risk, remove the planner and watch first-pass, remove all the agents and watch the line run at human speed.

The lineclick a station to add or remove it · 2 of 2 human gates on the line
Under the floor · harnessdeterministic code, no reasoning — the part nobody demos
Human gateIntake gate

A person reads the file-level plan and approves it — one card drag in the board. Reject unclear tickets here, not after they are built.

18tickets a nightof 20 claimed
90%completion ratemodeled tickets completed
$1.70per ticket≈ $31 a night
3.8human minutes / ticketspec-and-inspect
10risk · under control

18 tickets a night for about $31, with 68 human minutes. A modeled result; the human minutes are assumptions.

Where the humans go · minutes per ticket

  • Planning1.5approving a plan
  • Verifying0agent + evidence
  • Reviewing1.5reading evidence, merging
  • Incidents0.83% escaped defects × 25 min
  • Forensics0a folder per ticket

01 · Calibration

One real week, one real night

July 20–27, 2026: 50 tickets, 48 completed, 2 rejected by the factory, 390 engine calls, $145 ≈ $2.90 per ticket. One night, 01:12–05:47: 20 claimed, 18 done, $31, 42 tests green. Those are the anchors.

02 · Rules, not a simulation

How removing a block scores

No planner: −25 pts completion. No isolation: −15 (the fourth parallel ticket). No verifier: +18 pts escaped defects. No inspection gate: +25 risk. Humans without agents do the stage by hand: 20, 12 and 15 minutes. A night has 480 human minutes.

03 · What it doesn't say

Most of this is reliability engineering

The prompts were marginal. The queue, the budgets, the contracts and the artefact trail were the work — and none of it is a product you can buy. That is the part a team builds once, with someone who has done it before.

Model A vs Model B

Same tools. Different factory.

Model A is what nearly everyone has: AI speeds up fragments of an unchanged process. Model B redesigns the process around agents and puts humans where judgement is needed. The difference is not the model you rent. It is who does the waiting.

01
Model A vs Model B

The flow.

Model ATicket → developer with copilot → review → QA → release. Coding got faster. Everything after it did not.
Model BPlan approved → build → verify → review → merge. Each step has a contract and a verdict. The human touches two of them. Waiting is what got removed, not people.
Copilot speeds the typing. A factory speeds the ticket.
02
Model A vs Model B

What the human does.

Model AWrites the spec, prompts, reads the generated code, tests it, opens the PR, reviews a colleague's PR, merges. Faster at every step. Busy at every step.
Model BApproves a plan (a card drag). Merges a PR with its evidence folder. Rejects what is unclear before it is built. Spec-and-inspect — closer to a security review than to code review.
The skill moved. Precise inputs and critical audit of outputs.
03
Model A vs Model B

What breaks.

Model AFaros: review time grows 91%, PRs get 154% bigger, bugs per developer rise 9%; no organization-level delivery gain is linked to AI adoption. In a separate METR trial, developers felt 20% faster but took 19% longer.
Model BFour tickets in parallel edit the same file. Two reviewers oscillate. A run hangs at 3am. A frontier model times out mid-plan. All four happened in the real week — and all four were fixed in the harness, not the prompts.
Model A fails quietly. Model B fails loudly, once.
04
Model A vs Model B

What you can measure.

Model ADeveloper sentiment. Lines generated. Seats activated. None of it survives a board question.
Model BCost per ticket. Time per stage. First-pass per role. A folder of artefacts per ticket. Without these you cannot tell adoption theatre from advantage.
If the copilot budget is spent and nothing moved, look in the process.
Design rules · six

What makes a factory boring enough to run at night

The six rules the real week was built on. Every one of them is a block in the composer above — try running the line without it.

01
agents only reason

Deterministic code does the infrastructure

Git, CI, budgets, routing, artefacts — all handled by plain code in the harness. Agents never touch a repository directly. The builder may edit only the files named in the approved plan.

— Pirxey case study · AI software factory · rule 1
02
ambiguous = STOP

Fail closed, everywhere

Every verdict is a strict JSON contract. If a planner, verifier or reviewer cannot say yes or no unambiguously, the workflow halts and a human looks. Nothing is ever guessed through.

— Pirxey case study · AI software factory · rule 2
03
plan approval · PR merge

Two human gates, zero auto-merge

Intake: a person approves the plan before anything is built. Inspection: a person merges, with the evidence trail open. Both are minutes per ticket. Neither is optional.

— Pirxey case study · AI software factory · rule 3
04
+ merge queue, re-verify

One ticket, one worktree

Isolation per ticket with file-level reservations, and a merge queue that re-verifies every waiting ticket when main moves. Added after semantic conflicts appeared at four parallel tickets.

— Pirxey case study · AI software factory · rule 4
05
planner ≠ builder ≠ reviewer

Different models for different roles

A frontier model plans, a faster model builds, a different model reviews, and the verifier can be a local one. Routing is per role, per domain, per project — cost and independence in one setting.

— Pirxey case study · AI software factory · rule 5
06
$/ticket · $/hour breaker

Economics built in from day one

Per-ticket budgets, a circuit breaker on spend per hour, plan reuse, and full metrics per invocation. The week cost $145 because someone could see what every call cost while it ran.

— Pirxey case study · AI software factory · rule 6
Evidence · receipts

One week, counted.

The composer's gauges are anchored to these numbers. The experiment was one person, two repositories, one week — run and written up by our AI R&D lead. The full write-up is in our case studies.

50
tickets processed in a week, 48 completed.
July 20–27, 2026. Two rejected by the factory itself. 390 engine invocations, about 18.6 hours of engine time, $145 at API-equivalent rates — about $2.90 per ticket.
Pirxey case study — AI software factory (July 2026)
18/20
tickets completed in a single night run.
01:12 to 05:47 on July 22: 20 tickets claimed, 18 completed, 2 rejected, 166 engine invocations, $31 — about $1.70 per completed ticket. 42 tests run, all passed, including one correct BLOCKED for 'already implemented'.
Pirxey case study — AI software factory (July 2026)
94/95
planner invocations returned a valid result.
Role invocations without error, timeout or restart: planner 94/95, builder 78/87, reviewer 76/76. These are not ticket completion or PR approval rates. The composer's completion gauge starts at the separate 18/20 night-run result; penalties for removed blocks are assumptions.
Pirxey case study — AI software factory (July 2026)
90% / 30%
use AI at work; almost a third don't trust what it writes.
DORA 2025, nearly 5,000 respondents: 90% use AI, more than 80% feel more productive, 30% report little or no trust in AI-generated code. Adoption correlates with throughput and still with lower stability. Model A in one dataset — and the reason Model B has gates.
Google Cloud — 2025 DORA report
+91%
review time in high-AI teams.
Faros: developers in high-AI teams interact with 47% more PRs daily, with 154% larger PRs and 91% longer review time. AI adoption was not correlated with organization-level delivery gains. An observational study, not a comparison of this composer's configurations.
Faros AI — Lab vs Reality
4
real failures, all fixed in the harness.
Semantic conflicts at four parallel tickets → merge queue with re-verify. Oscillating review loops → prior-round memory and oscillation detection. Zombie runs → hard timeouts and orphan adoption. Budget timeouts killing frontier models → per-model limits, planning 5 → 12 minutes.
Pirxey case study — AI software factory (July 2026)
Mission control standing by

One repeatable workflow. One harness. Built with someone who has run one.

We write custom software, some pieces are ready-made, and we join your team and work alongside it. For a factory that means: map your SDLC, pick the one stream worth automating, design the gates and the harness for your stack, and measure cost, time and first-pass from day one. Two to three weeks. No slide deck.

Pirxey · Aleja Grunwaldzka 472, 80-309 Gdańsk, Poland·130+ engineers · 100+ missions delivered