All Briefs
Mission Brief · 010—·updated

AI writes the options.
Somebody still has to choose.

Before AI a developer made a few dozen decisions a day. Now the ceiling is gone: every prompt, every accept or reject, every 'good enough, ship it' is one. The model does not make them. A person does — on the strength of what they know, how fast they think, how carefully they read and how clearly they asked. This ledger weighs that person. We hire for those four things and test for two of them: our people sit the CCAT, and 80% of the crew score above the 90th percentile.

20 → 99%
modeled odds a decision is right · profile 1 → 10
1 → 40
modeled decisions built on a wrong one · same hour → production
80%
of our crew above the 90th percentile on the CCAT · cognitive aptitude, tested at hiring
Open the ledger
Ledger · interactive

Set the person. Watch the month.

Four sliders describe who is deciding. One control sets how fast a wrong decision comes to light. The grid is thirty days of decisions — white ones right, red ones wrong, orange ones built on a wrong one before anyone noticed. The balance at the bottom prices the difference between this person and an average one.

The person deciding

6
knows the toolknows the domain, the system and the history
6
needs a meetingdecides between two prompts
6
skims the diffcatches the one line that is wrong
6
'make it better'names the constraint and the definition of done

Decisions a day

60one person, AI in the loop

Every prompt, every accept or reject, every "ship it". Before AI a developer made a few dozen a day. Now the ceiling is gone.

How fast do you find out you were wrong?

Until you find out, the next 4 decisions are built on the wrong one. Fixing it then costs about 40 min.

Thirty days · 1,800 decisions

rightwrongbuilt on a wrong one

86%chance each decision is right

Showing one in 6 decisions. Same inputs always draw the same picture.

Wrong, and what it dragged along

169wrong decisions
676decisions built on them before anyone noticed
47%of the month's decisions have to be revisited

What that costs

225 hof rework in the month
281 hwith an average decider (profile 5 / 10)
56 h savedwhat this person's judgement is worth against an average one, per month

01 · The rule

Quality is a curve, not a switch

The four traits average to a profile. All ones decide right 20% of the time, an average profile about 77%, all tens about 99%. Multiply that by two thousand decisions a month and the gap is a team's month. The Pirxey preset is a translation of a test result into sliders, not a measurement — the CCAT covers speed and reading, not your domain.

02 · The lag

Wrong decisions have children

Until a wrong decision is caught, the next ones are built on it. Same hour: one. Next sprint: fifteen. In production: forty. Cost of the fix grows the same way — the old 1-10-100 rule, with AI's volume behind it.

03 · What it doesn't say

AI is not the decider

The model writes the option. A person picks it, on the strength of what they know, how fast they think, how carefully they read and how clearly they asked. That is what the sliders weigh.

The four things · what a decision needs

To decide well, a person has to already be four things.

None of these is a prompt technique. They are who the person is on the day the question arrives — and AI made the questions arrive faster than ever.

01
the domain, the system, the history

Knowledge — wide, not just deep

The model knows the framework. The person has to know that the client's invoices run on the 25th, that the last migration broke reporting, and that 'customer' means three different tables. A decision made without that context is a guess with confidence.

— Pirxey delivery playbook · 2024–2026
02
between two prompts

Speed of thought

A factory that plans, builds and verifies in minutes gives the person minutes to decide. Slow deciders do not make worse decisions — they make fewer, and the line waits. Speed is part of quality now. We test for it: our people sit the CCAT, a cognitive aptitude test of problem solving and learning speed, and 80% of the crew score above the 90th percentile.

— Pirxey case study · AI software factory · Pirxey hiring, CCAT results
03
+91% review time, +154% larger PRs

Reading with comprehension

The output is fluent, long and confident. The one wrong line looks exactly like the fifty right ones. Reading it properly is the whole job of the inspection gate — and it is the skill that separates 'looks fine' from 'is fine'.

— Faros AI — Lab vs Reality
04
constraint + definition of done

Clarity of expectations

'Make it better' produces a random improvement. 'Keep the API stable, cut p95 under 200 ms, add a test for the midnight case' produces the improvement you meant. The person who can say the second sentence is worth several who can only say the first.

— Pirxey field study · prompt-to-outcome reviews
05
q^n

The compounding

Decisions stack. If each one is right 95% of the time, a chain of twenty holds together 36% of the time. At 99%, 82%. Small differences in the person become large differences in the product — which is what the ledger shows in hours.

— Mission Brief 001 — the same math, applied to people
06
1 → 10 → 100

The lag

The cost of a wrong decision grows with how long it survives — the oldest rule in software economics. AI did not change the rule. It changed the volume flowing through it: more decisions, less time between them, more built on each one.

— Boehm, Software Engineering Economics (1981) · still true
Evidence · receipts

Why we think this is the constraint now.

The ledger is a model. The pressure it models is measured — by other people, on other teams, and by us on our own product.

39pt
gap between how fast developers feel and how fast they are.
METR's trial: developers expected AI to make them 24% faster, believed afterwards it had made them ~20% faster, and were measured 19% slower. Confident decisions, wrong direction — the ledger's red cells, in the wild.
METR (2025) · arXiv 2507.09089
+154%
larger pull requests. Same eyes reading them.
Faros AI: high-AI teams ship PRs 154% larger and spend 91% longer reviewing them. Reading with comprehension went from a nice-to-have to the gating skill of the whole pipeline.
Faros AI — Lab vs Reality
2
human decisions per ticket in a working factory.
Approve the plan, merge the PR — minutes each. Fifty tickets a week means a hundred decisions that nobody else will make, from a person who has to be right nearly every time. That is where the four traits get paid — we counted it in our own factory week.
Pirxey case study — AI software factory (July 2026)
80%
of our crew score above the 90th percentile on the CCAT.
The Criteria Cognitive Aptitude Test measures problem solving, critical thinking and how fast someone learns new information — two of the four sliders. Our people sit it, and eight in ten land in the top decile of the test's norm group. It does not measure domain knowledge or how clearly someone asks, which is why the other two sliders exist.
Pirxey hiring · CCAT results
46%
of our own commits were fixes — decisions revisited.
PirxeyOS, our own company operating system (time tracking, HR, skills matrix, employee records, onboarding), 587 commits: 269 of them fixed something already decided. Most were small. All of them were somebody finding out, later than ideal, that an earlier call was wrong.
Mission Brief 004 — PirxeyOS ship's log
95/5
the part of the product that is all decisions.
The invisible 5% under the waterline — edge cases, permissions, integrations — is not typing. It is knowing what to ask for and recognising when the answer is wrong. This brief is about the person doing that.
Mission Brief 003 — The new Pareto
Mission control standing by

The bottleneck moved upstream. So did the hiring bar.

We write custom software, some pieces are ready-made, and we join your team and work alongside it. The people we bring are the ones the ledger rewards: they know the domain, decide fast, read the whole diff and say exactly what done means. Tested, too: our people sit the CCAT, and 80% of the crew score above the 90th percentile. We validate your team's decisions — we do not replace the team. Free first conversation. No slide deck.

Pirxey · Aleja Grunwaldzka 472, 80-309 Gdańsk, Poland·130+ engineers · 100+ missions delivered