All Briefs
Mission Brief · 011—·updated

The demo worked.
Will the feature hold?

An AI feature is easy to demo and hard to keep. The difference is a short list nobody shows on stage: validating what the model returns, having a path for when it is wrong, testing the whole journey on real cases, capping the bill, and knowing what data may leave the building. Put your feature on the bench and switch that list on and off.

8
safeguards · two days to two weeks each in this model
3×
a runaway month without a cost cap · model assumption
0
models needed when a rule does the job — we say so
Put it on the bench
Test bench · interactive

Pick the feature. Switch the safeguards.

Five archetypes cover most AI features people ask us for. Each starts with the accuracy a decent prompt gets on its own. The eight switches are the work between a demo and a feature — each one raises a gauge, and each one costs days. The verdict at the bottom is blunt on purpose.

The feature

Traffic

2,000feature requests a day

~6,000 tokens a call for this archetype · blended $3 per million tokens — an order of magnitude, not a quote.

Does it need a model at all?

If users ask the same eight questions, a good search box and eight answers beat an assistant.

Safeguards · switch them on and off

Each one is real work. Each one is also the difference between a feature and a demo.

79%right or safely handled21% of answers shown wrong
$1,080a month, at this trafficcapped · uncapped runaway would be $3,456
4.2stypical responseneeds streaming or a loader
38trust score · 0–100what security and legal will ask about
4.4weeks to production3 for the feature, the rest for the safeguards
VerdictIt will hold in the demo

Without validation and a path for wrong answers, 21% of answers shown are confidently wrong in this model. That is the version that gets quietly switched off after launch.

01 · Rules, not a benchmark

How the gauges are built

Each archetype starts with a base accuracy a decent prompt gets. Validation, evals and feedback raise it; a fallback and escalation catch part of what is still wrong. Cost is calls × tokens × a blended price; the cap keeps a bad month from tripling.

02 · Where it comes from

Production, not demos

We shipped a production feature on a language model with output validation, handling of wrong answers and tests of the whole path (a healthcare client), and took an assistant in someone else's product from prototype to usable (a client's SaaS product). The switches are that list.

03 · What it doesn't say

Sometimes the answer is a rule

The bench will tell you when the feature does not need a model. We say the same thing in the first call — AI is often overkill, and a plain rule may be enough.

Demo vs feature

Same model. Different product.

Four things every AI feature does. On the left, how they look in the demo — honestly good. On the right, what the same thing has to do once it sits in a product with paying users, support tickets and a legal team.

01
Demo vs Feature

Answer.

DemoThe model replies fluently to the three questions you rehearsed. Everyone in the room nods.
FeatureEvery reply is checked against a schema before a user sees it. What does not parse is retried, routed to a rule, or handed to a person — with the context attached. Fluent is not the same as right.
The demo shows the best answer. The feature has to survive the worst one.
02
Demo vs Feature

Be wrong.

DemoIt isn't, because you asked the questions it is good at.
FeatureIt will be, a few percent of the time, confidently. The feature knows its own confidence, shows the user what it based the answer on, and escalates below a threshold. Wrong-and-caught is fine. Wrong-and-shown is a churned customer.
Nobody demos the fallback. Everybody needs it.
03
Demo vs Feature

Change.

DemoOne prompt, one model version, one afternoon.
FeatureThe provider updates the model on a Tuesday. An evaluation set of two hundred real cases runs against every prompt and model change and tells you what broke before users do. Without it, the first sign of a regression is a support ticket.
A prompt is not a test. An eval set is.
04
Demo vs Feature

Cost money.

DemoTwenty calls. Nobody looked at the bill.
FeatureTwo thousand calls a day, six thousand tokens each. Budgets per user and per day, a circuit breaker for the loop nobody meant to write, caching for the question everyone asks. Two days of work; the difference between a line item and an incident.
The demo is free. The feature has a monthly invoice — cap it.
Proof · features we shipped

What the switches look like in real products

The bench is a rule set. The switches on it are the list we actually worked through — on products that are live, with users we do not control.

01
validated · handled · tested end to end

A production feature on a language model

Not a demo: a feature that runs for real users, with validation of everything the model returns, explicit handling of wrong or empty answers, and tests covering the whole path from input to what the user sees. It is the reason the bench has those three switches at the top.

— A healthcare client's product · live feature
02
in someone else's product

An assistant taken from prototype to usable

The assistant existed. It was not something people could rely on. Getting it to usable was not a model change — it was grounding, guardrails, feedback, and a definition of what 'useful' meant for that product's users.

— A client's SaaS product · AI assistant
03
product debt, not AI debt

When the model suggests things nobody wanted

In two products the model suggested things users did not want and asked questions too broad to help. That is not a prompt bug — it is a missing product definition and a missing test of real scenarios. It gets fixed by product and QA people, not by a bigger model.

— Elite Medical Prep · a real estate product · qualitative
04
a rule ships in a week

When AI is overkill

Classification by three keywords. Extraction from one fixed template. The next step that follows from the last by a rule. We have talked clients out of a model more than once — a rule is cheaper, deterministic, testable, and never hallucinates.

— Pirxey viability checks · 2025–2026
05
classify · redact · route

The PII conversation happens before launch or after

Which fields may go to which model is a one-week job before launch and a one-quarter job after the first customer audit. The bench scores it under trust because that is who asks: security, legal, procurement.

— Pirxey delivery playbook
06
your team + ours, in your product

Who builds it

The feature has to live in your codebase, your deploy pipeline, your on-call. We join the team, build it there, and leave the eval set and the runbook behind — so it keeps holding after we are gone.

— Pirxey engagement model
Evidence · receipts

Why 'it worked in the demo' is not a plan.

The bench is a model. The gap between demo and production is measured — by people who tried to skip it.

19%
slower — with AI, on real code, in an unchanged process.
METR's early-2025 trial: experienced developers took 19% longer with AI allowed while believing they were faster. It measured development tasks, not the effect of safeguards on AI features. Our recommendation: evaluate the feature on real cases before calling it done.
METR (2025)
+9%
more bugs in high-AI teams.
Faros: bugs per developer up 9%, review time up 91%, no organization-level delivery gain linked to AI adoption. Our recommendation for generated answers is validation and evals; Faros measured software delivery, not answer accuracy.
Faros AI — Lab vs Reality
fail closed
the one rule a working AI system runs on.
In our factory week — 50 tickets, two human gates — every agent verdict was a strict contract and an ambiguous result stopped the line. Anthropic's guidance on agent systems says the same: guardrails, sandboxed testing, and a pause for a person at checkpoints. The bench's 'validate the output' and 'handle wrong answers' switches are that rule, one feature at a time.
Pirxey case study — AI software factory · Anthropic, Building effective agents
46%
of our own commits were fixes — most of them in the invisible layers.
PirxeyOS: data consistency, permissions, integrations, real devices. An AI feature adds a fifth invisible layer — the model's behaviour — and it is the one that changes without a deploy. Hence the eval set.
Mission Brief 004 — PirxeyOS ship's log
95/5
the safeguards are the invisible 5%.
Everything on the bench is under the waterline of the demo. If you want the general rule for why the last stretch costs most of the work, it is the brief that started this series.
Mission Brief 003 — The new Pareto
$3
per million tokens — the order of magnitude behind the cost gauge.
Blended, mid-size model, 2026. Your provider and model will differ; the shape will not: calls × tokens × price, minus caching, times whatever a runaway loop multiplies it by. Cap it.
Bench assumption · order of magnitude
Mission control standing by

We check whether it makes sense, plan it, build it and ship it into your product.

We write custom software, some pieces are ready-made, and we join your team and work alongside it. For an AI feature that means: a viability check first — technical and financial — then a plan, then the build with the safeguards on the bench, wired into your product with tests of the whole path. And we say out loud when AI is overkill and a plain rule will do. Free first pass. No slide deck.

Pirxey · Aleja Grunwaldzka 472, 80-309 Gdańsk, Poland·130+ engineers · 100+ missions delivered