Pick the feature. Switch the safeguards.
Five archetypes cover most AI features people ask us for. Each starts with the accuracy a decent prompt gets on its own. The eight switches are the work between a demo and a feature — each one raises a gauge, and each one costs days. The verdict at the bottom is blunt on purpose.
The feature
Traffic
~6,000 tokens a call for this archetype · blended $3 per million tokens — an order of magnitude, not a quote.
If users ask the same eight questions, a good search box and eight answers beat an assistant.
Safeguards · switch them on and off
Each one is real work. Each one is also the difference between a feature and a demo.
Without validation and a path for wrong answers, 21% of answers shown are confidently wrong in this model. That is the version that gets quietly switched off after launch.
01 · Rules, not a benchmark
How the gauges are built
Each archetype starts with a base accuracy a decent prompt gets. Validation, evals and feedback raise it; a fallback and escalation catch part of what is still wrong. Cost is calls × tokens × a blended price; the cap keeps a bad month from tripling.
02 · Where it comes from
Production, not demos
We shipped a production feature on a language model with output validation, handling of wrong answers and tests of the whole path (a healthcare client), and took an assistant in someone else's product from prototype to usable (a client's SaaS product). The switches are that list.
03 · What it doesn't say
Sometimes the answer is a rule
The bench will tell you when the feature does not need a model. We say the same thing in the first call — AI is often overkill, and a plain rule may be enough.
Same model. Different product.
Four things every AI feature does. On the left, how they look in the demo — honestly good. On the right, what the same thing has to do once it sits in a product with paying users, support tickets and a legal team.
Answer.
Be wrong.
Change.
Cost money.
What the switches look like in real products
The bench is a rule set. The switches on it are the list we actually worked through — on products that are live, with users we do not control.
A production feature on a language model
Not a demo: a feature that runs for real users, with validation of everything the model returns, explicit handling of wrong or empty answers, and tests covering the whole path from input to what the user sees. It is the reason the bench has those three switches at the top.
An assistant taken from prototype to usable
The assistant existed. It was not something people could rely on. Getting it to usable was not a model change — it was grounding, guardrails, feedback, and a definition of what 'useful' meant for that product's users.
When the model suggests things nobody wanted
In two products the model suggested things users did not want and asked questions too broad to help. That is not a prompt bug — it is a missing product definition and a missing test of real scenarios. It gets fixed by product and QA people, not by a bigger model.
When AI is overkill
Classification by three keywords. Extraction from one fixed template. The next step that follows from the last by a rule. We have talked clients out of a model more than once — a rule is cheaper, deterministic, testable, and never hallucinates.
The PII conversation happens before launch or after
Which fields may go to which model is a one-week job before launch and a one-quarter job after the first customer audit. The bench scores it under trust because that is who asks: security, legal, procurement.
Who builds it
The feature has to live in your codebase, your deploy pipeline, your on-call. We join the team, build it there, and leave the eval set and the runbook behind — so it keeps holding after we are gone.
Why 'it worked in the demo' is not a plan.
The bench is a model. The gap between demo and production is measured — by people who tried to skip it.

