All Briefs
Mission Brief · 008—·updated

You know what.
Now design how.

The process is already chosen: invoice intake, support triage, a weekly report. Now someone has to decide what reads the input, what may change a system, what checks the result and what happens when a step fails. Draw that workflow here. Then take the blueprint into your pilot.

3
patterns in the worksheet · rules, LLM, agent
6
steps on the drawing · trigger to audit trail
1
workflow · measured from input to accepted result
Draw the workflow
Design one workflow · interactive

Draw the workflow. Keep the controls in the drawing.

Your process has already passed qualification. Start with an example, choose a pattern and edit the six steps. Try removing Verify: the machine cost barely changes, but the design risk jumps. Every estimate follows the rulebook inside the drawing.

Drawing 008 · one workflow
Start with an example
01 / Choose how the work gets done

Variable machine cost only; licenses, setup and labor are extra. Changing a pattern resets Reason and the pre-action gate. Your other choices stay on the drawing.

02 / Flow card · design, not a running integration

Invoice intake

Email → parse or OCR → extract fields → accounting draft. No payments.

Blueprint
Rev. A
Model / USD
~$0.06 / item0.5 human min50/100 risk3 weeks to pilotModel
Modeled machine cost / item
~$0.06

Per attempted item. Review time is separate.

Human minutes / item
0.5 min

0.5 review + 0 pre-action approval

Design risk · moderate
50 / 100

A rule score, not a failure probability.

Time to pilot · model
3 weeks

3 base · connections assumed available

The control points are named. Now test them on real examples, including the ones that should stop.

Trigger → Retrieve → Reason → Act → Verify → Log

  1. An email or system event arrives. Expect retries and duplicate deliveries.

  2. Parse the document; use OCR for scans. Check that amounts, dates and source references survived.

  3. The model returns data only. It has no tools or authority to write or send.

    Model · no tools
  4. Fixed application code applies the output after the gate below. The model never calls the destination.

    Pre-action gate specified below ↓
  5. Check every receipt in code, then sample completed items for human review. This happens after Act.

  6. Input and output references, versions, cost, action receipt and who approved. Apply redaction and retention.

Code validates the schema, business rules and allowed action before execution. Failed checks stop here.

The choice adds handling tasks below; it does not authorize sending data to a model.

Open the rulebook · every number in this drawing

Machine cost = sum of steps

Illustrative USD per attempted item, before retries. The small charges stand in for parsing, checks and storage; they are not vendor prices.

Reason · Extract fields
$0.0400
Retrieve
$0.0200
Act
$0.0020
Verify · automated checks
$0.0010
Log
$0.0010
Unrounded total
$0.0640

Reason: rules $0; LLM $0.03 / $0.04 / $0.06; agent $0.75 / $1.50 / $2.50. Retrieve: fields $0, document $0.02, API $0.002. Act: draft $0, write $0.002, send $0.001. Verify: $0.001 unless off. Log: full $0.001, errors $0.0002, none $0.

Human time and pilot

Review: none or schema = 0 min; sample = 0.5 min averaged over all items; every result = 2 / 3 / 4 min for rules / LLM / agent. Pre-action human approval adds 1 min. Manual intake, exceptions and waiting are excluded.

Pilot = 2 / 3 / 6 weeks by pattern + 3 weeks when Retrieve needs an API that does not exist. Other connections, access and a bounded scope are assumed. These are planning placeholders.

Risk = weighted gaps, capped at 100

LLM extraction / classification base
+15
Trigger can repeat or arrive late
+5
External text reaches a model
+10
Personal data reaches a model
+5
Action changes another system
+10
Verification leaves gaps
+5
After the cap
50 / 100

Base: rules 5, LLM 15, agent 25. Open-ended reasoning adds 0–10; scheduled / event triggers add 3 / 5; external text to a model +10; personal data to a model +5. Write / send adds 10 / 15.

Verify: every 0, sample 5, schema 10, none 45. Log: full 0, errors 10, none 20. An external action without a gate adds 10, or 25 for an agent. An unverified external action has a minimum score of 80. Bands: below 30 lower, 30–69 moderate, 70–100 high.

The score orders design conversations. It is not a measured likelihood, a security audit or a release approval.

03 / Installation notes

Guardrails for this drawing

To implement and test · selecting them is not implementation
  • Trigger: make repeats harmless

    Carry the source item ID through the flow. A repeated event or manual retry must refer to the same item.

  • Retrieve: put a date on context

    Keep the source ID, version and timestamp. Stop on missing or stale context instead of filling gaps from memory.

  • Prompt injection: input is not an instruction

    Separate instructions from retrieved text; sanitize rendered content. Validate outputs and restrict tool permissions. Filtering alone is not a boundary.

  • Personal data: classify before sending

    Remove fields the model does not need. Check the approved processor, access and retention rules, including the audit log.

  • Reason: check the contract before Act

    Validate fields, allowed values and business rules in code. Missing or invalid output goes to an exception queue; a valid schema alone does not make a value true.

  • Act: write once, even after a timeout

    Use a deduplication key at the destination. After an uncertain response, look up the outcome before retrying. Define a reversal or correction path.

  • Verify: inspect the actual result

    Check every receipt in code; route exceptions and a representative sample to a person. Sample review finds defects after the action.

  • Log: leave enough to reconstruct a run

    Record input/output references, versions, action receipts, cost and who approved. Redact sensitive data, restrict access and set retention; do not dump everything forever.

  • Cost: stop the loop before the bill grows

    Set limits per item and per day, plus a maximum number of tool calls and retries. Alert a named owner when a limit stops the workflow.

  • Exceptions: keep a manual way through

    Test bad input, a timeout, a duplicate and an unavailable system. Route unresolved items to a named person; give them a stop switch and a replay procedure.

For the model and tool boundary, see the OWASP prompt injection guidance. The score above is our illustrative rule set.

04 / Acceptance card

Measure the whole item.

Before connecting a live action, agree on the pass/fail thresholds and who can stop the pilot.

  1. Keep a baseline. Time the current workflow on the same class of items. Keep examples with agreed answers, including duplicates and invalid input.
  2. Count accepted outcomes. Track first-pass acceptance, corrections, human minutes and end-to-end elapsed time. Divide all run costs, including failed runs, by accepted items.
  3. Start in shadow mode. Compare proposed results with the real process. Then permit a limited batch of live actions, with an owner for exceptions and a stop switch.
Calibration / 01

Rules, not a forecast

This is a design worksheet. Cost, review time, risk weights and pilot weeks are declared assumptions. They help compare choices; replace them with measurements from your workflow before making a budget.

Calibration / 02

Where the numbers come from

Our factory week reported about $2.90 in API-equivalent inference per ticket. That is one software experiment, not an invoice price list.

Our work for a healthcare client put an LLM feature into production with response validation, invalid-response handling and tests of the whole path. For another client's SaaS product, we made an AI assistant usable. Those projects inform the control points. The coefficients here are model assumptions.

Calibration / 03

What it doesn’t say

No forecast of savings, staffing, throughput or error rates. No license, setup or labor cost. Document length, retries, review queues and exception work can dominate the bill. A lower score still needs tests and a responsible owner.

Six steps · six failure points

The six steps, and what goes wrong in each

Your team knows the process. The design work is turning that knowledge into contracts, permissions and checks that still hold when a system times out or an input changes.

01
a retry is normal

Trigger: the same item arrives again

A webhook retries, a scheduler catches up, someone clicks twice. Give the item a stable ID and carry it through the flow. The destination must recognize a repeated write. A scheduled report also needs a time zone, a cutoff and a rule for missed runs.

— Design rule · identify the item before processing it
02
source · version · time

Retrieve: the context is out of date

The account changed yesterday; the workflow reads last week's export. Keep the source and timestamp with the data, define how old is too old and check access at retrieval. Missing context goes to an exception queue. Do not ask the model to fill the gap from memory.

— Design rule · stale input can produce a plausible wrong answer
03
unknown stays unknown

Reason: a plausible value is invented

A model supplies an invoice field that was never on the page. Require source references, allowed values and a way to say 'unknown.' Validate the output contract before Act. A JSON schema checks the shape; labeled examples and business rules check whether the values make sense.

— Design rule · a model output is an input to your application
04
check before retrying

Act: a timeout becomes a duplicate

The accounting system saves the record, but its reply never arrives. Retrying blindly creates another record. Use an idempotency key, look up uncertain outcomes and limit which actions the executor can take. Put approval before the side effect. Plan how to reverse or correct a wrong action.

— Design rule · authority belongs in the executor
05
a receipt is a starting point

Verify: the run is green, the result is wrong

The API returns success, but the ticket lands in the wrong queue. Check the saved result, then test its meaning against agreed examples. Route exceptions and the chosen review sample to a person. A review after sending is detection, not permission to send. Account for the correction work.

— Design rule · no verification means defects can stay silent
06
trace the whole item

Log: nobody can reconstruct the mistake

An error-only log misses a wrong result that looked successful. Keep input and output references, model or rule versions, the action receipt, cost and who approved. Make one item traceable across all steps. Redact sensitive fields and define access and retention instead of storing everything forever.

— Design rule · a usable audit trail is part of the workflow
Evidence · why measure the whole flow

A fast step is useful. An accepted result is the measure.

These studies and one factory run concern software delivery. They motivate measuring your workflow; they do not supply the worksheet's coefficients. The builder's numbers remain explicit model assumptions.

Company-wide readiness is covered in Brief 009: How to approach AI transformation. For the agents, gates and failures of a working software flow, see Brief 005: AI software factory.

Mission control standing by

Your process. Designed, built and shipped.

We check whether it makes sense, plan it, build it and ship it into your product — and we say out loud when AI is overkill and a plain rule will do. We write custom software. Some pieces are ready-made. We join your team and work alongside it.

Pirxey · Aleja Grunwaldzka 472, 80-309 Gdańsk, Poland·130+ engineers · 100+ missions delivered