Draw the workflow. Keep the controls in the drawing.
Your process has already passed qualification. Start with an example, choose a pattern and edit the six steps. Try removing Verify: the machine cost barely changes, but the design risk jumps. Every estimate follows the rulebook inside the drawing.
Variable machine cost only; licenses, setup and labor are extra. Changing a pattern resets Reason and the pre-action gate. Your other choices stay on the drawing.
02 / Flow card · design, not a running integration
Invoice intake
Email → parse or OCR → extract fields → accounting draft. No payments.
Rev. A
Model / USD
Per attempted item. Review time is separate.
- Human minutes / item
- 0.5 min
0.5 review + 0 pre-action approval
- Design risk · moderate
- 50 / 100
A rule score, not a failure probability.
- Time to pilot · model
- 3 weeks
3 base · connections assumed available
The control points are named. Now test them on real examples, including the ones that should stop.
Trigger → Retrieve → Reason → Act → Verify → Log
An email or system event arrives. Expect retries and duplicate deliveries.
Parse the document; use OCR for scans. Check that amounts, dates and source references survived.
The model returns data only. It has no tools or authority to write or send.
Model · no toolsFixed application code applies the output after the gate below. The model never calls the destination.
Pre-action gate specified below ↓Check every receipt in code, then sample completed items for human review. This happens after Act.
Input and output references, versions, cost, action receipt and who approved. Apply redaction and retention.
Code validates the schema, business rules and allowed action before execution. Failed checks stop here.
The choice adds handling tasks below; it does not authorize sending data to a model.
Open the rulebook · every number in this drawing
Machine cost = sum of steps
Illustrative USD per attempted item, before retries. The small charges stand in for parsing, checks and storage; they are not vendor prices.
- Reason · Extract fields
- $0.0400
- Retrieve
- $0.0200
- Act
- $0.0020
- Verify · automated checks
- $0.0010
- Log
- $0.0010
- Unrounded total
- $0.0640
Reason: rules $0; LLM $0.03 / $0.04 / $0.06; agent $0.75 / $1.50 / $2.50. Retrieve: fields $0, document $0.02, API $0.002. Act: draft $0, write $0.002, send $0.001. Verify: $0.001 unless off. Log: full $0.001, errors $0.0002, none $0.
Human time and pilot
Review: none or schema = 0 min; sample = 0.5 min averaged over all items; every result = 2 / 3 / 4 min for rules / LLM / agent. Pre-action human approval adds 1 min. Manual intake, exceptions and waiting are excluded.
Pilot = 2 / 3 / 6 weeks by pattern + 3 weeks when Retrieve needs an API that does not exist. Other connections, access and a bounded scope are assumed. These are planning placeholders.
Risk = weighted gaps, capped at 100
- LLM extraction / classification base
- +15
- Trigger can repeat or arrive late
- +5
- External text reaches a model
- +10
- Personal data reaches a model
- +5
- Action changes another system
- +10
- Verification leaves gaps
- +5
- After the cap
- 50 / 100
Base: rules 5, LLM 15, agent 25. Open-ended reasoning adds 0–10; scheduled / event triggers add 3 / 5; external text to a model +10; personal data to a model +5. Write / send adds 10 / 15.
Verify: every 0, sample 5, schema 10, none 45. Log: full 0, errors 10, none 20. An external action without a gate adds 10, or 25 for an agent. An unverified external action has a minimum score of 80. Bands: below 30 lower, 30–69 moderate, 70–100 high.
The score orders design conversations. It is not a measured likelihood, a security audit or a release approval.
03 / Installation notes
Guardrails for this drawing
Trigger: make repeats harmless
Carry the source item ID through the flow. A repeated event or manual retry must refer to the same item.
Retrieve: put a date on context
Keep the source ID, version and timestamp. Stop on missing or stale context instead of filling gaps from memory.
Prompt injection: input is not an instruction
Separate instructions from retrieved text; sanitize rendered content. Validate outputs and restrict tool permissions. Filtering alone is not a boundary.
Personal data: classify before sending
Remove fields the model does not need. Check the approved processor, access and retention rules, including the audit log.
Reason: check the contract before Act
Validate fields, allowed values and business rules in code. Missing or invalid output goes to an exception queue; a valid schema alone does not make a value true.
Act: write once, even after a timeout
Use a deduplication key at the destination. After an uncertain response, look up the outcome before retrying. Define a reversal or correction path.
Verify: inspect the actual result
Check every receipt in code; route exceptions and a representative sample to a person. Sample review finds defects after the action.
Log: leave enough to reconstruct a run
Record input/output references, versions, action receipts, cost and who approved. Redact sensitive data, restrict access and set retention; do not dump everything forever.
Cost: stop the loop before the bill grows
Set limits per item and per day, plus a maximum number of tool calls and retries. Alert a named owner when a limit stops the workflow.
Exceptions: keep a manual way through
Test bad input, a timeout, a duplicate and an unavailable system. Route unresolved items to a named person; give them a stop switch and a replay procedure.
For the model and tool boundary, see the OWASP prompt injection guidance. The score above is our illustrative rule set.
04 / Acceptance card
Measure the whole item.
Before connecting a live action, agree on the pass/fail thresholds and who can stop the pilot.
- Keep a baseline. Time the current workflow on the same class of items. Keep examples with agreed answers, including duplicates and invalid input.
- Count accepted outcomes. Track first-pass acceptance, corrections, human minutes and end-to-end elapsed time. Divide all run costs, including failed runs, by accepted items.
- Start in shadow mode. Compare proposed results with the real process. Then permit a limited batch of live actions, with an owner for exceptions and a stop switch.
Rules, not a forecast
This is a design worksheet. Cost, review time, risk weights and pilot weeks are declared assumptions. They help compare choices; replace them with measurements from your workflow before making a budget.
Where the numbers come from
Our factory week reported about $2.90 in API-equivalent inference per ticket. That is one software experiment, not an invoice price list.
Our work for a healthcare client put an LLM feature into production with response validation, invalid-response handling and tests of the whole path. For another client's SaaS product, we made an AI assistant usable. Those projects inform the control points. The coefficients here are model assumptions.
What it doesn’t say
No forecast of savings, staffing, throughput or error rates. No license, setup or labor cost. Document length, retries, review queues and exception work can dominate the bill. A lower score still needs tests and a responsible owner.
The six steps, and what goes wrong in each
Your team knows the process. The design work is turning that knowledge into contracts, permissions and checks that still hold when a system times out or an input changes.
Trigger: the same item arrives again
A webhook retries, a scheduler catches up, someone clicks twice. Give the item a stable ID and carry it through the flow. The destination must recognize a repeated write. A scheduled report also needs a time zone, a cutoff and a rule for missed runs.
Retrieve: the context is out of date
The account changed yesterday; the workflow reads last week's export. Keep the source and timestamp with the data, define how old is too old and check access at retrieval. Missing context goes to an exception queue. Do not ask the model to fill the gap from memory.
Reason: a plausible value is invented
A model supplies an invoice field that was never on the page. Require source references, allowed values and a way to say 'unknown.' Validate the output contract before Act. A JSON schema checks the shape; labeled examples and business rules check whether the values make sense.
Act: a timeout becomes a duplicate
The accounting system saves the record, but its reply never arrives. Retrying blindly creates another record. Use an idempotency key, look up uncertain outcomes and limit which actions the executor can take. Put approval before the side effect. Plan how to reverse or correct a wrong action.
Verify: the run is green, the result is wrong
The API returns success, but the ticket lands in the wrong queue. Check the saved result, then test its meaning against agreed examples. Route exceptions and the chosen review sample to a person. A review after sending is detection, not permission to send. Account for the correction work.
Log: nobody can reconstruct the mistake
An error-only log misses a wrong result that looked successful. Keep input and output references, model or rule versions, the action receipt, cost and who approved. Make one item traceable across all steps. Redact sensitive fields and define access and retention instead of storing everything forever.
A fast step is useful. An accepted result is the measure.
These studies and one factory run concern software delivery. They motivate measuring your workflow; they do not supply the worksheet's coefficients. The builder's numbers remain explicit model assumptions.
Company-wide readiness is covered in Brief 009: How to approach AI transformation. For the agents, gates and failures of a working software flow, see Brief 005: AI software factory.

