Give the agent a smaller place to look.
Start with the Pirxey preset. Change the task, then turn typed contracts off: the reading set spreads across the map. Compare the shapes and run a seeded PR fault through the checks. Every number is an explicit teaching model, not a scan of your repository.
Context to load / model
~20k tokens
Modular monolith with contracts
Eight feature modules expose ports. An accounts change can stay in accounts; callers use its public contract.
The spotlight covers 1 of 8 modules.Ports keep this task local. A pricing or payment change still needs its callers checked; types do not prove business behavior.
- Blast radius
- 1 / 8
13% of the map’s modules. Potential impact, including unchanged callers.
- Verifiability
- 90 / 100
Support for checking a change. Not a probability of correctness.
- Agents in parallel
- 6 max
Builders on disjoint task scopes. Shared files still serialize work.
- Review per change
- 9 min
Modeled reading and inspection effort. Human review stays.
- Practice setup
- 4.5 weeks
Selected practices only. Rebuilding the architecture is extra.
Read the sets and the rules behind these numbers
Reading set and potential impact
- Read
- Accounts
- Impact
- Accounts
- Base reading pack
- 6 files
- After repo docs
- 6 files
The three tasks have explicit reading and impact sets for each shape. Without typed contracts, read two extra dependency hops and widen potential impact by one hop, counting callers and dependencies. This is an assumption about missing boundaries, not a static analysis.
Sum the selected parcels’ base files. ADRs and repo docs reduce assumed searching by one sixth: round up to a whole file. Tokens ≈ files × 3,333, rounded to the nearest thousand; conversation and tool overhead are excluded.
Blast radius counts modeled modules, not code volume. Counts across different shapes are not comparable on their own. A regression can travel farther than the selected impact set.
Checks, parallel work and review
Verifiability base: mud 5, layered 10, modular 15, services 10. Add 20 for contracts, 30 for module tests, 25 for CI, 5 for docs and 5 for observability; cap at 100. Security has its own flight-check gate and setup cost.
Parallel builders = 1 unless contracts, module tests and CI are on. Then use floor(module count / impacted modules), capped at 1 for mud and shared layers, 6 for modular and services. This assumes unrelated tasks, isolated worktrees, approved file scopes and a merge queue.
Review minutes = max(5, ceil(12 + 1.2 × files + 20 × impact share + network allowance − practice credits)). Services add 6 minutes. Credits: contracts 4, tests 6, CI 3, docs 2, observability 2.
Setup = sum of the enabled practice allowances: contracts 1, tests 2, docs 0.5, CI 0.5, security 1, observability 1 week. The chosen structure already exists. Migration, service operations, staffing and calendar scheduling are outside this model.
PR flight check / from a scoped plan to a held merge
Give the gates something to catch.
The planner names the allowed files; a person approves that scope. Follow the resulting PR through its checks.
A response field changed shape while the caller still expects the previous version.
Illustrative trace of a seeded example. No real build, scan or review runs here.
Build
In the planMissing imports, invalid references and build-time incompatibilities.
What this looks like
Compile the exact commit under review, including generated API clients.
Quality check
In the planLint and types, plus the module and contract tests you enabled above.
Module tests: included · Contract tests: included
What this looks like
PirxeyOS's ‘version skew’ fix suggests an old-client/new-response test. ‘Scope member filter by access’ suggests a negative authorization test.
Security check
In the planConfigured SAST, secret and dependency checks. Add adversarial tests when the product path uses an LLM.
What this looks like
The secret example matches a configured scanner rule. Business access rules still need explicit tests; a scanner does not know who may see a member.
Agent review
In the planA different model challenges the plan, diff and evidence. It can still miss a defect.
What this looks like
Reviewer gets the approved scope and verifier artifacts. The builder's own confidence is not an acceptance test.
Human review
In the planBehavior, scope, migration risk and whether the evidence supports the claim.
What this looks like
Plan approval happened before building. Here a person inspects the current commit and signs the merge: the second human gate.
Merge
Held for sign-offStale evidence after the target branch moves, and conflicts between isolated changes.
What this looks like
Use a merge queue and reverify on the new base. This walkthrough holds at human review; it never approves or merges anything.
Choose a seeded fault, run the walkthrough, then remove its check and replay.
Fix examples come from the PirxeyOS log. The proposed checks are our interpretation, not a record of which check caught the original bug. Role orchestration is covered in Brief 005.
Rules, not a benchmark
The map is an invented application with declared task scopes. The file counts reflect rough experience from our rescue projects, not a measured sample. Use the comparison to ask what a change needs to read, touch and prove in your repository.
Where the numbers come from
Six versus sixty files is a teaching example. Faros and METR motivate measuring review and the whole task; neither benchmarks these architectures.
In our factory week, the builder may edit only planned files and the verifier uses a fresh checkout. Clear scopes make that restriction practical.
What it doesn’t say
Microservices are not a shortcut to a small context. Shared databases, lockfiles, migrations and semantic dependencies can defeat a drawn boundary. Tests may miss cases; a different reviewer model may share the builder’s mistake. Verify the real repository and the current commit.
A cake, or cupcakes?
Two drawings from a deck we have shown clients for years. A wedding cake looks impressive, but five people around one stacked thing get in each other's way, and one bad tier spoils everything above it. A cupcake tower is less dramatic and far easier to build in parallel, taste piece by piece, and fix. In the age of agents the same drawing says something new: give each agent one cupcake — a module with a boundary, a contract and a context it can actually hold. The tray they all share is the architect's job, and it matters more now, not less.
- Blast radius
- —
- Can keep working
- 5 of 5 agents
- One agent's context must hold
- the whole repository
- Blast radius
- —
- Can keep working
- 5 of 5 agents
- One agent's context must hold
- one module + its contract
03 · How the pieces talk
Choreography over orchestration
Left: every module publishes what happened to a bus and reacts to what it cares about; nobody is in charge. Right: a conductor calls every module, and modules call each other. Take one module down.
- Still serving
- 10 of 10 modules
- Who has to know
- —
- Still serving
- 10 of 10 modules
- Who has to know
- —
In the age of agents the drawing gets a second reader. One agent per cupcake is the natural shape of agentic work: a module with a clear boundary, a contract it must keep, and a context window that only has to hold that. The tray — the shared contract between pieces — is where the architects earn their keep.
One bad tier, whole cake
A monolith without internal boundaries is the wedding cake: every change touches the whole thing, every worker needs the whole recipe, and a mistake in one layer is a mistake in the product. It is also the shape an agent handles worst — its context has to hold everything to change anything.
One agent, one cupcake
A module with an explicit boundary and a typed contract is a job an agent can be given whole: the files it may touch, the promise it must keep, the tests that will reject it. Five agents on five cupcakes do not collide. Five agents on one cake do — the map above shows how fast.
Choreography over orchestration
We have preferred modules that react to events over a conductor that calls everyone since our first deck. A conductor is a single place that has to know everything — and a single place that stops when anything does. With agents building the modules, a bus with clear event contracts is also the cheapest way to keep their work independent.
It is easy to build a primitive application fast
The old deck's other warning still holds, now with faster typing: initial cost is small next to the cost of change. Cloud bills can differ tenfold depending on who set them up, and maintenance is where low quality and wrong boundaries get paid for. Agents make the first version cheaper. They do not make a bad shape cheaper to live with.
Six patterns that pay twice with agents
The point is to make a change understandable and testable in its own scope. Start with a bounded feature in the first week; the work below is a suggested pilot sequence, not a promise to rebuild an application in that time.
Hexagonal ports and adapters
A port describes a capability the application needs or offers. REST and persistence adapters implement the conversation with the outside world. The agent can change a rule without learning every database detail. First week: take one endpoint, move its rule into a use case, define its outbound port and test it with an in-memory adapter. Keep framework types outside the domain.
Typed contracts, checked at the boundary
A shared schema or OpenAPI definition gives both sides the same vocabulary. Generated types narrow what the agent has to infer; runtime validation checks what actually arrives. First week: define one response schema, generate the client types and add fixtures for a missing field and an older client. A successful compile does not establish compatibility with a deployed consumer.
Feature modules with an owner
A bounded context keeps a business rule and its language together. A feature module gives that boundary a home in code. The agent can receive an allowed file list instead of discovering ownership during the edit. First week: put one feature's UI, validation and client calls together, expose its public API and add an import-boundary check. Separate deployment is an additional decision.
Tests as executable specifications
A test states an observable promise. It gives the builder a target and the verifier an independent way to reject the result. First week: write the pricing edge case, the request from a user without access and the failed dependency response. Run the relevant module tests in CI and an e2e test through the actual user path. Assert the saved result, not just a success message.
Idempotent operations and feature flags
An idempotency key makes a repeated operation recognizable. A feature flag lets the team control exposure while it checks real behavior. They give an agent concrete failure semantics to preserve. First week: test a payment timeout after the write, replay the same key and verify that the side effect happens once. Add a flag, a rollback path and a trace ID. A flag does not undo a payment.
ADRs beside the code
An architecture decision record preserves the constraint behind a choice. The next agent can read why a shared transaction matters before proposing a service split. First week: record the current decision, alternatives, consequences and the condition that would justify revisiting it. Link it from the module's README and keep the test command there. This is how we make 'no trend candy' a working rule.
A first-week slice in our Java/Spring projects
The request path is explicit. These are responsibilities and package boundaries, not separate deployable services:
adapter/in · RESTapplication/port/inusecaseapplication/port/outadapter/out · JPA
The domain model uses records. An application command carries the input to a use case; the use case calls an outbound port, whose JPA adapter handles persistence. Flyway owns schema migrations. Test the domain rule, the use case against a fake port, the REST contract and the JPA mapping against the migrated database.
Our feature skill can generate that complete vertical slice because the boundaries and test responsibilities are already defined. The engineer still checks the behavior, transaction boundaries and migration. On the frontend, the matching feature owns its form and validation, consumes types generated from the API schema, and has an e2e test for the user path.
This is our implementation of ports and adapters. The diagram shows a request path; source-code dependencies point toward application ports and domain code. Start with a slice small enough to review. Keep code that uses another module's work behind its interface.
Checks and collaboration
A planner, builder, verifier and reviewer need shared contracts for the work itself: the allowed scope, acceptance criteria and the commit each result describes. These are gates you can inspect in a PR.
Quality: check the promise, not just the build
Security: scanners plus your access rules
Review: give the reviewer evidence
Collaboration: allocate scopes before builders
See Brief 005: AI software factory for role orchestration and the two human gates. The map here asks whether those roles have clear places to work.
Read the evidence. Keep its scope.
Our project records describe how we work. The external studies measure delivery outcomes in particular settings. None validates the map's file counts, review formula or architecture ranking.

