One engineer in the middle. Five agents around them.
Click an agent to see what it does, what the engineer does, and who signs. Then type a commit message and watch the hook that guards our repositories accept or reject it. Below: the report a client received for the week of August 3–9, and what our own team says when asked anonymously.
Who does what · Reviewer
- The agent
- Reads the diff adversarially, a different model than the one that wrote it.
- The engineer
- Reads every line before merge. Not most lines. Every line.
- Who signs
- The engineer's name is on the merge. Always.
Responsibility does not orbit. It sits in the middle, with a name on every merge.
The commit prefix · try one
Every commit on the project starts with [AI:x%/model/tool]. A hook validates it; the weekly report is generated from git log. Edit the message or pick a real one from the report.
- AI share
- 88%50–89% · mixed
- Model
- claude-opus-5recorded per commit
- Tool
- vscoderecorded per commit
- Subject
- feat: excel-style cell selection in tableswhat the client sees in the report
Human-supervised AI in project delivery
Weekly report · August 3–9, 2026 · one client project · generated from git log
AI share distribution
- 90–100% 72
- 50–89% 6
- 1–49% 3
- 0% or unmeasured 11
Average AI share by area
Commits by model · by tool
57,854 lines added, 1,570 removed across merged commits. 81 of the 85 measured commits include AI-authored code (95%). The distribution groups 4 explicitly declared 0% commits with 7 unmeasured commits; their AI share is unknown. Model and tool metadata rolled out mid-week and covers 22 commits.
Team pulse · survey of August 28, 2026
62 of 130 people answered, 81% of them from engineering — so read this as our engineering team, not the whole company.
- 92%work in an agentic tool — Claude Code, Codex, Cursor, Copilot — not a chat window
- 73%use two or more model families; no single-vendor dependency
- 87%of the 30 heaviest users (80%+ of work with AI) rate their trust 4–5 of 5 — self-reported trust, not an accuracy test
- 79%report a positive hours balance; the rest did not report a positive balance
What the team says is in the way
- 29no time to learn
- 20no time to understand someone else's code
- 19quality and trust in AI output
Review capacity, not AI, is the bottleneck — which is why we do not sell "faster delivery" on the strength of these numbers.
AI writes the code. A person answers for it.
Adoption is easy to claim and hard to prove. These are the six habits behind the numbers above — none of them a tool, all of them a way of working you can inspect.
Measure it in the commit
The prefix records the declared AI share and, as metadata rolls out, the model and tool. The inspector rejects a missing prefix. In the historical report, 81 of 85 measured commits included AI-authored code (95%); seven more commits were unmeasured. Their AI share is unknown.
Frame before you prompt
The engineer writes what must be true when the change lands: the constraint, the definition of done, what must not move. That is the difference between a plan the agent can execute and a wish it will interpret.
Pick the model for the job
In the report week engineers routed backend infrastructure to Codex, QA hardening to Claude Code and frontend interaction work to Claude in VS Code. Two model families, four tools, chosen per task — and recorded, so the choice can be judged later.
Read every line
Generated code is fluent, long and confident. The one wrong line looks like the fifty right ones. Nothing merges unread; the reviewer is a person, and their name is on the merge. Review capacity — not model capacity — is our real bottleneck, and we say that too.
Keep some things deliberately manual
Credentials, migrations on production data, a handful of decisions where a wrong keystroke costs a customer: reasons to choose manual work. Four commits declared 0% AI; seven had no measurement. The report groups those eleven together, but missing data does not tell us how the work was done.
Count the losses
Our own survey asks how many hours AI wasted, not only how many it saved. The ratio is 5.6 to 1; 79% report a positive balance, while 21% did not report a positive balance. That remainder may include zero as well as losses. We publish both hours gained and hours wasted.
Everything above is counted.
Two of our own instruments, and the public research that says why supervision is the point.

