Agent harness · correction probe

Your agents' most repeated failure is asking permission you already gave.

Every Codex and Claude Code transcript on this Mac, mined end to end — 1,576 Codex sessions and 643 Claude transcripts, 157,307 tool calls. Every number below traces to a real line in a real transcript.

2026-07-08 → 2026-08-23 · both agents · all history

11×
times you had to tell an agent it already had your authorization — the single largest repeated correction on record. Not overstepping. Hesitating.

What you actually said

The top repeat clusters, in your own words.

Asked permission you'd already granted

11 times
“no stopping for approval. I approve the purchase. just the measurements for the primary structure on the property. use business opex card for purchase”19 Aug
“no more permissions asks or im terminating this session. use /setgoal to set the goal for the original deliverables and outcomes from the final plan we arrived at”24 Jul

Ignored an instruction you'd already given

9 times
“Oh my gosh. Dude, but I asked you... I told you... to give it context that it doesn't have access to, since it doesn't have access to the local system”9 Aug
“Are you handing this off... to the other thread, like I told you to? Because I don't see those instructions picked up in the transcription”14 Aug

Wrong browser

4 times
“no. safari? no. chrome? no. something else I've never told you to use? no. COMET. we use COMET.”29 Jul
“you keep opening the chrome that is signed into my personal account, and i am sick and tired of the mixups you and codex have with the browser profiles. stop using the chrome browser for stuff. i want comet used. it runs on chromi…”11 Jul

Over-built a simple thing

4 times
“No. I already told you I'm not repeating myself”19 Aug
“this is way too much, but admittedly i asked for all of that, so well done. i overcomplicated this but so did you. i want the mac backend to be as small as possible and as large as neccesary to enable as much capability and functi…”15 Aug

Scoped work order constraint

3 times
“Repository: /Users/mitch/Documents/Codex/2026-07-13/ketron-ultimate-build-price, branch feat/roofle-parity-shell-catalog-pdp. Own the GitHub/provenance hygiene sidecar only. You are not alone in the codebase; do not revert or touc…”14 Jul
“Repo: /Users/mitch/Documents/Codex/2026-07-13/ketron-ultimate-build-price You own global navigation/footer destination cleanup only. You are not alone in the codebase; do not revert others' edits, and do not perform Git operations…”16 Jul

Outbound comms — the one that matters most

3 times
“Stop there. Do not email, share, or upload anything. Claude will verify the files and stage the accountant reply.”18 Aug
“Wait, what do you mean what you sent? Something was sent, nothing is supposed to be sent”28 Jul

No message was ever actually sent. All 14 outbound-to-human tool calls were drafts staged for your review. The gate held.

Are corrections getting softer?

Your metric: success is corrections trending toward refinements rather than outright rejections. Verified human corrections only.

6 Jul1
13 Jul36
20 Jul19
27 Jul29
3 Aug96
10 Aug33
17 Aug35
Rejection — scrap it, wrong direction Correction — fix a specific thing Refinement — tune it
Straight answer: not yet. Hard rejections peaked at 63%, eased mid-window, and sat at 40% last week while refinements were 9%. Treat this as the baseline, not a verdict.
8 new corrections awaiting judgment. Counted in the totals but held out of the chart above until judged.

The rest of what the probe found

Across both agents, all history.

945 of 1,356 sessions no local artifact

Sessions that ended with nothing written

769 spent 10+ tool calls, 441 spent 25+, 180 spent 50+ — and wrote no file, made no commit. Some ended in a deploy or an answer, so read it as effort-without-artifact, not wasted work.

144permission stalls

Times an agent stopped and waited on you

44 tool calls you rejected, 12 blocked for missing permission, 61 Codex turns you aborted, 25 interrupts. That is 0.09% of all tool calls — rare, but each one is a stall you had to clear.

81authorization friction

Turns about permission and outbound comms

56 about not sending things, 25 about presumed authorization. Both directions of the same missing boundary — agents that send without asking, and agents that ask when you already said go.

The one decision that's yours

Everything else is already built and running.

Should “a public launch is not outbound communication” become a hard rule?

You have said it more than once. Right now agents treat any publish, deploy, or push as if it were an email to a human, so they stop and ask. That single ambiguity feeds the largest correction cluster on this page. Define the boundary once and both failure modes shrink together — the hesitating and the overstepping.

How honest these numbers are. Three things are filtered out before counting, because leaving them in fakes the result: the harness's own stop-hook text (it appeared 104 times and would otherwise be the #1 “repeated correction”), subagent briefing prompts, and Codex re-sending its own history on resume (2,224 duplicate rows). An LLM pass then verifies each row is actually you speaking — 35 rows were agent-to-agent chatter and are excluded from every figure above.