Agent harness · correction probe

Your agents' most repeated failure is asking permission you already gave.

Every Codex and Claude Code transcript on this Mac, mined end to end — 1,566 Codex sessions and 629 Claude transcripts, 154,707 tool calls. Every number below traces to a real line in a real transcript.

2026-07-08 → 2026-08-21 · both agents · all history

11×
times you had to tell an agent it already had your authorization — the single largest repeated correction on record. Not overstepping. Hesitating.

What you actually said

The top repeat clusters, in your own words.

Asked permission you'd already granted

11 times
“no stopping for approval. I approve the purchase. just the measurements for the primary structure on the property. use business opex card for purchase”19 Aug
“no more permissions asks or im terminating this session. use /setgoal to set the goal for the original deliverables and outcomes from the final plan we arrived at”24 Jul

Ignored an instruction you'd already given

9 times
“Oh my gosh. Dude, but I asked you... I told you... to give it context that it doesn't have access to, since it doesn't have access to the local system”9 Aug
“Are you handing this off... to the other thread, like I told you to? Because I don't see those instructions picked up in the transcription”14 Aug

Wrong browser

4 times
“no. safari? no. chrome? no. something else I've never told you to use? no. COMET. we use COMET.”29 Jul
“you keep opening the chrome that is signed into my personal account, and i am sick and tired of the mixups you and codex have with the browser profiles. stop using the chrome browser for stuff. i want comet used. it runs on chromi…”11 Jul

Over-built a simple thing

4 times
“No. I already told you I'm not repeating myself”19 Aug
“this is way too much, but admittedly i asked for all of that, so well done. i overcomplicated this but so did you. i want the mac backend to be as small as possible and as large as neccesary to enable as much capability and functi…”15 Aug

Scoped work order constraint

3 times
“Repository: /Users/mitch/Documents/Codex/2026-07-13/ketron-ultimate-build-price, branch feat/roofle-parity-shell-catalog-pdp. Own the GitHub/provenance hygiene sidecar only. You are not alone in the codebase; do not revert or touc…”14 Jul
“Repo: /Users/mitch/Documents/Codex/2026-07-13/ketron-ultimate-build-price You own global navigation/footer destination cleanup only. You are not alone in the codebase; do not revert others' edits, and do not perform Git operations…”16 Jul

Worked from stale context

3 times
“I told you I'm not home. Why did it become unaccessible? What does that even mean”4 Aug
“No, we already finished that yesterday”19 Aug

Are corrections getting softer?

Your metric: success is corrections trending toward refinements rather than outright rejections. Verified human corrections only.

6 Jul1
13 Jul36
20 Jul19
27 Jul29
3 Aug96
10 Aug33
17 Aug38
Rejection — scrap it, wrong direction Correction — fix a specific thing Refinement — tune it
Straight answer: not yet. Hard rejections peaked at 63%, eased mid-window, and sat at 40% last week while refinements were 8%. Treat this as the baseline, not a verdict.

The rest of what the probe found

Across both agents, all history.

936 of 1,337 sessions no local artifact

Sessions that ended with nothing written

761 spent 10+ tool calls, 437 spent 25+, 179 spent 50+ — and wrote no file, made no commit. Some ended in a deploy or an answer, so read it as effort-without-artifact, not wasted work.

134permission stalls

Times an agent stopped and waited on you

43 tool calls you rejected, 12 blocked for missing permission, 61 Codex turns you aborted, 16 interrupts. That is 0.09% of all tool calls — rare, but each one is a stall you had to clear.

81authorization friction

Turns about permission and outbound comms

56 about not sending things, 25 about presumed authorization. Both directions of the same missing boundary — agents that send without asking, and agents that ask when you already said go.

The one decision that's yours

Everything else is already built and running.

Should “a public launch is not outbound communication” become a hard rule?

You have said it more than once. Right now agents treat any publish, deploy, or push as if it were an email to a human, so they stop and ask. That single ambiguity feeds the largest correction cluster on this page. Define the boundary once and both failure modes shrink together — the hesitating and the overstepping.

How honest these numbers are. Three things are filtered out before counting, because leaving them in fakes the result: the harness's own stop-hook text (it appeared 104 times and would otherwise be the #1 “repeated correction”), subagent briefing prompts, and Codex re-sending its own history on resume (2,217 duplicate rows). An LLM pass then verifies each row is actually you speaking — 35 rows were agent-to-agent chatter and are excluded from every figure above.