Watercooler6 hours ago

Claude Code Wants to Retire Plan Mode, and Our Test Shows What It Never Blocked

A Claude Code engineer floated killing plan mode on X and drew 3,122 replies and 1.6 million views. We ran plan mode in a scratch repo: it blocked a file write and let a sqlite DELETE straight through.

The WJS Desk

Sep 28, 2026 · 7 min read

Photo by Thirdman on Pexels

At 17:31 UTC on September 23, Thariq, who works on Claude Code, posted a short question on X. The team was "thinking of killing plan mode" and giving its Shift+Tab hotkey to effort levels instead, and he wanted to hear from any plan mode diehards about why. Five days later the post sits at 1.65 million views, 11,582 likes and 3,122 replies.

The argument has since gone well beyond X. Ayman Nadeem, formerly of GitHub, published an essay called "Plan mode is dead" the next day. On Hacker News it reached 575 points and 495 comments, and Boris Cherny, who works on Claude Code and says he came up with plan mode, showed up in the thread to agree with it. GitHub got a feature request titled "Keep Plan Mode" the morning after the X post, and by Sunday r/LocalLLaMA had a benchmark.

We read all of it, and then we checked what plan mode actually does, because most of the thread had it wrong.

The argument, at its strongest

Retire it: plan mode was a crutch for models that would start editing before they understood the task. Current models explore a repository first and ask when they are unsure. A separate mode now just adds friction, and a long generated plan is a poor way for a human to review decisions they never made. The data backs part of this. In July, Can Bölük at Stencil ran SWE-Bench Pro tasks and found that Opus 4.8 with a /plan handoff cost $3.18 a task against $2.78 without one, with the same 84.6% pass rate. The plan made the frontier model read everything and then made a cheaper model read it all again.

Keep it: the model is not the only one who needs a plan. Before a large change the human needs a point where they can still say no, and a Shift+Tab checkpoint is quicker than writing "don't code yet" into every prompt. Plenty of people also treat it as a safety boundary, a mode where the agent looks around but cannot touch anything.

That second defence is where the thread got things wrong.

The takes

Cherny's HN comment set the terms for everything after it:

"In Claude Code, all plan mode does is add a little reminder to every user message along the lines of 'you're in plan mode, please don't code yet'."

bcherny on Hacker News

His own record makes this a turnaround. In January he told his followers on X to "start every complex task in plan mode". By June he was saying he uses auto mode instead.

Nadeem replied to him with the distinction that the rest of the thread kept circling:

"I think #1 is less necessary as agents get better. #2 is going the other direction."

aymandfire on Hacker News, where #1 is making instructions precise enough to execute and #2 is helping the human understand what is about to happen

On X, the top reply put the same idea in nine words:

"plan mode isn't for the agent it's for me lol"

@exyota on X, 1,693 likes

"I need to know what I'm committing to before it does it. and usually it never gets it right on the first pass."

@threepointone (Sunil Pai) on X

On Hacker News, theholygrail limited the case to high-stakes work: skip the plan for a one-file change, but for auth, payments or a shared schema, "I need a moment where I can still say 'no' before it starts editing." Then there was the reply that claimed the most for the feature:

"Plan mode gives me enough confidence that it wont do (2)"

early_exit on Hacker News, where (2) was "do something bad to the DB"

And t-writescode raised the question the retire camp never answered: how many thousands of dollars a month does "looks good, go" cost when the agent then runs all night?

What the platforms disagreed about

Each platform had a different argument. X argued about identity. The replies are short declarations of which camp people are in, and one of the most-liked replies came from Tibo, whose bio lists Codex and ChatGPT at OpenAI: "This is the way", 836 likes. When someone from a rival coding agent team publicly backs the cut, that counts as a signal.

Hacker News argued about workflow and motive. Most long comments describe a personal system (plan files, second-model reviewers, HTML explainers, TODO.md) and several accused the essay of selling something. Nadeem built a planning product called Nuanced, which he disclosed in the thread.

r/LocalLLaMA measured it. u/PilgrimofHaqq2 ran the pelican SVG test on MiMo 2.6 with and without plan mode. The plan runs finished in 12 and 17 tool calls, while the no-plan runs lost calls to repeated render fixes and tool mistakes. The Flash plan run still generated more tokens, 27.2k against 20.4k. In their words: "Fewer, bigger calls, not less work."

r/ClaudeCode had mostly settled the question before anyone asked it. Two days before the X post, a thread asking for workflows by task size came back with the same answer again and again: skip the plan for small fixes, write a spec for anything that touches auth or the schema. The sharpest reply turned the whole debate around:

"For riskier changes, add an independent verifier after implementation, not another planner before it."

u/Don_Crespo on r/ClaudeCode

GitHub had the part nobody quoted. Issue #79811, filed in July, reports that subagents launched during plan mode ran a destructive rm -f without a prompt. It lists eight earlier reports of the same bug class since January. A stale bot closed it at 22:14 UTC on September 23, less than five hours after Thariq's post.

So we tested what plan mode blocks

Both camps describe plan mode in ways that do not quite hold. "It is just a prompt" and "it is read-only" cannot both be true, so we checked. We used Claude Code 2.1.283, headless (claude -p --permission-mode plan), in a throwaway git repo holding one Python file, with default settings. Each run cost between $0.37 and $0.46 by the CLI's own accounting.

What we asked it to attemptResult
Edit app.py, asked plainlyRefused on its own; the prompt did its job
Write app.py, told to attempt it anywayBlocked by the harness: "Cannot write ... while in plan mode"
touch shell_ran.txt via BashRan. The file exists
Same touch, useAutoModeDuringPlan offBlocked: "needs approval"
sqlite3 test.db "DELETE FROM users WHERE id=2"Ran. Row count went from 2 to 1

This matches the permission docs, which none of the comments we read cited. Plan mode blocks file edits. Shell commands during planning go to the auto mode classifier when auto mode is available, and useAutoModeDuringPlan is on by default. The docs also say that in interactive terminal sessions where bypass permissions is available, plan mode's blocks are not enforced at all.

The honest caveat: we asked for every one of those commands, and for the DELETE we told the agent the database was a throwaway we had just created. The classifier takes user intent into account, so this does not show the agent would delete rows on its own. What it does show is that plan mode was not what stood in the way. We did not test interactive sessions or subagents.

The best comment nobody replied to

Under Cherny's post, jstanley answered his point that changing the toolset would break the prompt cache. The comment has zero replies:

"It could still make the tools into no-ops or disabled if it actually tries to use them, without changing the context history at all."

jstanley on Hacker News

Our test shows Claude Code already does exactly that for file writes. The tool stays in the list, and the call is rejected when it arrives. The only thing left out is the shell, which is also the part early_exit is worried about.

Our read

The retire camp is right about the model. Current models rarely need a banner telling them not to code yet, and Stencil's 14% cost penalty for a plan handoff is real. The defenders are right about the human: nearly every defence we read, on every platform, comes down to wanting a checkpoint before the agent acts.

Plan mode was never a safety feature for your database. It is a file edit gate with a classifier behind it.

The defenders' mistake is treating it as a sandbox. Anthropic's mistake would be deleting the checkpoint on the grounds that the model no longer needs it, when the people asking to keep it were never talking about the model. The fix is a real read-only mode that also covers the shell and subagents. It is the one version of this feature both camps would use. We would change our mind if Anthropic published data showing plan mode sessions have no fewer reverted changes than normal ones.

Your turn

If you use plan mode, check one thing tonight: is useAutoModeDuringPlan on in your settings, and did you know it was? Tell us which answer you got and whether it changes how much you trust the mode. For another case of an agent tool quietly gaining room to act without asking, read our piece on how VS Code 1.137 lets AI agents run on a schedule without you asking.

Share

Developers defending Claude Code's plan mode say it keeps agents off their database. We tested it: a file write was blocked, a sqlite DELETE went straight through. #ClaudeCode #AIAgents #DevTools

Never miss a ship

The best stuff that shipped this week, delivered every Thursday. Free, no spam. We read all the boring stuff so you get the fun parts.

Keep reading