How to stop a coding agent from taking actions without asking
A question on Hacker News this week asked how to stop a coding agent from taking actions without asking first. It is a fair worry. You ask for a small fix and the agent edits ten files, runs a migration, or commits on its own.
Here is the short answer. Then five free skills that make "ask first" the default instead of a hope.
The short answer: put the gate outside the prompt
Telling an agent "please ask first" works until a long session buries the instruction. Stack three layers instead, from hardest to softest:
- Tool permissions. In Claude Code, keep the default mode that asks before edits and shell commands, add deny rules for commands you never want run, and use plan mode when you only want a proposal. Never skip permission prompts outside a throwaway machine. Read your tool's permissions page before changing a setting.
- Isolation. Work on a branch, in a read-only sandbox, or in a container. A wrong action then costs a deleted branch, not a broken main.
- Workflow rules. Skills that make the agent stop at fixed points: plan, wait for a yes, review, report.
None of this is a guarantee. It moves the default from "act, then tell you" to "tell you, then act".
Make it write a plan and wait for your yes
Claude Superpowers splits a coding task into four stages: Plan, Isolate, Test First and Double Review. Before any code, the agent writes a plan listing every file it will create, change or delete, plus edge cases and assumptions. It ends with "Confirm this plan before I start coding." Per its file, the agent must not go on until you say yes.
Stage two keeps work off your main code: a branch named superpowers/[task-slug], no files outside the plan, and a question to you if new scope turns up mid-task.
Good for: anyone whose agent "helpfully" rewrites more than was asked.
Default to read-only when you hand work to Codex
If you run OpenAI's Codex CLI from inside Claude Code, the Codex skill sets tight defaults. It picks --sandbox read-only unless the task needs edits or network access. Its safety section says not to pick --full-auto or a wider sandbox than the task needs. It also says not to overwrite files, create commits or do destructive operations without your explicit approval.
Before any high-impact flag, such as --full-auto or danger-full-access, it asks you through a question prompt, unless you already said yes. After every Codex run, it asks you to confirm the next step.
Catch: you need the Codex CLI installed and set up. The skill only runs it when you ask for Codex by name.
Limit which tools the agent can reach at all
An agent cannot misuse a tool it does not have. The Docker MCP Gateway runs each local MCP server in its own Docker container. Per its README, npx and uvx servers get minimal host privileges, and API keys stay in Docker Desktop's secrets store instead of environment variables.
The key part here is the tool allowlist. Per profile, turn on single tools such as --enable github.list_repos and turn off others such as --disable github.search_code. Logging and call tracing show what the agent called.
Catch: it needs Docker Desktop 4.59 or later with the MCP Toolkit feature turned on.
Check a skill before it can tell your agent what to do
Sometimes the action you did not approve came from a skill you installed. Skills are plain text, but text can tell a model to run commands. Skill Security Auditor reads a SKILL.md and every bundled script before you install it. It looks for destructive shell such as rm -rf /, dd or chmod 777, data sent to outside URLs, hidden Unicode, and instructions that try to widen permissions.
It gives one of three verdicts: safe to install, install with caution, or do not install. Its rule is strict: any high-severity finding means do not install. Each finding must quote the exact line as evidence.
Good for: anyone pulling community skills into a repo where the agent has shell access.
Ask for a review that reports instead of fixing
The last gate sits after the work. Code Review checks changes for defects, with correctness as the default and security as an optional pass. One line in its file matters most here: "Review and report. Apply fixes only when asked." It also states its scope and baseline before it starts, and asks instead of guessing when your choice of checks is unclear.
So you get a short list of findings and the agent keeps its hands off the code until you decide.
Catch: Claude Code has its own built-in /code-review. To call this one, use the plugin form shown in its file.
Summary
Keep your tool's approval prompts on, isolate the work, and add skills that stop at fixed points: a plan gate before, a sandbox and allowlist during, an audit before install, and a report-only review after. Every skill above is free, with the full source on its page. New to skills? Start with how to install a skill.