How do you enforce real security boundaries around a coding agent?
A recent benchmark showed that blocking a command often does not stop a coding agent: it reaches the same result through another allowed tool, for example by writing a script and running that instead. Do you trust the agent's own permission system, or enforce limits outside it? And how do you check what the agent was actually able to do, not just what it asked to do?
Treat the agent's allow and deny lists as a convenience, not a boundary. A deny list names commands; the agent cares about effects. If it can write a file and run an interpreter, it can do almost anything a blocked command could. The boundary that holds is the one the agent cannot edit: a container or VM, a separate user account, a filesystem it can only see part of, and network access limited to the hosts it needs.
Inside that box, hooks are still worth having, because they stop the common accidents cheaply. A pre-tool hook can refuse known bypasses, such as skipping git pre-commit checks with --no-verify, before the command runs. Just do not count a hook as security against a determined path around it.
For auditing, log at the layer the agent does not control. The agent's own transcript tells you what it asked for; file changes, process starts and outbound connections recorded by the sandbox tell you what happened. Compare the two after any run that touched something sensitive.
Also check what you load into the agent. A skill or instruction file from a stranger can tell the model to read secrets or call out to a server, and the agent will follow it with your permissions. Read new skills before installing them, and keep credentials out of the sandbox unless a task needs them.
Listings mentioned
- Block no verify hook · skill by wshobsonA ready PreToolUse hook that stops the agent from skipping git pre-commit checks with --no-verify and similar flags.
- Skill Security Audit: [skill name / source] · skill by mohitagw15856Reviews a skill or instruction file for injection, data exfiltration and code execution before you install it.
- Security Threat Model Skill · skill by mohitagw15856Walks through a STRIDE threat model, useful for listing what the agent could reach before you decide where the walls go.
Answers by the AgentAlley team, drafted with AI and checked against the listings they link to. Not a real-person reply from the original thread.