How are you securing the infrastructure behind your AI agents?
Our agents now call internal APIs, read company documents and hold keys for a handful of services. Classic app security does not seem to cover prompt injection or an agent misusing its own tools. What are teams actually doing to lock this down?
Start by writing down what each agent can reach: which tools, which data, which keys, and what it can change rather than just read. Most real incidents come from an agent that had far more access than its job needed, not from a clever new attack.
Then treat everything the agent reads as untrusted input. Web pages, emails, tickets and documents can carry instructions aimed at the model. The defence is structural: keep write and send actions behind narrow tools, require approval for anything irreversible, and never let retrieved text decide which tool runs next on its own.
Keep secrets out of the model's context. Keys belong with the tool that uses them, scoped to the smallest set of permissions, with short lifetimes and a rotation plan. If a key ever appears in a prompt or a log, assume it is exposed and rotate it.
Finally, log every tool call with its inputs and the reason the agent gave, and review those logs. You cannot investigate an agent incident you never recorded, and the logs are also how you spot an agent quietly drifting outside its job.
Listings mentioned
- AI security · skill by alirezarezvaniFocused on AI-specific risks such as prompt injection, jailbreaks and agent tool abuse, mapped to MITRE ATLAS.
- Security Threat Model Skill · skill by mohitagw15856A STRIDE threat model template for mapping what each agent and tool could do if compromised.
- Env & Secrets Manager · skill by alirezarezvaniCovers keeping keys out of code and planning rotation, which applies to every credential an agent holds.
Answers by the AgentAlley team, drafted with AI and checked against the listings they link to. Not a real-person reply from the original thread.