AI in engineering companies · Security

Agents with keys

An AI assistant that answers questions has a limited capacity for harm. An agent with access to your systems is different. It holds keys: credentials, permissions and tools that let it read and change things on the organisation's behalf. If it can be misled, misconfigured or simply make a mistake, those keys go with it.

In the agents series, I argued in Guardrails: autonomy is earned that agents should be given the least freedom that lets them be useful, and more only when the record justifies it. This article looks at the same subject from the security side: how agents with access get misused, and how to design so that when something goes wrong, the damage is contained.


The confused deputy

There is an old idea in computer security called the confused deputy: a program that holds legitimate authority and is tricked into using it on someone else's behalf. Agents are a textbook case.

An agent working for an engineer may be able to read telemetry, query build records and draft messages. If content it reads, such as a customer ticket or a supplier document, contains instructions, as described in When the data gives the orders, the agent may try to use its legitimate access for the attacker's purpose. The agent is not compromised in the conventional sense. It is doing what agents do: following instructions it has read. It has simply been given the wrong ones.

The defence is not to hope the agent can tell the difference. It is to make sure that even a thoroughly confused agent cannot do much harm.


Excessive permissions

The most common weakness in agent systems I see is not a clever attack. It is an agent with far more access than its task requires, usually because it was easier to set up that way.

An agent given a broad database credential "so it can answer anything" can also change anything. An agent given access to email "so it can draft responses" can also send them. An agent running with an administrator's identity inherits everything that administrator can do.

What to do. Give each agent its own identity, not a shared or personal one. Grant only the specific tools its task needs, and make those tools narrow: a defined query rather than open database access, drafting rather than sending. Where an agent works for a person, it should act with that person's permissions, never more. Review agent permissions as you would review any privileged account.


Credentials

Agents need credentials to reach systems, and credentials are among the most valuable things an attacker can obtain.

What to do. Never put credentials where the model can read them, including in its instructions, in documents it can retrieve, or in its working memory. Keep them in the surrounding system, which uses them on the agent's behalf when a permitted tool is called. Use short-lived credentials scoped to the task. Rotate them. And make sure logs, which record everything the agent does, do not record the secrets it used.


Runaway loops and runaway bills

Not every harm is a breach. An agent stuck in a loop, repeatedly retrying a failing step, can consume computing resources and money very quickly. An attacker can sometimes cause that deliberately, by feeding an agent content designed to keep it busy. This is sometimes called denial of wallet.

What to do. Put hard limits on every agent run: number of steps, time, retries and spending. As I described in The cost and infrastructure of agents, the platform I am developing checks the budget before every AI call and reconciles actual usage afterwards. That control is as important for security as it is for cost.


Actions that cannot be taken back

Some actions are easy to reverse: a draft can be deleted, a record corrected. Others are not: a message sent to a customer, a firmware release pushed to the field, data sent outside the organisation, a payment made.

What to do. Classify every action an agent can take by whether it can be reversed. Irreversible actions should always need human approval, with the approver shown exactly what will happen. The approved action should be verified before it runs, so that what is executed is exactly what was approved. Actions should be designed so that repeating them by accident is safe, or so they are never repeated without checking.


Many agents, many paths

Where several agents work together, as discussed in Running agents in production, each handoff is another place where untrusted content can travel and permissions can blur. An agent that only reads may pass manipulated content to one that can act.

What to do. Treat content passed between agents as untrusted in the same way as content from outside. Keep permissions separate for each agent, and make sure an agent that reads untrusted material cannot, directly or through another agent, trigger an action that has not been approved.


Contain, detect, stop

Pulling this together, the design goal for any agent with keys has three parts.

Contain. Assume the agent will at some point try to do something it should not. Limit its access so that the worst it can do is small.

Detect. Record every step, tool call and action. Watch for unusual behaviour, such as an agent using a tool it rarely uses, reaching for data outside its task, or running unusually long.

Stop. Make sure any agent can be halted immediately, and that its credentials can be revoked in one place.

These are the same principles I have applied to every system I have designed that operates remotely and unattended. They simply need applying to a new kind of actor.


Four things worth taking seriously

For security teams: add AI agents to your inventory of privileged identities, with their own credentials, permissions and review cycle.

For engineering leaders: ask, for each agent, what the worst thing it could do is if it followed the wrong instructions. If the answer is serious, reduce its access.

For anyone building agents: keep credentials out of the model's reach entirely, and put a person in front of every action that cannot be reversed.

For everyone: an agent is a deputy with your keys. Give it only the keys it needs.


I would be interested to hear how your organisation manages the credentials its automated systems use, and whether AI agents have been added to that picture.


Further reading in this series

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.