AI in engineering companies · Agents
Guardrails: autonomy is earned
Every conversation about AI agents eventually reaches the same question: how much should we let it do on its own?
My answer is that autonomy should be earned, not granted. An agent should start with the least freedom that lets it be useful, prove itself on a well-understood task, and be given more only when the record shows it has earned it. That is how we treat new engineers, and it is how we should treat agents.
This article sets out the guardrails that make that possible. None of them are exotic. Most are the same controls any safety-conscious engineer would apply to an automated system. What is different is that an agent's behaviour is shaped by what it reads, which makes some of them more important than they first appear.
Start with the tools, not the agent
An agent can only do what its tools allow. That makes tool design the first and most powerful guardrail.
Least privilege. Give each agent only the tools its task needs. The fault investigation agent described in Agents at work in engineering needs to read build records and telemetry. It does not need to write to them, send email, or deploy firmware.
Narrow tools. A tool that runs a specific, defined query is far safer than one that runs any query. "Fetch the fault history for this unit" is a tool. "Run this database command" is an invitation to trouble.
Act as the user. An agent working for a person should see only what that person is allowed to see, as I described in From what the manual says to what the product is doing. It should not have a master key.
Read before write. Keep agents read-only until there is a clear reason to give them the ability to change things, and then give that ability one action at a time.
Put people at the points of consequence
Not every step needs approval. Steps that change the world do.
Approval gates. Before an agent sends anything outside the organisation, changes a record, releases anything or spends significant money, a named person approves it. This is the principle of AI drafts, engineers decide, extended to actions.
Show what will happen. The approval request should show exactly what the agent intends to do, with the evidence behind it. An approver who is shown only "approve?" is not really approving.
Verify what was approved. The action carried out should be exactly the action that was approved. If anything changes between approval and execution, it should need approving again.
In the opportunity platform I am developing, the design follows these rules strictly. AI roles work within defined limits, nothing is sent externally without my approval, the approved content is checked before it is dispatched, and uncertain outcomes are handled explicitly rather than retried blindly. It is still in development, and I describe it as such, but building it has made the value of these controls very concrete.
Limit what can go wrong
Budgets. Every agent run should have limits on time, number of steps, retries and cost. An agent that is stuck will keep trying, and without limits it can spend a surprising amount achieving nothing. I return to this in The cost and infrastructure of agents.
Bounded retries. Retrying a failed step can be sensible. Retrying an action that may already have succeeded, such as sending a message, can cause real harm. Actions with consequences need to be designed so that repeating them is safe, or so that they are never repeated without checking.
A stop button. Any agent must be able to be stopped immediately, by a person, with its state preserved so the stop can be understood afterwards.
Treat what the agent reads as untrusted
This is the guardrail that is most specific to AI, and the one most often missed.
An agent's behaviour is shaped by the text it reads. If an agent reads customer tickets, supplier emails, web pages, documents or even code comments, then anyone who can write those things can potentially influence what it does. Text containing instructions, whether malicious or accidental, can steer an agent that has not been designed to resist it.
This cannot be fully solved by better prompts. It has to be solved by design:
- keep a clear separation between instructions from the organisation and content the agent is processing
- assume any content from outside may contain instructions, and never let it grant the agent new permissions
- make sure agents that read untrusted content cannot take consequential actions without human approval
The security series that follows this one covers this in depth, in When the data gives the orders and Agents with keys.
Keep a record
Log every step. What the agent was asked, what it read, which tools it called with what inputs, what it concluded and what it did. Without that record, an agent is impossible to debug, impossible to audit and impossible to trust.
Record approvals. Who approved what, when, and on the basis of what evidence.
Review the record. Logs that nobody reads are only useful after something has gone wrong. Regular review of what agents actually did is how you find problems early, and how you build the evidence to widen their autonomy.
Widening autonomy, one task at a time
With those guardrails in place, autonomy can grow safely.
Start an agent in a mode where it proposes and a person approves everything. Measure how often its proposals are accepted unchanged, how often they are corrected and how often they are rejected. When, for a specific and well-understood task, the record shows it is consistently right, consider letting it act without approval for that task alone, with monitoring and the ability to reverse.
That is earned autonomy: granted task by task, on evidence, and withdrawn if the evidence changes. It is slower than switching everything on. It is also the only approach I would stand behind.
Four things worth taking seriously
For engineering leaders: before approving any agent, ask what it can reach, what it can change, who approves its actions and where the record is kept. If those answers are unclear, it is not ready.
For security teams: treat every agent as a new kind of user with its own identity, permissions and audit trail. And treat everything it reads from outside as untrusted.
For anyone building agents: design the guardrails first. They are much harder to add afterwards.
For everyone: autonomy is not a feature to switch on. It is a privilege an agent earns, one task at a time.
I would be interested to hear how your organisation decides what an automated system is allowed to do on its own, and whether that thinking has caught up with AI.
Further reading in this series
- AI drafts, engineers decide (the five levels series)
- Workflows before agents
- Running agents in production
- When the data gives the orders (security series)
- Agents with keys (security series)