Read-only first: a permissions model for production AI agents

Published Updated 7 min read

Start a production AI agent with permission to read and nothing else, then grant writes one narrow scope at a time. Give the agent its own identity, keep each credential to the least it needs, log every call, and make revoking access one quick step. What an agent cannot reach, no instruction can make it change.

Why an instruction is not a control

A language model can be told “do not change anything in production”, and it will usually comply. Usually is not a security property. The model can misread the instruction, lose it in a long session, or be steered by text it reads, because an incident agent reads a lot of text it did not write: log lines, ticket bodies, commit messages, chat. Any of that can carry instructions of its own. This is often called prompt injection.

So a production agent needs two layers. Instructions shape what it tries to do. Permissions decide what it can do, and they hold when the instructions fail. The rest of this page is about the second layer, and it follows one rule: least privilege. Give the agent the least access that lets it do the job, and add access only when the job needs it.

Reads come first because they are where most of the value is. Triage, investigation and summaries all need access to data, and none of them needs a tool that changes anything. The four jobs of an AI SRE all start with reading.

Give the agent its own identity

Before any scope, the agent needs an identity of its own: a service identity, not a person’s login and not a shared admin key. Three things follow from that.

  • Attribution. Every call shows up in logs as the agent, so a review can tell what the agent did from what a person did.
  • Independence. Its access can be narrowed or revoked without affecting anyone’s work.
  • Short-lived credentials. Tokens that expire on their own limit how long a leak is useful, and make revocation the default rather than a chore.

One identity per agent role is better than one for all agents. An agent that only reads dashboards should not share a credential with one that can restart services.

Not all reads are equal

“Read access” is not one permission. What an agent reads can leak, and some reads carry more risk than others.

TierWhat it readsMain riskStarting point
R1Metrics, service topologyLowGrant
R2Deploy history, code, pull requestsSecrets committed to code, private logicGrant, with secret scanning in place
R3Application and system logs, tracesPersonal data and secrets in logs, and in trace attributesGrant with redaction, or a filtered view
R4Databases and customer recordsPersonal data, contractual limitsWithhold until a specific need is agreed

The tiers give a team a vocabulary for the argument it has to have anyway: what is the agent allowed to see, and who decided. Logs deserve care because they were never designed as an access-controlled store, and they often hold more than anyone remembers. Where the agent needs a log, give it a filtered view rather than the raw stream.

Writes: narrow tools, not broad access

When an agent needs to change something, resist giving it a general credential and a shell. Give it named tools that each do one thing: restart this service, scale this group, toggle this flag. A tool built that way has four advantages.

  • Arguments are checked. The tool accepts only the services and values it is meant to, not whatever the model produces.
  • The blast radius is the tool’s. A tool that restarts one service cannot delete a database, whatever the model is told.
  • It is reviewable. A list of twelve named tools is something a security team can read. “Access to the cloud account” is not.
  • It can carry a gate. Each tool can require approval, or run alone within limits. How to place those gates is the subject of human-in-the-loop for incident agents.

Group writes by reversibility. A restart or a scale-up can be repeated. Changing data or configuration that is not versioned cannot be reliably undone, and belongs with a person.

An example: scopes for one agent

This is illustrative. A team grants an investigation agent access to a search service in stages.

ScopeAccessGateReviewed
Metrics, searchReadNone, loggedQuarterly
Deploy history and pull requestsReadNone, loggedQuarterly
Logs and traces, search, redactedReadNone, loggedMonthly
Restart one instanceNamed tool, search onlyOn-call engineer approvesMonthly
Anything on the databaseNoneNot available to the agentn/a

The last row matters as much as the others. Writing down what the agent cannot do, and why, keeps the list from growing by accident.

Audit: a record you can read

Every tool call should leave a record: which identity, which tool, with which arguments, on whose request, at what time, and what came back. Keep the record apart from the agent, so the agent cannot edit its own history, and link each entry to the incident it belongs to. Two uses follow. During an incident, the record tells a responder exactly what the agent has already tried. Afterwards, it is the evidence for whether the permissions were right.

Revocation: one step, tested

Assume that some day access will need to be taken away fast: the agent is misbehaving, a credential leaked, or the team has simply lost confidence. Plan for it.

  1. One switch. Revoking the agent’s access is a single action that an on-call engineer can take.
  2. Short-lived tokens. Revocation then happens by default, and a leaked token soon stops working.
  3. A graceful stop. When access is removed mid-incident, the agent stops, saves its notes and hands the incident to a person, instead of failing silently.
  4. A rehearsal. Practise the revocation before it is needed, as you would test an escalation policy.

A checklist before granting anything new

  • Does the agent need this to do a job it has today?
  • Is there a narrower way to get the same result?
  • Who owns this permission, and when is it reviewed?
  • What is the worst the agent can do with it, and can that be undone?
  • Is the call logged, and can it be revoked in one step?

If the answers are unclear, the permission waits. Read the systems the agent can reach on the integrations page of any tool you evaluate, and compare that list with what your job requires.

Frequently asked questions

Why not just tell the agent not to change anything?

An instruction is a request to the model, and models can misread it, forget it or be steered away from it by text they read. A permission is a fact about the system: if the credential cannot delete, the agent cannot delete. Use instructions to guide behaviour and permissions to bound the damage.

Can an agent do useful work with read access alone?

Yes, and most of the value of incident response is there. Reading metrics, logs, traces, deploy history and code lets an agent triage an alert, form a hypothesis and hand a person the evidence. Writes save time later, and only after the reads have proved trustworthy.

What is prompt injection, and why does it matter here?

It is text inside the data an agent reads, such as a log line or a ticket, that tries to give the agent new instructions. An incident agent reads a lot of text it did not write. Narrow permissions mean that even a successful injection can only reach what the agent is allowed to touch. Watch the outbound path as well: anything the agent can post to chat, a ticket or email can carry data out, so limit where it can write, not only what it can read.

Should the agent use a person's credentials?

No. Give it its own identity, so its actions are attributable, its access can be scoped and revoked without affecting a person, and its calls are separate in the audit log. A shared admin key is the worst option: it hides who did what and grants far more than any task needs.

How do we widen access safely over time?

One scope at a time, each with an owner, a reason and a review date. Start with a write that is easy to undo, watch how the agent uses it, and keep the approval step in place. Widen only when the record shows the agent proposing the right action and people approving it unchanged.