LAB-007SECURITY

Why autonomous systems need independent authorization

PUBLISHED

It’s tempting to let an agent’s reasoning double as its permission check — if the agent concluded an action makes sense, why not let that conclusion also be the approval. The problem shows up the first time the agent’s reasoning is wrong, and there’s nothing else in the system positioned to catch it.

Working through this concretely, the useful distinction is between two different questions that are easy to collapse into one: “is this action a good idea” and “is this action permitted.” An agent is well suited to answer the first — that’s the reasoning it’s built for. It’s poorly suited to answer the second, because answering it honestly sometimes means concluding “no” about an action the agent has already decided, through its own reasoning, that it wants to take. Asking a component to override its own conclusion is a weak control.

Separating the two means the policy check has to live somewhere the agent doesn’t control and can’t reason its way around — evaluated against rules that exist independently of whatever the agent currently believes is correct. In practice that’s a boundary the request has to pass through, with its own allow/deny logic, before anything happens.

What we observed building toward this: the hard part isn’t the boundary itself, which is a fairly ordinary policy check. The hard part is making sure every path an agent could use to take an action actually goes through that boundary, rather than one being added for the obvious path while a less obvious one quietly bypasses it.


← All lab notes