The Autonomy Dial: How Much Authority Should Your Agents Actually Get?
The debate about whether to deploy AI agents is over. Analyst houses now expect 80% of enterprise applications to embed agents by the end of 2026, and the argument inside companies has moved one level up: how much authority does each agent actually get. Can it send the email or only draft it. Can it issue the refund or only recommend it. Can it commit the change or only propose it. That question gets settled ad hoc in most organisations, one nervous meeting at a time, which is why the same company often has one agent that is dangerously over-trusted and another that is uselessly hobbled.
We think about it as a dial, and after thirty-plus production deployments we can say the dial has five positions, and that the right position is a property of the workflow, never of the agent. This piece lays out the five levels, the two variables that decide the setting, and the mechanism that moves an agent up the dial safely.
The five positions
Level 0: Observe
The agent reads and reports. It summarises the inbox, flags the anomaly, drafts nothing that leaves the building. There is no action to be wrong about, only analysis. This is where every deployment should spend its first weeks, because it builds the evidence you need for every later setting.
Level 1: Suggest
The agent proposes, a human executes. It drafts the reply, recommends the price, prepares the journal entry, and a person clicks send. The value is real but capped, because the human is still the throughput bottleneck. Most companies park here permanently out of caution, which is the useless-hobble failure: paying for an agent while keeping the labour.
Level 2: Act with approval
The agent executes after a named human approves each consequential action. The difference from level 1 sounds small and is not: the agent owns the whole workflow and the human owns a checkpoint, rather than the human owning the workflow with the agent as typist. This is the plan-to-act gap from the agent loop, made into a permission level.
Level 3: Act within bounds
The agent executes on its own inside explicit limits, and escalates the rest. Refunds under €200, replies inside approved policy, reorders below a stock threshold. The bounds are written down, the escalation path lands on a named person, every action is reversible or logged. This is where the economics change, and it is where most of our production agents run.
Level 4: Own the outcome
The agent owns the full loop for a defined outcome, with periodic human review instead of per-action gates. Tier-one support resolution, invoice reconciliation, catalogue pricing inside guardrails. Level 4 is earned, never granted on day one, and even here the governance layer stays: observability, a kill switch, an audit trail, a human accountable by name.
The autonomy dial: five levels from observe to own, set per workflow by verifiability and consequence.
The two variables that set the dial
Strip away the anxiety and the setting is decided by two questions we have written about separately, now working together.
The first is verifiability: can the output be checked cheaply, quickly and against a clear standard. That is the verifiability test. A workflow whose output verifies fast can run at a higher level, because wrongness gets caught by the loop itself rather than by an angry customer.
The second is consequence: what does a wrong action cost, and can it be undone. A mispriced wine on a catalogue for an hour is a shrug. A wrong regulatory disclosure is a crisis. Reversible plus cheap tolerates autonomy. Irreversible plus expensive demands a human gate no matter how clever the model.
Plot any workflow on those two axes and the dial sets itself. High verifiability, low consequence: level 3 or 4. High verifiability, high consequence: level 2, the checkpoint earns its latency. Low verifiability, low consequence: level 3 with sampling review. Low verifiability, high consequence: level 1 and honest scepticism about whether an agent belongs there at all, which is the same warning we gave in the 5% problem about pilots pointed at ungradable work.
The ratchet, and why day-one settings are wrong by design
The dial is set per workflow, but it moves with evidence. The pattern that works in production is a ratchet. Every agent starts at level 0 or 1 regardless of how good the demo looked, because the first weeks are for building the track record: resolution rates, error rates, escalation quality, the boring numbers. When the numbers hold for a defined period, the dial moves one position, never two. When an incident happens, it moves back one position, immediately and without debate, while the failure mode gets designed out. The asymmetry is deliberate: slow up, fast down.
The ratchet does something organisational as well as technical. It converts the scary one-time question, do we trust the AI, into a routine operational one, did this workflow earn its next level this quarter. Teams that argue about trust in the abstract stall for months. Teams that review a ratchet table in a monthly meeting move faster and sleep better, and the table itself becomes the artefact a regulator or a board can actually read. It is the same discipline as onboarding a person: nobody gives a new hire the company credit card in week one, and nobody keeps a proven performer on supervised probation for three years, at least nobody who wants to keep them.
The two failure modes the dial prevents
Companies without an explicit dial land in one of two ditches. The over-trust ditch: an agent gets broad authority on launch week because the demo impressed a director, and the first bad Friday produces the incident that freezes the whole programme, the trust collapse we described in the governance gap. The under-trust ditch: every action needs approval forever, the human checkpoint becomes a queue, and the agent quietly becomes an expensive autocomplete while the business case evaporates. Both ditches are the same mistake, a dial set once by feelings instead of continuously by evidence.
The Greek-market angle
The ratchet is easier to run in the organisations we serve than in the multinationals it was designed for. Moving a workflow from level 2 to level 3 in a large matrix means a committee, a risk review and a quarter of calendar time. In a 150-person Greek firm the accountable human, the workflow owner and the decision maker are usually two people who share a corridor, and the monthly ratchet review is a twenty-minute meeting. Smaller firms tend to assume enterprise-grade agent governance is beyond their weight class. On this specific discipline, they are structurally better at it.
Where to start on Monday
Take your live agents, or the one you are about to launch, and write a one-page table: workflow, current level, verifiability, consequence, evidence collected, next review date. Most companies discover in an hour that they have never actually decided any of these settings, they inherited them from launch-week defaults. Deciding them is the whole exercise. The dial costs nothing to draw and repays itself the first time someone asks, why is this agent allowed to do that, and the answer is a row in a table instead of a shrug.
We set the dial with clients as part of every deployment, and the agents we ship (AI Customer Support, AI Contract-to-Cash, AI-Powered CRM and the rest of the product family) come with the levels, the bounds and the ratchet review built in, because authority design is product design. If your agents' permissions were set by vibes in launch week, get in touch at inbusiness.gr and we will draw the table with you.