Skip to content

Where to put agent approval gates in an event data workflow

AI for organisersUpdated 2026-08-238 min read

In short

Agent approval gates are checkpoints where an autonomous step stops and waits for a named person before it proceeds. Put them on database writes and on any action that reaches someone outside your organisation, such as an email or an invoice. Leave read-only queries ungated, because gating reads trains people to approve everything.

An agent built to help the sales team through renewal season does exactly what it is told. It reads the renewal risk scores, pulls last year's rebooking history, drafts 200 tailored notes to exhibitors who have not yet signed for the next edition, and sends them. Two land on accounts that cancelled after a dispute. One lands on a company acquired by an exhibitor in the same hall, addressed to a brand that stopped existing in March.

The drafting was good. The sending was the problem, and the design question that follows is where agent approval gates belong so the first keeps happening and the second stops.

Gate on writes and on egress, leave reads alone

The rule I would defend is a two line taxonomy.

An action is gated if it changes stored state, or if its output reaches a person outside your organisation. Everything else runs.

That puts warehouse writes, CRM field updates, badge record edits, emails, exhibitor portal messages, invoices, published pages and file exports on the gated side. It puts every SELECT against the warehouse, every aggregate computation, every draft written into a queue and every internal preview on the ungated side.

The reason for drawing it there rather than somewhere more cautious is that gates have a cost measured in reviewer attention, and attention spent on a harmless read is attention unavailable for a batch of 200 outbound emails. A team that has to approve the agent's queries learns within a week that approving is the default, and carries that habit into the approval that mattered.

What does OWASP call this failure?

OWASP lists Excessive Agency as LLM06:2025 in its Top 10 for LLM Applications, and its diagnosis is more precise than the usual worry about runaway autonomy. The entry names three root causes: excessive functionality, excessive permissions, and excessive autonomy. Those are separate faults with separate fixes, and most event data agents fail on the middle one.

Excessive functionality is a tool that can do more than the task needs. OWASP's example is giving an agent a general shell or file-writing capability when a "specific file-writing extension that only implements that specific functionality" would do.

Excessive permissions is the database connection. If an agent's job is to summarise exhibitor performance, it needs SELECT on a handful of fact and dimension tables. It very often gets the credentials of the application, which can write to everything, because that connection already existed. OWASP's recommendation on this is least privilege applied literally: read-only access when only data retrieval is needed.

Excessive autonomy is the missing approval. OWASP recommends explicit user confirmation before high-impact actions, and names deletions, financial transactions and external communications as the categories. An exhibitor email is an external communication, which puts the 200 renewal notes squarely inside the case the control exists for.

The reason to hold all three separately is that a gate is the most expensive of the three fixes and the least reliable, because it depends on a person. Narrowing the tool and narrowing the permission cost one afternoon each and never get tired. Put the gate where the first two cannot reach.

A worked example: 200 renewal notes that go nowhere

Rebuild the renewal agent with the taxonomy applied and count the gates.

The agent runs on a Monday. It queries renewal risk, rebooking history and contracted square metres for every 2026 exhibitor without a 2027 contract. That is three reads, ungated, logged, running under a role with SELECT on four tables and nothing else. It returns 214 accounts.

It filters out accounts flagged as in dispute and accounts whose parent company changed in the last twelve months. That is 14 accounts removed, leaving 200, and the filter runs before drafting rather than after, so the two cancelled accounts and the acquired brand never reach a draft at all. This is worth stating plainly: the cheapest gate is a filter that runs earlier, and a good deal of what teams try to catch with human approval is a WHERE clause they have not written yet.

It drafts 200 notes into a queue table. Writing to a queue is a write, so by the taxonomy it should be gated, and here I would carve out the exception deliberately. A queue that exists only to hold drafts and has no consumer other than the approval screen is not really stored state in the sense that matters, because nothing downstream reads it until a person releases it. Gate the release rather than the insert.

The sales director opens one screen. It says 200 notes, 214 accounts considered, 14 excluded and why, total contracted value represented, and a sample of five drafts with their source figures. She approves the batch, pulls two accounts out by name because she spoke to them on Friday, and 198 notes send.

One approval decision. Six seconds of clicking on top of maybe four minutes of reading. Compare that with 200 individual approvals at even ten seconds each, which is thirty-three minutes of work that produces a worse decision, because nobody holds 200 shallow reads in their head well enough to notice that eleven of them share a wrong figure.

How should a batch gate handle the odd one out?

Per batch approval with per item rejection, and a hold rule.

The batch screen gives the reviewer the shape of what is about to happen: how many items, how many accounts were considered and excluded, the aggregate figures the drafts rest on, and a sample large enough to be worth reading. Any individual item can be pulled from the batch without blocking the rest.

The hold rule is the part that earns its keep. If rejections in a batch exceed a threshold you set in advance, say three items or two per cent, the remaining items stop and the run is marked for investigation. Three independent bad drafts in a batch of 200 is unlikely. Three drafts sharing an upstream fault is the ordinary case, and the whole point of catching it at item three is that items four through 200 have not gone out yet.

That threshold is also a genuine number to report. Batches held per month, and the reason each was held, tells you more about the health of the pipeline than any accuracy figure you could put in a slide.

Gates that should exist and usually do not

Two more places deserve a stop, and neither is an approval in the usual sense.

The first is a spend ceiling per run. An agent looping over exhibitors with a retry on failure can consume a great deal more than anyone budgeted, and OWASP treats that case seriously enough to list Unbounded Consumption separately as LLM10:2025, with "Denial of Wallet" named among its attack vectors. A hard cap that aborts the run and pages someone is cheaper than a monthly invoice that starts a meeting.

The second is a volume ceiling on egress. If a run would send more messages than any previous run of that type, stop and ask. The failure it catches is the one where a join fans out and 200 notes become 6,400, and it catches it before the first message rather than after the last.

Both of those are conditions rather than screens, which is why they get skipped. They are also the two that would have saved the most embarrassment in the incidents I have seen, and the approval screen design work in Q21 assumes they already exist.

Writing the gate map down

The artefact worth producing is dull and short. One row per agent action, with the tool it calls, the permissions that tool holds, whether the action writes, whether the action reaches outside, and the gate if any.

Doing it exposes the mismatches immediately. You find the summarisation tool holding a write-capable connection. You find an export step that drops a CSV into a shared drive with no gate, which is egress that nobody classified as egress because it has no send button. You find two agents sharing a credential.

NIST's AI Risk Management Framework, published as NIST AI 100-1 in January 2023, organises this kind of work under the first of its four functions, Govern, alongside Map, Measure and Manage. The framework's own line on why the documentation matters is worth quoting: "Trustworthy AI depends upon accountability. Accountability presupposes transparency." A gate map is the cheapest transparency artefact available, and applying the framework properly across a stack is Q35's subject.

Where this stops

Gates control what the agent does. They say nothing about whether what it did was right, and they create a specific new risk that is easy to miss.

Every gate is a place where a person's approval becomes the record. Once a batch is approved, the approval is what the audit trail holds, and if the reviewer was rubber-stamping, you have manufactured evidence of a review that did not happen. A workflow with gates and inattentive reviewers is in a worse evidential position than one with no gates at all, because it now carries a signature. Q23 covers how to tell whether the approvals are real, and the honest answer is that you cannot tell without seeding errors and counting the catches.

The other limit is that batching trades depth for coverage on purpose. A reviewer approving 200 items after reading five has, by construction, not read 195 of them. That is the right trade when the 200 are generated by one process from one query, and it is the wrong trade when the batch is heterogeneous. If your queue mixes renewal notes, invoices and exhibitor performance reports, batch approval is approving three different risks with one click, and the fix is to split the queue rather than to add a longer sample. Whether the batch even needs regenerating is a separate saving, covered by caching on the aggregate hash in Q26.

Start this week by running one query against your own stack: list every database role an agent or AI service currently authenticates as, and for each one, count the tables it can write to. If any of those counts is greater than zero for a service whose job is summarising, you have found a permission fix that takes an afternoon and removes a class of incident no gate can. The wider picture of how these systems are put together is worth reading alongside it.

Questions people ask about agent approval gates

Which agent actions need a human approval gate?
Anything that changes stored state and anything that reaches a person outside your organisation. That covers writes to the warehouse or CRM, emails and messages to exhibitors or attendees, invoices, published pages and file exports. Read-only queries against the warehouse need logging and rate limits, and gating them adds clicks without reducing any real risk.
What is excessive agency in the OWASP list?
LLM06:2025 Excessive Agency in the OWASP Top 10 for LLM Applications describes harm caused when a model-driven system holds more functionality, permissions or autonomy than the task needs. OWASP names those three as the root causes and recommends least-privilege extensions, read-only database access where retrieval is all that is needed, and explicit confirmation for high-impact actions.
Should approval be per item or per batch?
Per batch, with per item override. A reviewer approving 200 items one at a time is doing 200 shallow reads. A reviewer approving a batch after inspecting a sample and a summary of what the batch will do is making one considered decision, and any item can still be pulled out and rejected on its own.

Related reading

All ai for organisers articles