Automation bias in AI review and how to keep approvals honest
Automation bias in AI review is the tendency to accept a system's output because the system produced it. Parasuraman and Manzey found in 2010 that it causes both omission and commission errors and resists training. Measure it by seeding known-wrong drafts into the review queue and counting how many reviewers catch.
The number that should worry an event data lead is a 98 per cent approval rate on AI drafts, and the reason it should worry them has nothing to do with the drafts. Automation bias in AI review is what happens when a reviewer's job quietly changes from checking the output to confirming that output exists, and the approval log looks identical either way.
Work the shape of it through with round numbers. A season produces 1,240 AI-drafted exhibitor summaries, all of them routed through review, and 22 come back rejected. Somebody later hand-checks 40 of the approved ones and finds 3 carrying a figure that does not match its source. Apply that share to the 1,218 approvals and you get roughly 90 wrong figures reaching exhibitors against 22 rejections. Nothing in that story requires a lazy reviewer. It works perfectly well with experienced show operations people who take the job seriously, which is what makes it worth taking seriously in return.
What does the research on automation bias actually say?
Parasuraman and Manzey published a review of the empirical work in Human Factors in 2010, and its findings are more awkward than the usual summary suggests.
On complacency, they report that it "occurs under conditions of multiple-task load, when manual tasks compete with the automated task for the operator's attention", that it appears in both inexperienced and expert participants, and that it "cannot be overcome with simple practice". On automation bias, they report that it "results in making both omission and commission errors when decision aids are imperfect", that it "cannot be prevented by training or instructions", and that it affects teams as well as individuals.
Read those two sentences slowly, because they close off the two responses every organisation reaches for first. Training the reviewers does not fix it. Telling them to be careful does not fix it. Adding a second reviewer does not reliably fix it either, since the effect operates on teams.
What the same review points at is the mechanism, which is attentional. Complacency shows up under multiple task load, when the manual work competes with the monitored task. That is a description of a show operations team in the fortnight before doors open, which is exactly when the AI-drafted output volume peaks.
The EU AI Act names the same failure. Article 14 of Regulation (EU) 2024/1689 requires that human oversight measures let the overseer "remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias), in particular for high-risk AI systems used to provide information or recommendations for decisions". The Regulation says to remain aware of it without saying how, which is the gap this post is about.
Seed known errors and count the catches
If training cannot fix the effect, the remaining move is measurement, and the measurement has to be adversarial because a reviewer who knows which drafts are being checked reviews those drafts differently.
The method is old, dull and effective. Insert a small number of deliberately wrong drafts into the ordinary review queue, indistinguishable from the real ones, and count how many get rejected.
Five seeds a month is a reasonable starting rate. They should be varied by error class, because catch rates differ enormously by class and an aggregate figure hides that. Four classes worth seeding:
- A wrong figure with plausible magnitude. Scans reported as 486 when the source says 468. A transposition, in range, checkable only against the panel.
- A right figure with the wrong period. This year's number described as last year's, or a three day total compared against a four day one.
- A confident claim with no source aggregate behind it. A sentence asserting that international attendance improved when no international field was in the input at all.
- A wrong direction. Down described as up, where the operands are present and the word contradicts them.
Run the seeds through the same queue, the same screen and the same reviewers as everything else. Record the catch by class.
The arithmetic on a quarter of seeding
Five seeds a month across three months is fifteen, and fifteen is a small number, so be honest about what it supports.
Suppose over that quarter the reviewers catch twelve of fifteen. That is a catch rate of 80 per cent. With fifteen trials the interval around that estimate is wide, roughly 55 to 93 per cent at conventional confidence, so 80 per cent should be read as "somewhere between half and nearly all" rather than as a precise figure. It is still worth having, because the interesting comparison is against zero measurement and against the class breakdown.
The class breakdown is where it earns its cost. Say the twelve catches split as four of four on wrong direction, four of four on the unsourced claim, three of four on the wrong figure, and one of three on the wrong period.
That pattern is legible. Reviewers catch contradictions and inventions almost every time, because both are visible in the prose alone. They catch a transposed digit most of the time, because the panel puts the source value next to the claim. They miss period mismatches, because checking a period means comparing two date ranges that look similar and neither one is wrong on its own.
The response to that is a design change rather than a training session. If period mismatches are the class getting through, put the period comparison on the screen as a single derived field that reads "same period" or "3 days vs 4 days", and the check stops being an act of attention. That is the general shape of the fix: every class with a poor catch rate is a candidate for being computed rather than noticed, and the layout of the approval screen in Q21 is where that computation goes.
What else in the log tells you a review is real?
Catch rate is the direct measure. Three indirect ones are free, because your approval log already holds them. All three assume the gate is somewhere sensible in the first place, which is Q22's territory on where an approval step belongs.
Time per item, computed as the elapsed gap between consecutive approvals within a batch. A queue cleared at under five seconds an item has not been read, whatever the reviewer believes. Watch the distribution rather than the mean, because a handful of long pauses will drag a mean up over a run of instant clicks.
Edit rate. The share of approved items where the reviewer changed the text before approving. A reviewer who edits nothing over hundreds of items is either working with an excellent model or has stopped engaging with the prose, and the seeded errors tell you which.
Rejection reasons, if your screen collects them. A queue where every rejection reason is the same one is a queue where only one class of error is visible.
None of the three is proof on its own. Together with a catch rate they give you a defensible answer to the question a finance lead will eventually ask, which is how you know the review step does anything.
Do not fix this by demanding rejections
There is a failure mode in the response worth naming, because I have watched it happen.
Once approval rate becomes a monitored metric, people manage it. Reviewers start rejecting borderline items to keep their number away from 100 per cent, drafts go back for cosmetic edits, and the queue grows. You have replaced an unmeasured review with a theatre of review, and the seeded errors will still get through because rejecting a good draft and catching a bad one are unrelated skills.
Approval rate is a diagnostic to look at alongside catch rate, never a target to hold anyone to. A team catching four of five seeds while approving 96 per cent of real drafts is doing the job correctly, and a team approving 80 per cent while catching one of five is doing it badly at greater cost. Which model produces drafts good enough to justify a high approval rate is a separate question, settled by grading candidates on your own eval set in Q27.
Where this stops
Seeding measures the reviewers against errors you already know how to make. It says nothing about the errors you have not thought of, and by construction it cannot.
The four classes above came from failures somebody had already seen. A model that begins making a fifth kind of mistake, subtler than any of them, will pass through a queue whose catch rate looks healthy, and the catch rate will keep looking healthy because it is measured against the old four. Refresh the seed classes whenever you change the model or the prompt, and treat any drop in catch rate as a signal about the seeds as readily as a signal about the people.
There is also a limit that comes from Parasuraman and Manzey directly. Their finding is that these effects arise from the interaction of personal, situational and automation-related characteristics, with attention at the centre. Seeding does not remove the effect. It measures it, and it gives you a reason to move a check from a person's attention into the interface, which is the only intervention their review supports. Anyone hoping that a quarterly briefing on the risks of over-reliance will do the work should read the 2010 paper and give up on that plan.
The first step is one you can take before you build anything. Take the last fifty AI-drafted items your team approved, pick ten at random, and hand-check every figure in them against its source table. Whatever share comes back wrong is your current error rate reaching the reader, measured on your own output, and it took an afternoon. If you want the wider context on how these systems are built and reviewed, the pillar covers the surrounding decisions.
Questions people ask about automation bias in ai review
- What did Parasuraman and Manzey find about automation bias?
- Their 2010 review in Human Factors reported that automation bias produces both omission and commission errors when a decision aid is imperfect, appears in expert and inexperienced participants alike, and cannot be prevented by training or instructions. They also found that automation complacency arises under multiple task load and cannot be overcome with simple practice.
- Is a high approval rate always a sign of automation bias?
- No. A mature pipeline producing good drafts should have a high approval rate, and demanding rejections would be worse. The signal is a high approval rate with no measurement behind it. Once you know your reviewers catch four of five seeded errors, a 96 per cent approval rate is evidence the drafts are good rather than evidence nobody is looking.
- How many seeded errors should go into a review queue?
- Enough to give a usable rate and few enough to avoid poisoning the queue. Five a month against a few hundred reviewed items gives a catch rate with a wide interval, which is why the number to watch is the running total over a quarter. Fifteen seeds and twelve catches is a defensible eighty per cent.
Related reading
- Designing human in the loop approval that people actually read
- Where to put agent approval gates in an event data workflow