The automated help that raises confidence without raising accuracy

Manos sobre el mostrador de madera de una caja pasando un marcador detector sobre un billete, con la marca ya trazada y el billete a punto de volver al cliente, mientras la lámpara para mirar la marca de agua queda a un costado sin usarse

The automated help that raises confidence without raising accuracy

The automated help that raises confidence without raising accuracy

The email came from paypal.co.uk, a genuine PayPal domain. Its main button pointed to a paypal.com page that did not exist. Anyone who clicked landed nowhere, which left within reach the only thing in the message that did work, a phone number to call. That was where the fraud lived.

An automated analysis read that email and reached a reasonable conclusion. The action the message asked for was a click, and the click led nowhere dangerous, so the email was clean. The system got the link right and got the email wrong.

In a 2025 experiment with 489 participants, the people who saw that report classified the email correctly 33.1% of the time. The people who had only the usual generic advice, with no analysis at all, got it right 66.9% of the time. They got it right twice as often.

The system being wrong was to be expected, because any probabilistic classifier gets things wrong. What explains the result is what happened around that error. With a report in view, the checking the programme took for granted stopped happening.

In why phishing is not a knowledge problem I wrote that much of this decision-making runs on autopilot, with no deliberation in between. Now there is a new layer sitting on top of that same problem. Something decides before the person does, and says so out loud. The name cognitive psychology gives to what happens next is automation bias.

What is automation bias?

It is the tendency to accept the output of an automated system as a replacement for attentive information seeking, instead of treating it as one more piece of evidence. It shows up in two forms. One is following incorrect advice. The other is failing to act because nothing prompted you.

Both names matter because the two forms do not look the same from inside a programme. Following wrong advice leaves a recorded decision that can be reviewed afterwards. Not looking because nothing prompted you leaves nothing, because a check that never happened produces no data.

The distinction comes from a 2012 systematic review of clinical decision support systems, which kept 74 studies and measured how much wrong advice weighs. When the system erred, the risk of an incorrect decision rose by 26% compared with deciding without it. The ground is clinical and a long way from an inbox, and the stretch of work being delegated is the same, deciding how much attention each case deserves and how much information needs looking up before resolving it.

The tendency is already named in law. The EU AI Act names it in Article 14, on human oversight, when it asks that whoever supervises the system be able “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)”. That article applies to systems classified as high risk, which a mail filter does not fall into by default, so what is useful for an awareness programme is the underlying assumption. The legislator treats the tendency as expected in anyone operating a system, to the point of requiring that it be worked on with that person.

What changes when the analysis is wrong?

On average, accuracy does not move and confidence does. The PayPal email was one of the stimuli in a 2025 experiment that tested three ways of helping people decide, and that is its most uncomfortable result, because it describes a group of people who ended up more certain of their judgements without having got any better at them.

The design was simple. The 489 people judged emails across two rounds, first with no help and then with one of three kinds. One group got generic advice, the sort circulating in any campaign. Another got a report specific to each email, always correct. The third got that same report carrying the errors a real system makes.

All three started in the same place, around 0.78 accuracy on a 0 to 1 scale. With generic advice they rose to 0.83. With the always-correct report, to 0.93. With the realistic report they stayed at 0.78, and their confidence rose from 3.52 to 3.74 on a 1 to 5 scale.

The accuracy difference between the generic advice group and the realistic report group did not reach statistical significance. A specific analysis of every email, with everything it costs to produce, performed no better than the usual advice and left people more certain.

That average hides the real movement, because the system’s errors were concentrated in a handful of emails. On the PayPal one, accuracy fell to 33.1%. On a legitimate email about an internship opportunity, which the report flagged as suspicious, it fell to 22.2% against 50.3% in the generic advice group.

The mechanism shows up in which signal ended up carrying the decision. On the PayPal email, the people who got it right said afterwards that the scam classification was what counted, while those who got it wrong pointed to the link destination, which is precisely the signal the report had misread. The report changes the order in which the email’s signals get looked at, and the checking stops at the one the machine put first.

The experiment measured with a UK panel and showed the report next to each email, when a real client would require asking for it, two limits its own team declares.

Does experience protect you?

It cushions the effect rather than preventing it. The clearest evidence on that point comes from outside security, from a 2023 study in which 27 radiology professionals read mammograms alongside suggestions from an artificial intelligence system.

When the suggestion was correct, all three experience levels scored similarly, between 79.7% and 82.3%. When the suggestion was incorrect, the least experienced fell to 19.8% and the most experienced to 45.5%. Experience halved the damage, and the result was still a collapse.

It is a laboratory study with simulated reading, so it holds as a demonstration of the mechanism rather than a measurement of its size. Even with that caveat it leaves something solid for programme design. If the effect survives years of task-specific training, more task-specific training will not fix it.

In deepfakes and the two opposite ways a programme fails I argued that training perception does not make it more accurate. Here the problem arrives from another direction and ends the same way. Training sharpens the judgement of the person deciding, and it does not change the weight a machine’s assertion carries when it lands before that judgement.

A green flag flying alone on the mast of an empty lifeguard tower, facing a choppy sea with a current running, nobody watching the water

What stops being looked at when nothing is flagged?

It is the part of the bias that leaves no trace, and the question almost nobody asks is what went unchecked on the emails where the system said nothing.

When a filter blocks a message, there is a visible decision that can be argued with. When it lets one through, no decision is visible, because the email simply appears in the inbox. Out of that comes an inference nobody states out loud and which governs behaviour anyway, that if the filter did not warn you the message must be fine.

The system’s silence says far less than that. It says no alert was raised, and that includes the case where the problem was there and went undetected. Automation does not need to assert that an email is safe in order to change what the person receiving it does. It is enough that it says nothing, and with that a check that used to happen falls away.

That is the omission form the 2012 review separates from the other, and the one that reaches no dashboard. In security fatigue I looked at the warning that repeats until it stops producing a response. This is the reverse, because here there is no warning and the missing warning teaches something all the same.

That silence also does something to measurement. A simulation needs to be let through in order to be delivered, so it lands in inboxes where the filter has had its voice switched off, and the rate coming out of that describes a condition unlike the ordinary day. It is another reason to look closely at how a simulation is delivered before reading its click rate.

What the programme has to decide

Three decisions move once you accept there is an automated layer that rules first.

  1. Separate detecting from checking. Classifying an email correctly and independently checking what the system said are two different things, and an accuracy percentage mixes them. The person who reviewed the sender and the person who accepted an automated verdict on a day it happened to be right look identical on the dashboard. The useful reading sits in the cases where that verdict was wrong, because that is the only moment when checking changes the outcome. That checking cannot be observed directly. What can be observed is somebody correcting a wrong recommendation, and that is the most accessible evidence that some independent check took place.
  2. Show the signal the verdict rests on. In the experiment, the people working with error-free reports leaned mostly on concrete signals, such as checking whether the sender’s domain matched the organisation the message claimed to represent. When the automated layer shows only its conclusion, whoever reads it has nothing to weigh it against. When it shows the signal it based itself on, it leaves a claim that can be tested in a few seconds.
  3. Give correcting the machine a visible consequence. When somebody reports something the filter let through and nothing happens afterwards, the filter’s silence weighs a little more next time. That loop, seen from the operational side, is worked through in the reporting habit with no response.

The SMARTFENSE platform settles part of those three decisions out of the box. The phishing report button returns instant recognition when what was reported turns out to be a simulation, so in that case the answer arrives on the spot rather than waiting for somebody to write it. And the reporting hub includes attack response curves showing how fast the organisation falls against how fast it reports, along with the exposure window of each campaign.

That same hub includes Smart Triage, which classifies reported emails by risk level. It is exactly the kind of layer this article has been looking at carefully, so it is worth saying plainly. A classification like that organises the response team’s work, and it works better read as a priority than as a sentence.

An automated layer that is rarely wrong is harder to audit than one that is often wrong, because it almost never gives a reason to doubt it. What the programme does control is how much of that layer’s reasoning stays in view.

Tatiana Stacul

Psicóloga cognitivo-conductual enfocada en comportamiento humano en entornos digitales: estudia cómo la atención, la carga cognitiva y la respuesta emocional al riesgo condicionan la toma de decisiones frente a la pantalla. Colabora con SMARTFENSE en el diseño de contenidos de concienciación en ciberseguridad y divulga sobre ciberpsicología y bienestar digital en Código Calma. Forma parte de Women4Cyber Sweden y Cibervoluntarios.

Leave a Reply