Why a “repeat-clicker” list says more about the lures than about the people
Two simulated emails went out from the same organisation, to the same employee population, inside the same eight-month programme. One warned that an Outlook password was about to expire. It was failed by 1.82% of the people who received it. The other announced a change to the vacation policy. It was failed by 30.80%.
Same people. One email turned out to be nearly seventeen times more deceptive than the other.
The figure comes from a controlled experiment published in 2025 covering 19,789 employees of a large healthcare organisation, who received ten different lures over eight months. And it has an uncomfortable consequence for any dashboard that ranks people by how many times they clicked.
If the outcome depends that much on the email, then part of each person’s score does not belong to that person.
What does a repeat-clicker list actually measure?
It measures two things added together: what somebody did, and how hard the thing in front of them was. The dashboard hands them over fused and offers no way to separate them.
The table in that study ranks the ten lures by failure rate and the full range runs from 1.82% to 30.80%. The variation is not explained by the topic. Two messages from the same family, both about changes to benefits, produced 7.62% and 30.80%.
The authors draw a methodological warning from this that is worth quoting as they wrote it, because it is not my reading of the data. They note that a randomised comparison and a careful statistical analysis are needed to evaluate whether a change in a user’s failure is due to the training they received or to other confounding factors, such as subsequently receiving weaker phishing messages that are easier to identify.
Translated into programme operations: when somebody “improves” between two campaigns, the available data does not say whether the person improved or the difficulty dropped. And when they get worse, it does not say either.
How to build a score that survives that objection is a data engineering problem, and it is covered in what a reliable human risk score needs. What interests me here is what happens afterwards, when that number turns into a name on a list.
Why does clicking twice say less than it appears to?
Because the repeat-clicker label depends on how many chances there were to earn it.
In the same study, 9.7% of the population had failed at least once by the end of the first month. By the eighth, 56% had, which is 11,077 people out of 19,789. A further 25.9% failed at least two simulations, 9.8% at least three and 3.5% at least four. One person failed every single one.
The conclusion the authors reach is that users who avoided the earlier emails will not necessarily avoid the later ones. With enough time and effort, they say, an attacker would likely fool a large fraction of an organisation’s employees.
A second figure shows the same thing from another angle. A 15-month study at one company, with 14,733 participants and eight simulations each, found that 1,448 people clicked twice or more. That number is 30.62% of those who had already clicked at least once, not of the total. The distinction matters, because both ways of counting circulate as if they were the same and they produce very different headlines.
That is where the cut-off falls apart. What the list separates out are the people who accumulated unlucky exposure first, and that looks very little like a stable group of susceptible people.
Which bias makes the dashboard read like a trait of the person?
Correspondence bias: the tendency to infer a stable disposition in somebody from behaviour that the situation explains just as well.
It is one of the sturdiest findings in social psychology and it was first demonstrated in 1967, with a design that looks a lot like what a dashboard does. A group of people were given an essay to read, either supporting or opposing the Cuban regime, and were asked to estimate what the writer actually believed. Even when told that the essay’s position had been assigned, and that the writer had chosen nothing, they still attributed the opinion to them.
The 2018 large-scale replication repeated the design with a different debate topic and obtained the effect again across 7,197 participants, at a size of 1.82 on Cohen’s scale, where 0.8 already counts as large. Knowing that the situation compelled the behaviour did not stop it being read as character.
The problem with a repeat-clicker dashboard lies in how it gets interpreted. A record of behaviours gets read as a record of people, and that reading happens even when the person looking knows the lure was a hard one.
What happens to the programme when the label becomes the intervention?
The 15-month study measured exactly that and the result went the other way. Participants who were shown the training page immediately after clicking ended up with higher click and dangerous-action rates than those who were not.
The 2025 experiment measured the same thing with a different design and found an effect in the right direction, though a very small one. The absolute difference between the control group and the trained groups across all emails was 1.7 percentage points. That same work also found no relationship between having recently completed the mandatory annual training and failing less often.
Both results leave teaching on the table and narrow down where. The usual circuit, flagging whoever clicked and showing them a page at that moment, is the part of the programme with the least evidence behind it, and in one case with evidence against it.
There is a limit worth declaring about what the label does to the person receiving it. What has been measured is the effect on later behaviour. The internal mechanism that explains it falls outside the scope of both studies. Whether the effect comes from embarrassment, from rushing to close the window, or from having learned that the consequence of clicking is a two-minute formality, the available evidence does not settle.
What is clear is the asymmetry of costs. A label wrongly placed on somebody who received hard lures costs more than the number it corrects, because the channel the programme needs most is the voluntary one. The habit of reporting with no response covers what happens to that channel when the organisation gives nothing back.
What can you do with that list without turning it into a label?
The list is useful, under three conditions you can check before using it.
- Record the difficulty of the lure alongside the result. Without that field, a person’s number is not comparable even against their own number in another campaign. It is the requirement that enables everything else, and it is usually missing.
- Compare each person against their own history, under equivalent lures. Ranking colleagues against each other mixes the two variables and returns an order that depends on what each one happened to get. Somebody’s own series, measured at comparable difficulty, is a far cleaner signal.
- Make the consequence of appearing assigned content rather than a mark. The operational difference is who sees it and what for. A plan that changes because that person is missing a topic is not the same as a condition that follows them from campaign to campaign.
The SMARTFENSE dashboard is a good place to apply the argument, because it has both pieces. The reports hub includes a 360° user profile combining risk, resilience, simulations, training and proactive activity, and also a prioritised list of users to act on. That list is exactly the object this article is about, and the third condition is what decides whether that list helps or gets in the way.
The way out lies in how the intervention is resolved. Smart Path evaluates each person once a day and assigns them content without repeating what they have already done, within a monthly budget of minutes. The criterion is what each person is still missing from the plan. It decides person by person rather than by group, which is the way of personalising that does not need to label anybody.
One question stays open, and it belongs to product design more than to psychology. The difficulty of a lure does get measured and published. The 2025 study reports it campaign by campaign, and any programme that reuses a lure already has it to hand. What needs checking is whether that figure travels attached to each person’s record, which is what would make two results comparable. It is worth asking of any dashboard, ours included.
Which indicators to watch instead of the click rate was already covered by Nicolás in is your awareness programme measuring what matters?, and the variable none of those indicators captures is the one I worked on in the variable that predicts secure behaviour. This piece deals with the step before both. Before asking which indicator to look at, it helps to know who the one we already have is talking about.
A dashboard that ranks people by how many times they clicked is describing, without saying so, the catalogue of emails it chose to send.
Leave a Reply