Awareness validation questions: what they actually measure and what they don’t

Al amanecer, una persona arrodillada mide con una cinta métrica la sombra alargada que proyecta una gran estructura, mientras la estructura misma se alza detrás, en penumbra.

Awareness validation questions: what they actually measure and what they don’t

Awareness validation questions: what they actually measure and what they don’t

At the end of nearly every awareness module there is a question. The person answers it correctly, the system records a pass, and the program adds one more data point to its dashboard. That data point later travels to a management report under a reassuring heading, the one for training results. The trouble starts when someone mistakes that number for what they actually wanted to know, which is whether the organization decides better in the face of risk. A validation question measures something real, but it rarely measures that.

The gap between answering well and acting well comes from how the mind works. It is worth looking closely at what falls within a quiz’s reach and what falls systematically outside it, because that reading is what determines whether an awareness program takes its own measure with honest data or with a mirage.

What does a validation question actually measure?

A validation question measures whether the person can retrieve or recognize the correct answer at the moment they are asked. That is all, and it is not nothing. It verifies that the content left an accessible trace, that faced with the prompt the person can tell the right option from the rest.

The point is that this act happens under very particular conditions. The person knows they are being tested, the topic is fresh because they have just seen it, their full attention is on answering, and the options already come laid out on the table. None of that resembles the moment a message with feigned urgency arrives in the middle of a task. The question captures a performance in an exam setting, not a stable disposition to behave a certain way when no one announces that this is the test.

Why does passing a quiz fail to predict behavior?

Because knowing and doing are governed by different systems. A person can flawlessly identify the signs of a fraud in a test and fall for an equivalent one weeks later, when their attention is elsewhere. We have seen it already in this same series, where we argued that phishing is not a knowledge problem, but a matter of the conditions under which someone decides in front of the screen.

There is also a more technical reason, and it has to do with transfer. What is learned in one context tends to stay anchored to that context, and shows up only with difficulty in situations that do not resemble it. The quiz is a classroom context: an explicit prompt, closed options, a clear signal that thinking is required. The moment of risk is the opposite. That is why strong performance on the first tells us little about the second, and why content designed to change behavior works on transfer, not just on retention.

Reading a pass as proof of safe behavior is to commit, unintentionally, the same error made when misreading the results of a simulation, which is taking a narrow indicator for the whole phenomenon.

What can well-designed questions actually measure?

A validation question stops measuring memory and starts measuring understanding when it forces the person to apply judgment rather than recognize a fact. The difference lies in the design. A question that asks for the definition of phishing tends to stay in the realm of recall, though good design can take it further. Presenting an ambiguous message and asking the person to decide how they would act and why leans toward testing judgment, one step closer to real behavior. What each format measures in practice depends on how it is built.

Scenario-based questions measure that ability to reason within a plausible situation. They do not guarantee the person will act the same way in their own inbox, but at least they probe the same kind of mental operation the moment of risk demands: reading signals, weighing them, resolving under some ambiguity. A question that merely asks the person to retrieve a definition never reaches that ground.

There is a second, less obvious value, and it is that answering is not a passive act. Cognitive psychology describes the testing effect, studied by Henry Roediger and Jeffrey Karpicke. Retrieving something from memory strengthens that trace more than reading it again. A well-crafted question, then, does not only measure; it consolidates. It stops being a toll at the end of the module and becomes part of the learning. That is its legitimate role: to check understanding, surface mistaken ideas, and cement what matters through the very act of recalling it.

What should a CISO report when they present “training results”?

Completion rate and average score are metrics of activity and knowledge, not of risk. They say the content was delivered and that, under exam conditions, it was understood. They are legitimate and worth reporting for what they are, without asking them to answer a question they cannot answer. Presenting them as evidence that the organization is safer inflates an expectation the first incident takes care of disproving.

The evidence of behavior lives elsewhere. It appears when the person faces a realistic stimulus without knowing it is a test, and that is where a well-designed simulation measures behavior in the context that matters. The fine detail of which behavioral indicators are worth tracking, and how to read them without falling into new mirages, deserves its own treatment. What matters to settle here is the boundary. A quiz score measures that the answer was known, not that the right thing will be done. An honest report keeps those two things in separate columns.

How SMARTFENSE works with this

On the SMARTFENSE platform, validation questions live inside the exams and play the role that belongs to them, which is to check that the content was understood and to reinforce it through retrieval, not to stand in for behavior. When the design calls for it, they lean on scenarios instead of definitions, to move closer to the kind of judgment the real world demands.

Measuring behavior runs on a separate track. The simulation tools bring the person to decide under conditions close to the real ones, and the reports make it possible to show each thing in its place, with what was understood on one side and what was done on the other. Keeping them apart is not a methodological detail; it is the condition for a program to truly know where it stands.

A validation question is a good tool as long as it is asked for what it can give. Awareness gains precision the day it stops asking how many passed and starts asking what each number it reports is measuring.

Tatiana Stacul

Psicóloga cognitivo-conductual enfocada en comportamiento humano en entornos digitales: estudia cómo la atención, la carga cognitiva y la respuesta emocional al riesgo condicionan la toma de decisiones frente a la pantalla. Colabora con SMARTFENSE en el diseño de contenidos de concienciación en ciberseguridad y divulga sobre ciberpsicología y bienestar digital en Código Calma. Forma parte de Women4Cyber Sweden y Cibervoluntarios.

Leave a Reply