Security
Do phishing simulations work? The evidence says barely
Do phishing simulations work? Almost everyone in security behaves as if the answer is obviously yes. Compliance frameworks expect them, cyber insurers ask about them, and a large awareness-training industry depends on them. Yet the best evidence we now have says the honest answer is barely, and whoever you run IT for, an MSP's clients or your own colleagues, that should change where your hours and budget go.
The result the industry does not want to discuss
In 2025, researchers at UC San Diego Health published one of the largest real-world tests of phishing training ever run: a randomised trial across roughly 19,500 employees, with ten simulated campaigns over eight months. The headline finding was blunt. There was no significant relationship between having recently completed the organisation's annual security training and the likelihood of falling for a phishing email. The embedded lessons that appear after someone clicks a test lure moved failure rates by about 1.7 percentage points. That is the entire measured benefit of the thing your whole programme is built around.
The behavioural detail explains why. More than three quarters of staff spent under a minute on the training material, and between a third and a half closed the page almost immediately. People are not learning from this content because they are not reading it, and no amount of gamification has changed that.
This was not a one-off. A 2021 ETH Zurich study followed around 14,000 employees of a large company for fifteen months and found that embedded training after a failed test did not reduce future clicks, with some evidence it made users more likely to fall for later attacks, possibly by leaving them feeling protected. Two large, careful, independent studies, years apart, pointing the same way.
Someone always clicks, and the maths does the rest
Even if training worked better than it does, simulations would still be aimed at the wrong number. A phishing programme tries to lower your average click rate. An attacker does not care about your average. They need one person, on one bad day, reading one well-timed email on a phone between meetings.
Suppose training gets your click rate down to a widely envied three percent. Across 200 users, one campaign still lands six footholds. Run a campaign a month and the question is not whether someone clicks but how quickly. The UC San Diego data shows exactly this compounding: in the first month about one in ten employees clicked a lure, but by month eight more than half had clicked at least once. Individually rare events, repeated across enough people and enough emails, become a certainty.
You cannot train your way out of that arithmetic. Every user you support will eventually click something. A security posture that depends on that never happening is not a posture, it is a hope.
A click rate measures your lure, not your people
Here is the quieter problem with phishing simulations: the metric they produce is adjustable at will. Send an obviously clumsy lure and your click rate looks wonderful. Send a convincing one about a bonus scheme or a missed parcel and it looks terrible. The number the board sees is mostly a property of the test, not of the workforce.
That makes the click rate almost uniquely bad as a security KPI. It can be steered up to justify budget or down to show progress, and neither movement tells you anything about whether a real, targeted attack would succeed. A metric you can set to whatever you need it to be is not a measurement. It is theatre with a percentage sign.
The same distortion runs through the industry's favourite statistics. Vendor case studies showing dramatic improvement almost always compare an easy later campaign with a hard early one, or measure users who were about to improve anyway. The controlled studies, the ones with randomisation and a control group, are the ones that keep finding almost nothing. When the rigour goes up, the effect goes away. That pattern has a name in every other field, and it is not a flattering one.
The costs are real even if the benefit is not
None of this would matter much if simulations were free. They are not. There is the direct cost of the platform and the thousands of staff-hours spent on lessons the data says are skimmed in under a minute. There are the trust costs: employees learn that IT sends them traps, legitimate emails get reported and sit in quarantine while someone in finance chases an unpaid invoice, and the fake-bonus-email genre has produced public HR disasters that outlived any security benefit.
The biggest cost is the least visible. A phishing programme lets an organisation feel it has dealt with phishing. That feeling absorbs budget, attention and board goodwill that should have gone to controls which actually determine whether a click matters. Training is not just weakly effective. It is a sedative.
Assume the click: the controls that actually stop phishing
The alternative is not resignation, it is engineering. Design the estate so that a click is survivable, and the tail-risk arithmetic flips to your side:
- Phishing-resistant MFA. Passkeys and FIDO2 keys cannot be tricked into signing in to a fake site, and they do not push approval prompts that tired thumbs accept. This one control removes the value of most credential lures outright.
- Close the tenant's side doors. Legacy authentication, unrestricted OAuth app consent and silent forwarding rules are how a phished password becomes a persistent breach. Our Microsoft 365 security checklist covers the ten settings to fix first.
- Patch the software the click lands in. A malicious page or attachment needs a vulnerable browser or reader to become code execution. Keeping those current is unglamorous and hugely effective, and as we have written before, third-party patching is harder than Windows Update and needs its own mechanism.
- Least privilege. A payload that runs without local admin rights lands in a puddle, not a lake.
- Fast reporting and response. A one-click report button, and a desk that can pull a lure from every mailbox in minutes. Time-to-quarantine is a real security metric. Click rate is not.
Every item on that list is boring, measurable and indifferent to human error. None of them can be skimmed in under a minute, forgotten by Friday or quietly closed like a training tab, because none of them lives in anyone's head. That is precisely why each is worth more than another round of training videos, and why the teams with the calmest phishing story are usually the ones that spent their awareness budget here instead.
A fair test: if your organisation spends more per year on awareness training than on making MFA phishing-resistant and keeping browsers patched, your phishing budget is upside down. Fix the controls first; they work on the day someone clicks, which training demonstrably does not.
What simulations are still good for
This is an argument for demotion, not abolition. Keep a light simulation programme, but change what it measures. Report rate and time-to-report are worth tracking, because a workforce that forwards suspicious emails quickly genuinely shortens incidents. An occasional exercise that tests your response pipeline, how fast the lure was reported, triaged and purged from mailboxes, tells you something real about your defences. And never publish league tables or punish clickers: the moment reporting feels risky, your best early-warning system goes silent.
So, do phishing simulations work as a way to stop people clicking? Barely, and the largest studies we have say so plainly. Treat them as a smoke detector test for your response process, spend the reclaimed money on controls that assume the click, and your next real phish will meet an estate that does not need anyone to be perfect.
Where this fits with Helios
Helios is built for the assume-the-click side of this argument. The platform continuously checks the controls that decide whether a phish matters: protection-gap detection flags devices where Defender or another AV is missing or unhealthy, third-party patching keeps browsers and readers current alongside Windows, and the Microsoft 365 integration surfaces risky tenant settings before an attacker finds them. When users do report something suspicious, Helio's auto-triage helps the desk separate the real lure from the noise quickly, which is the metric that actually counts.
Harden the estate, not just the inbox
Helios is an AI-native platform for MSPs and in-house IT teams: monitoring, patching, security and service desk in one place, with a 14-day trial and no feature gating.
Start free