Behavioral security rests on a simple promise: watch what people and accounts normally do, then flag the departures. Two large field studies, five years and a continent apart, put the human half of that promise to a controlled test. ETH Zurich ran 14,733 employees of a Swiss public company through 15 months of simulated phishing. UC San Diego Health randomised more than 19,500 staff over eight months. Neither found what the training market sells, and one found the opposite. The machine half of the promise has never been benchmarked in public at all.
Embedded Training Made Clicking Worse in a 14,733-Person Trial
Daniele Lain, Kari Kostiainen and Srdjan Capkun of ETH Zurich reported their results at the IEEE Symposium on Security and Privacy in 2022. The setting was a large Swiss public company of more than 56,000 staff spanning logistics, finance, transport and IT services, and the study ran from July 2019 to October 2020, sending 117,864 simulated phishing emails into normal working traffic rather than into an announced exercise.
The base rates alone are worth having. There were 6,680 clicks across 117,864 emails, a rate of 5.67 percent, and 4,729 participants, 32.10 percent of the population, clicked at least once. Dangerous actions, meaning entering credentials or the equivalent, numbered 4,885, which is 4.14 percent of emails sent and 73.13 percent of clicks, and 3,747 participants took one at least once. Repetition was common: 1,448 people clicked more than one simulated phish, 896 took a dangerous action more than once, and one participant fell for six of eight simulations.
The counter-finding is the one the industry has not absorbed. Employees who received contextual embedded training, the teachable-moment page shown immediately after a click, performed worse than those who did not. Clicks numbered 3,593 with training against 3,087 without. Dangerous actions were 2,730 against 2,155. Repeat offenders were 801 against 647. A Welch-corrected ANOVA gave F(1, 14592) = 18.37 with p less than 0.001, so this was not a marginal wobble in the data.
Something in the study did work. Crowdsourced reporting produced 14,401 reports from 4,260 participants, 11,035 of them on the simulated mails, plus a further 4,075 reports from 1,543 employees outside the study, at an operational load the authors judged acceptable. Detection through people reporting things beat correction through people being trained.
Nineteen Thousand Staff, Eight Months, a Two Percent Reduction
Grant Ho, Ariana Mirian, Stefan Savage and Geoffrey M. Voelker ran an eight-month randomised controlled experiment across more than 19,500 employees of UC San Diego Health, using ten distinct phishing campaigns, and presented it at the 46th IEEE Symposium on Security and Privacy in San Francisco in May 2025 and again at Black Hat that August.
Annual mandatory training showed no significant relationship with whether a user fell for a phishing email. Embedded training, the same teachable-moment mechanism the Swiss study tested, reduced the likelihood of clicking a phishing link by 2 percent. The engagement data explains the size of that effect: 75 percent of users spent a minute or less on the training material, and a third closed the page immediately without engaging at all.
The design of the lure mattered far more than the state of the learner. An Outlook password-update message drew a 1.82 percent click rate. A vacation-policy-update message drew 30.8 percent, a spread of roughly 17 times, from the same population inside the same study. The authors’ own summary is that anti-phishing training programmes, in their current and commonly deployed forms, are unlikely to offer significant practical value in reducing phishing risks.
One number from that study is widely misquoted, so it is worth stating carefully. In the first month only 10 percent of employees clicked a phishing link, and by the eighth month more than half had clicked at least one. That is cumulative individual exposure accumulating over ten campaigns, not a monthly click rate rising sevenfold. Two independent field studies, different countries, different sectors, five years apart, both landing between weakly positive and actively negative, is the finding worth carrying forward.
Baselining Behaviour Is a Regulated Activity in Europe
The vendor phrase for behavioral analytics is establishing a baseline of normal activity. In a European workplace that description already carries an eight-figure penalty. France’s data protection regulator, the CNIL, fined Amazon France Logistique 32 million euros for excessively intrusive employee monitoring, announced on 23 January 2024. The monitoring covered more than 6,000 permanent employees plus a significant number of temporary workers across French warehouses.
The CNIL also held the 31-day retention of this behavioral data excessive in light of the commercial interests pursued, and found breaches of GDPR Articles 6, 5(1)(a), 5(1)(c), 12, 13 and 32. Any organisation deploying user behaviour analytics in Europe is deploying a regulated processing activity, and granularity that reads as useful telemetry to a security team can read as unlawful surveillance to a regulator.
As summarised by the law firm Fieldfisher, the indicators the regulator struck down were behavioral metrics of exactly the kind a monitoring product generates automatically:
- The stow machine gun indicator, which raised an error whenever a worker scanned an item less than 1.25 seconds after the previous scan.
- Idle time, flagging any scanner downtime of ten minutes or more.
- Latency under ten minutes, flagging scanner interruptions of between one and ten minutes.
The False Positive Rate Nobody Has Published
Endpoint protection has an independent scoreboard. AV-Comparatives publishes false alarm counts next to protection rates, which is how anyone can see that near-identical detection can come with wildly different noise. Behavioral analytics has no equivalent. Research for this article located no independent, peer-reviewed measurement of production false positive rates for user behaviour analytics, and no named organisation publishing a deployment with before and after numbers. Every result found was published by a vendor. No freely available assessment of the market from a major analyst firm could be retrieved either, so no such position is paraphrased here.
What is documented is the environment those alerts land in. The Oxford security operations study presented at USENIX Security 2022 recorded a lead analyst’s estimate that 99 percent of generated alarms are false positives, and found that only 10 percent of its surveyed practitioners used machine learning based tooling of any kind. Arp and colleagues name the base rate fallacy explicitly among the ten pitfalls that inflate published anomaly detection results, and found that all 30 top-tier papers they reviewed were affected by at least three pitfalls. Robin Sommer and Vern Paxson made the underlying argument at the IEEE Symposium on Security and Privacy in 2010: anomaly detection transfers poorly from other machine learning domains into intrusion detection, because the cost of a mistake and the variability of normal are both unusually high.
The defensible position on behavioral analysis is narrow and still useful. Reporting works, as the Swiss study showed. Lure design predicts outcomes better than any training record does. Behavioral signals about accounts and machines may well earn their place, but until someone benchmarks them the way endpoint products are benchmarked, a vendor’s false positive claim is a marketing figure, and the training half of the field has two large field studies pointing the wrong way.
Sources: arXiv (Lain, Kostiainen and Capkun, IEEE S&P 2022) · UC San Diego Today · EurekAlert · Fieldfisher · USENIX Security Symposium (Alahmadi, Axon and Martinovic) · USENIX Security Symposium (Arp et al.)