Implementing AI Solutions for Real-Time Cyber Threat Detection

A person wearing a Guy Fawkes mask sitting in a dark room in front of glowing monitors

AV-Comparatives released its Business Security Test on 15 December 2025, covering 17 enterprise products run on Windows 11 between August and November against 461 malicious URL cases. Every product stopped between 96.3 and 100 percent of them, which makes the protection column close to useless for telling the products apart. The false alarm column, running from 0 to 24, is not. Anyone deciding how much automated detection to buy, and how much analyst time to spend on what it produces, should start with that table rather than with vendor language about adaptive defence.

Seventeen Products, Almost Identical Protection, a 24-Fold Spread in False Alarms

AV-Comparatives is an independent test lab based in Innsbruck, Austria, and it publishes the false alarm counts alongside the protection rates. In the August to November 2025 business test, the two columns tell opposite stories. Protection varied by fewer than four percentage points across the whole field. False alarms varied by a factor of 24 between the cleanest and the noisiest product.

The same lab’s Malware Protection Test of September 2025, run over 1,015 samples, produced a different picture again: every product recorded zero false alarms against common business software, and protection ran from 98.2 percent for ManageEngine to 100 percent for Cisco, Elastic, ESET, G Data and Microsoft. A product’s noise level is a property of the test conditions as much as of the engine, which is exactly why a single vendor-supplied detection percentage tells a buyer very little.

Selected results from the business test, protection rate first and false alarm count second:

  • Kaspersky, 100 percent protection with 0 false alarms
  • ESET, 100 percent with 6 false alarms
  • Bitdefender, 99.8 percent with 2 false alarms
  • CrowdStrike, 99.3 percent with 20 false alarms
  • Rapid7, 98.5 percent with 0 false alarms
  • ManageEngine, 97.4 percent with 24 false alarms
  • Cisco, 96.3 percent with 3 false alarms

Only Two of Twenty Surveyed Practitioners Used Machine Learning Tools at All

Bushra A. Alahmadi, Louise Axon and Ivan Martinovic of the University of Oxford presented a qualitative study of security operations centres at the 31st USENIX Security Symposium in August 2022, built on interviews with 21 practitioners across seven distinct SOCs, five of them managed service providers, plus a survey of 20 respondents. Its title quotes a lead analyst: “We know 99% of the alarms we generate are false positives”. One participant reported finding one real threat per 100 alerts investigated. Another put it at 50 legitimate alerts in every 200.

The finding that matters most for any article about AI detection is the adoption number. Two of the 20 surveyed practitioners, 10 percent, reported using machine learning based tools, against 90 percent using intrusion detection systems and 80 percent using SIEM. Fifty-five percent said they relied on tacit knowledge more than half the time. Whatever the market is buying, the analysts interviewed for that study were mostly not working with it.

The largest intrusion of the period was found the old way. Kevin Mandia told Dark Reading in January 2021 that FireEye traced the SolarWinds compromise from a severity-zero alert, a second device registered against an employee’s two-factor authentication, and his description of it was that somebody was accessing the network the way staff do, but with a second registered device. Analysts phoned the employee, who denied registering it. Backdoored Orion updates had gone to roughly 18,000 customers, of which FireEye said a very small number were actually targeted and compromised.

A Five Megabyte Append That Flipped Malware to Benign

On 18 July 2019 the Australian firm Skylight Cyber published a universal bypass of Cylance’s machine learning antivirus, later a BlackBerry product. The technique was to append roughly five megabytes of strings taken from a whitelisted video game executable to a malicious file. Skylight reported that all ten of the top malware families for May 2019 moved from scores between -999 and -826, firmly malicious, to scores between +625 and +999, firmly benign. Across a 384-sample set, 83.59 percent bypassed detection using the strings of one application and 88.54 percent using several.

This was not a laboratory curiosity that a vendor could dismiss. The CERT Coordination Center issued Vulnerability Note VU#489481 describing Cylance antivirus products as susceptible to a concatenation bypass. A model that scores a file on features an attacker can freely add is a model an attacker can steer.

Channel File 291 and the Cost of Shipping Content Automatically

The other documented failure mode is not evasion but the delivery pipeline itself. CrowdStrike’s own root cause analysis of the 19 July 2024 Falcon outage, reported by The Hacker News, attributed it to an out-of-bounds read caused by a mismatch between the 21 inputs passed to the Content Validator through the IPC Template Type and the 20 supplied to the Content Interpreter. The United States Cybersecurity and Infrastructure Security Agency issued a public alert about the resulting outage the same day.

Delta Air Lines told reporters the incident cost it an estimated $500 million in lost revenue and additional costs across thousands of cancelled flights. Rapid-response detection content is the mechanism that makes near-real-time protection possible, and it is also the mechanism that can push a defect to every endpoint at once. Any deployment plan for automated detection needs a staged rollout answer, not just a detection rate.

What a National Programme Publishes About Its Own Automation

The clearest public numbers on automated detection at scale come from a government body rather than a vendor. The UK National Cyber Security Centre reports that in the year from 1 September 2024 to 31 August 2025 its Active Cyber Defence programme removed 1.2 million cyber-enabled commodity campaigns, resolved 79 percent of confirmed phishing attacks within 24 hours and took half of them down in under an hour. Its Suspicious Email Reporting Service took 10.9 million reports in that year, more than 45 million since April 2020, and has removed 412,000 malicious URLs since 2020.

The rest of the programme is measured the same way. Early Warning had 13,178 registered organisations and alerted on 316,343 IP addresses in the year, reporting 131,000 addresses with suspected malware or compromise to 1,350 organisations and 187,000 vulnerable addresses to 4,030 organisations. Mail Check covered 13,193 organisations and 402,796 domains, raising 1,014,887 alerts; Web Check covered 4,624 organisations and 133,913 domains and URLs, raising 569,467.

Read against the Oxford study, those figures describe what automation is genuinely good at: volume, breadth and takedown speed on commodity attacks. They do not describe judgement. The honest procurement position is that automated detection buys coverage, that its false alarm cost is measurable and varies enormously between products, and that both of the failures above were found and documented by people.

Sources: AV-Comparatives · USENIX Security Symposium · Skylight Cyber · The Hacker News · UK National Cyber Security Centre · Dark Reading

More articles to read