Three papers presented at USENIX Security between 2019 and 2022 are the closest thing this market has to an independent audit. One compared 47 threat intelligence sources against a network telescope, one compared two leading paid vendors against each other, and one re-ran published machine learning models with their methodological errors corrected. None of the three flatters the product category, and together they set a floor under what a buyer can reasonably claim. Outside those papers, the efficacy of commercial threat intelligence platforms is essentially unmeasured in public.
Seventy-Three Percent of Scan Indicators Appear in Exactly One Feed
Vector Guo Li of UC San Diego and colleagues at Northeastern, Georgia Tech, NYU and the University of Illinois published a comparative analysis of threat intelligence at the 28th USENIX Security Symposium in August 2019. They examined 47 distinct IP address sources across six threat categories plus eight malware file hash feeds, and compared them against their own internet telescope. These are the measurements that bear directly on a purchase:
- 73 percent of all scan feed indicators were unique to a single feed, and 88 percent of brute force indicators were. In four of the six categories, more than three-quarters of pairwise intersection rates fell below 1 percent.
- Volume within one nominal category differed by orders of magnitude: 361,004 unique addresses in DShield’s scan feed against 1,572 in PA Analyst.
- The union of every scan address across all feeds covered less than 2 percent of the scanners the authors’ telescope observed, rising to roughly 10 percent when restricted to scanners of more than 10,000 hosts.
- Thirty-three feeds contained at least one unroutable address, and for 13 of them more than 1 percent were unroutable. Paid IP Reputation’s spam category reached a 78.7 percent unroutable rate.
- Twenty-one feeds included addresses belonging to top-Alexa domains, and several included content delivery network addresses.
- Median latency for scan feeds was one to three days against the telescope.
Two Paid Vendors, Twenty-Two Shared Actors, Four Percent Agreement
Xander Bouwman and colleagues at TU Delft, the Hasso Plattner Institute and Leiden asked a narrower question at USENIX Security 2020: what does paid threat intelligence add? For two leading commercial vendors they found almost no overlap between the vendors, nor with four large open feeds. Narrowing to 22 threat actors both vendors claimed to track, average indicator overlap was 2.5 to 4.0 percent, and the indicators that did appear in both arrived in the second vendor’s feed about a month later.
The interviews the authors ran with 14 paid subscribers explain why that gap has not hurt sales. Coverage and volume were not the customers’ main concern. The authors concluded that buyers optimise for the workflow of their scarce resource, analyst time, rather than for threat detection, and that they evaluate intelligence through informal processes and heuristics rather than the quantitative metrics research has proposed. A product bought on workflow grounds will not be dislodged by a coverage measurement.
Every One of Thirty Reviewed Papers Hit at Least Three Pitfalls
Daniel Arp of TU Berlin and co-authors from TU Braunschweig, King’s College London, Royal Holloway, KIT and UCL catalogued ten recurring pitfalls in security machine learning at USENIX Security 2022, later extended in Communications of the ACM in 2024. The ten are sampling bias, label inaccuracy, data snooping, spurious correlations, biased parameter selection, inappropriate baseline, inappropriate performance measures, the base rate fallacy, lab-only evaluation and an inappropriate threat model. Reviewing 30 top-tier security papers from the previous decade, they found every paper affected by at least three.
They then quantified the inflation by re-running the work. Correcting sampling bias cut DREBIN’s recall by 12.2 percent, from 0.964 to 0.846, and OPSEQS’ by 16.9 percent, from 0.883 to 0.734, with F1 falling 8.7 and 15.2 percent respectively. Source code authorship attribution accuracy dropped 48 percent once unused code was removed from the test set. In one comparison the deep learning system VulDeePecker scored an AUC of 0.984 and a true positive rate of 0.818 while a plain linear support vector machine baseline scored 0.986 and 0.963. The simpler model won.
That last result is the practical warning. A platform advertising a neural model has not thereby beaten anything, because in the published literature the trivial baseline was frequently not run.
Nobody Has Independently Evaluated a Commercial Platform
The two USENIX feed studies measured data, not products. Research for this article located no independent, peer-reviewed evaluation of any named commercial threat intelligence platform, whether Anomali ThreatStream, ThreatConnect, Recorded Future or OpenCTI. Everything comparative that turned up was vendor marketing, review-aggregator content or peer-review-site ratings. That absence is worth stating plainly rather than papering over with a survey figure.
Adversarial evasion, by contrast, is documented rather than hypothetical. Skylight Cyber’s 2019 work on Cylance’s machine learning antivirus flipped 83.59 to 88.54 percent of a 384-sample set to benign by appending strings from a whitelisted application. A deployed commercial security model was steered by an attacker-controlled input, and it was independent researchers who showed it.
What MISP Publishes, and What It Declines To
Not every part of this market is opaque. MISP, the Malware Information Sharing Platform, is free open-source software developed by CIRCL, Luxembourg’s national computer incident response centre. Access is free of charge, PGP-authenticated and governed by the Traffic Light Protocol. CIRCL states that it operates a sharing community of more than 1,100 international member organisations, and the project lists communities including NATO, FIRST, a law enforcement community, an industrial control systems community, the X-ISAC, a community for internet exchange points, national and governmental CSIRT communities, a financial sector community and Danish, Slovak, Kazakhstani and Japanese communities.
The project also says it does not maintain an exhaustive list of communities, which means no public figure for total MISP installations or users exists. That is an honest limitation, and it sits oddly against the money in the sector: Mastercard announced its acquisition of Recorded Future for $2.65 billion in September 2024 and confirmed the deal closed on 20 December 2024, describing the purchase as adding artificial intelligence driven threat intelligence capabilities.
Questions That Force a Vendor Onto Measurable Ground
None of this makes threat intelligence worthless. Sharing communities exist, they are used, and CIRCL’s is documented in public. It does mean that a claim of platform efficacy currently rests on the vendor making it, and that a buyer who wants evidence will have to produce it on their own network. Six questions move the conversation onto ground where that is possible.
- Ask which underlying sources the platform ingests, then ask what the measured overlap is with the feeds already in place. Li and colleagues showed that adding feeds mostly adds disjoint data, not corroboration.
- Ask for indicator latency against an independent observation point rather than against the vendor’s own collection time.
- Ask what proportion of delivered indicators are unroutable, expired or belong to shared infrastructure such as content delivery networks.
- For any machine learning claim, ask how the training labels were produced and whether the evaluation split respected time ordering, the sampling and snooping pitfalls Arp and colleagues named.
- Ask for the trivial baseline. If a linear model on the same features was never run, the reported advantage is unquantified.
- Ask for the base rate the product assumes, then compute the alert volume that rate implies for your network before signing anything.
Sources: USENIX Security Symposium (Li et al., 2019) · USENIX Security Symposium (Bouwman et al., 2020) · USENIX Security Symposium (Arp et al., 2022) · MISP Project · Mastercard Investor Relations · Skylight Cyber