there have been a number of posts about anti-virus/anti-malware testing recently... even i posted about testing in response to what has become a series of posts about anti-virus testing over on anton chuvakin's blog (1, 2, 3, 4)... well this post is a follow-up because anton has managed to post the original test paper that his series of posts were based on...
to say that i was unimpressed would be an understatement... lets start with the number of samples - you may recall from my previous post that i said that the minimum number samples needed to account for the 2% detection rate that was being claimed was 50... according to the actual paper
Of the 35 malware files, three invalid files were removed from the sample set, leaving 32 malware binaries used in the final tests and performance calculationsso if only 32 samples were used, how is it that the lowest scoring product only detected 2% of the samples? detecting just a single sample gives a detection rate of 3%, not 2%, and all products tested detected more than just one sample... it can't be blamed on anton misremembering the figure he was told either, since the actual test paper states
the lowest was tied between ClamAV and FileAdvisor with a 2% detection ratein one place and
two products tied for the lowest detection rate at 2%thankfully the chart with their results clears this up - it's 2 raw detections (not a 2% detection rate) which means a 6% detection rate (which was also correctly reported in that same chart)... now, you'll have to forgive me for calling a spade a spade, but this level of mathematical incompetence (recognizing that you can't have a 2% detection rate with only 32 samples is grade school math and simply reading the column marked "percent" in a chart takes even less skill than that) is inexcusable for people who wish to have their test taken seriously... given such a complete lack of mathematical acumen, it's almost understandable that they failed to realize that a test bed of 32 samples isn't anywhere near large enough to give statistically significant results...
bias also figured heavily in the test... not only because they only used samples that got through the layers of protection already present on the systems they culled the samples from (thereby missing a potentially huge chunk of what's really posing a threat to users and irrevocably compromising the results and the integrity of the test itself), but on a deeper level the test was written from the perspective of an incident response technician... what is the perspective of an incident response technician? well these are the people who spend their days dealing with the after effects of the failure of security software and/or preparing for the next failure so as to make cleaning up after that event easier than cleaning up after the previous one... all they see are the security product's failures because that's their job, and if it weren't for the immutable fact that all security products fail they wouldn't have that job and they'd have to find some other form of employment (such as performing and publishing dubious tests)... just as police (who deal with criminals on a daily basis) are prone to developing an imbalanced view of society if they're not careful, so too are incident response technicians prone to developing an imbalanced view of the efficacy of security products if they aren't exposed to the security product's successes (which are generally invisible by design)... this acute form of perceptual bias taints the entire test at fundamental levels - including the design and methodology of the test (as evidenced by their belief that they need only collect samples that successfully compromised production systems that already had protective measures in place)...
so far these problems would seemingly be attributable to the testers simply being inexperienced and/or suffering from false authority syndrome... it's time to shine a light on a part of the test that can't easily be attributed to that... there was one product in the test that didn't do too badly - in fact it did better than all other products, it detected 50% more than the next best product, and it was one of the few products in the test that weren't used through virustotal... that product was asarium by proventsure (no link for reasons that are about to be made clear)... go and read the paper carefully (it's only 4 pages) and tell me if you can see something a little off about it... yes, that's right - the product that did the best, the product that was better than all others by a wide margin was the product made by the company whose president helped write the paper and is the contact listed in the abstract... the test, which is titled "Antiviral shortcomings with respect to 'real' malware" by gary golomb, jonathan gross, and rich walchuck, is NOT independent... one or more of it's authors has a clear vested interest in making one product look good at the expense of all others... this puts all the other problems with this test into a new and decidedly unfavourable light... the bad math, the sample selection bias, the insignificant sample size, etc. - in light of this revelation they all point to a cooked test designed to make all products other than asarium look worse than they really are (FUD) in order to make asarium look better than it is in comparison to them (snake oil)... the test, therefore, becomes little more than a marketing stunt by a disreputable company whose product should probably be given a wide berth...
and poor anton chuvakin - though widely regarded as a security expert, not only is he clearly not an authority on malware himself but apparently he also can't recognize a fraud/pretender when he sees one... that doesn't bode well for average folks' ability to do the same, does it...