How to read an antivirus laboratory test

By Ashley Jackson|Published |Guide

What AV-TEST, AV-Comparatives and SE Labs actually measure, what their scores do not tell you, and how to avoid being misled by a single number.

No commercial links on this page

Background explainer, no partner links. How the site is funded: affiliate disclosure.

Why the number in an advertisement is almost meaningless

"99.7% detection rate" tells you nothing on its own. Detection of what, tested when, against how many samples, under what conditions, and how did every other product score in the same round? A figure without those five pieces of context is marketing, not evidence — and because the leading engines cluster very tightly at the top, the difference between a headline 99.7% and a headline 99.4% is frequently a handful of samples and no practical difference at all.

The good news is that the underlying work is public, free to read, and produced by organisations that publish their methodology.

Who actually runs the tests

  • AV-TEST (Magdeburg, Germany) — bi-monthly consumer rounds scored in three categories out of six points each.
  • AV-Comparatives (Innsbruck, Austria) — a series of distinct tests including a Real-World Protection Test and a Malware Protection Test, with detailed methodology documents and a published false-positive count.
  • SE Labs (United Kingdom) — emphasises full-chain attack replication rather than sample sets alone.

All three are independent of the vendors in their governance, and all three are partly funded by vendors who commission testing — which is normal in this field and is why methodology transparency, not funding purity, is the thing to judge them on.

Diagram of the three categories independent antivirus laboratories score - protection, performance and usability - with what each category measures and a caution that any single score is a snapshot of one test round.
Three categories, not one. The protection score gets quoted; the usability score — false positives — often matters more in daily use. Original diagram produced for novatova.online; not a vendor screenshot.

The three categories, and what each really measures

Protection

The share of malicious samples blocked. Note the distinction between a real-world test, where the product faces a live malicious URL exactly as a user would, and a file detection test against a static collection. The real-world test is the more demanding and the more representative; the two are often quoted interchangeably, and should not be.

Performance

Measured slow-down on standard tasks: copying files, installing and launching applications, browsing, downloading. Measured on the laboratory’s reference hardware. If you run a seven-year-old laptop with a mechanical disk, your experience will be worse than the score suggests; on a current machine with an SSD, the differences between products are mostly academic.

Usability / false positives

How often clean software and legitimate sites are wrongly blocked. This is the category most often ignored and it deserves more weight than it gets: a product that blocks your accounting software and your bank teaches you to click "allow" without reading, and a user trained to dismiss warnings is less safe than one who was never warned.

Seven questions to ask of any quoted score

  1. Who ran the test? If the source is the vendor’s own page and no laboratory is named, treat it as advertising.
  2. When? A round from two years ago describes a product that has since been rewritten.
  3. Which test? Real-world protection and static file detection are different measurements.
  4. How did everyone else do? If eleven of thirteen products scored above 99%, being one of them is not a distinction.
  5. What was the false-positive count? An aggressive product can buy a protection score with it.
  6. How many consecutive rounds? Consistency across four rounds means far more than a single excellent month.
  7. Is the award a certification or a ranking? Many badges are pass marks that most participants receive; they are not a first place.
Our own practice

This is why we do not print detection percentages on this site. We are not a laboratory, a figure we reproduced would be stale within weeks, and a number without its methodology invites exactly the misreading described above. We link to the laboratories instead, and we say so in our editorial policy.

What the tests cannot tell you

  • Whether the product suits your machine and habits.
  • What the renewal price is, and whether the subscription auto-renews at a higher rate.
  • What data the product sends to the vendor, and under which privacy notice.
  • Whether the support is any good when something goes wrong.
  • How intrusive the interface is — upsell prompts and notification nagging are not scored.

Sources


Related