Face verification and biometric comparison technology

White Paper · Standards Explained

NIST's Face Recognition Benchmark, Explained Simply.

Every face-matching vendor claims to be accurate. Only one test is run by a government lab with nothing to sell, on every algorithm that shows up — and the leaderboard changes every few weeks. Here is what it actually measures, and what to ask instead of trusting a single number.

One test, run by nobody with something to sell

The US National Institute of Standards and Technology (NIST) has been testing face recognition algorithms continuously since 2017. Any company can submit an algorithm; NIST runs it against the same fixed set of photographs everyone else was tested on, and publishes the result. There is no fee to a vendor for a good score and no way to buy a better one — which is precisely why procurement teams around the world treat it as the reference point, instead of a vendor's own marketing claims.

The report itself is a genuinely live document, not an annual release. The edition current at the time of writing is the thirty-seventh since testing began, updated roughly every few weeks as new algorithms are submitted. That matters more than it sounds — see the third section below.

Two ways a face match can go wrong

Strip away the terminology and a face-matching system can fail in exactly two directions. It can fail to recognise someone it should — turning away a genuine customer, or blocking a real citizen from a service they are entitled to. Or it can wrongly accept someone it should not — letting an impostor through as if they were the account holder.

NIST calls the first the false non-match rate and the second the false match rate. The reason this is worth understanding, rather than skipping past, is that no system sets both to zero. Every system sits somewhere on a dial between the two: tune it to reject fewer genuine people, and it will let more impostors through; tune it to catch more impostors, and it will turn away more genuine people. A vendor quoting one impressive number without saying which direction it was tuned for is not giving you the full picture.

Why the same vendor can rank differently on different tests

NIST does not test on one photo set. It runs algorithms against several — including deliberately difficult, realistic ones, such as poor-quality webcam photographs taken at a border crossing, alongside clean, studio-style portraits. That is not an accident of methodology; it reflects how photo quality actually varies in the field. A system that performs beautifully on a well-lit enrolment photo can fail noticeably on a grainy selfie taken on a low-end phone in poor light.

This is why a single overall "accuracy" figure is close to meaningless on its own. The honest question is not "how accurate is this algorithm" but "how accurate is this algorithm on photos that look like the ones my programme will actually capture." NIST's own ranking method reflects this: it rewards algorithms that perform consistently across several datasets, rather than ones that excel at one and struggle at the rest.

Why we won't quote you a "number one"

The leaderboard is not static. Recent editions of the report have each added new or updated results from a dozen-plus vendors, sometimes several dozen, roughly every month. A ranking that was accurate in one edition can move by the next — not because anyone got worse, but because the field keeps submitting improved versions. Quoting a "current leader" in a document like this one would be confidently wrong within weeks.

What actually matters for a real programme is narrower and more useful than a global rank: ask a vendor for their published NIST result on the dataset closest to your real capture conditions, not their best result overall. Ask when that specific algorithm was submitted — a result from several years ago tells you little about what you would actually be buying today. And ask whether the figure comes from a single dated submission or from NIST's "by developer" view, which only shows each company's best-ever entry.

Why we treat this as a procurement input, not a marketing line

This is also, plainly, how we approach matching technology ourselves. We do not bundle one face-matching engine by default and describe it as the best available. The matching layer underneath our enrolment and verification products is independently benchmarked and selected against the assurance level a specific programme requires — and revisited as the evidence changes, because the evidence does change, continuously.

If you are evaluating vendors for a national identity, SIM registration or KYC programme, treat a NIST citation the way this document has tried to model: ask which dataset, ask which submission date, and ask what the trade-off between the two error rates was tuned for. A vendor willing to answer all three in writing is telling you something a leaderboard position alone cannot.