# How accurate are AI voice detectors in real conditions?

Source: https://forum.global100.org/q/how-accurate-are-ai-voice-detectors-in-real-conditions/
Site: Global 100 Forum, category AI voice and audio
Published: 2026-09-19. Updated: 2026-09-22. Replies: 1.

## Question by Tomas Reyes (staff), 2026-09-19

In controlled benchmarks the best systems miss only a few percent of fakes, but accuracy collapses on audio from unseen generators, phone codecs and real-world recordings, where published tests show error rates of 15% and higher. Vendors advertise 99% figures from their own testing. Treat a detector as a screening signal, not a verdict.

## How accuracy is measured

Voice detection papers report equal error rate (EER): the point where the false alarm rate on genuine speech equals the miss rate on fakes. Lower is better, and the number only holds for the conditions it was measured under.

## What the ASVspoof challenges show

In [ASVspoof 2021](https://arxiv.org/abs/2109.00537) the best system on the logical access task reached 1.32% EER. The new deepfake task, built to mimic audio posted to social media with unknown compression, was far harder: 15.64% EER for the best submission and 22.38% for the best baseline.

[ASVspoof 5 (2024)](https://arxiv.org/abs/2408.08739) used crowdsourced speech from far more speakers, attacks tuned against surrogate detectors, and adversarial attacks for the first time. In the closed condition (challenge data only) the best system reached 8.61% EER and the AASIST baseline 29.12%. In the open condition (pre-trained speech models allowed) the best reached 2.59%.

## What happens outside the lab

[Müller and colleagues (Interspeech 2022)](https://arxiv.org/abs/2203.16263) re-implemented published detectors and tested them on 37.9 hours of found recordings of celebrities and politicians, 17.2 hours of them deepfakes. Performance degraded by up to one thousand percent relative to the benchmark results. Their conclusion: the field had tailored its solutions too closely to ASVspoof, and deepfakes are much harder to detect outside the lab. The same group's [earlier human study](https://arxiv.org/abs/2107.09667) found that people and a state-of-the-art detector struggled with the same attack types.

## Vendor claims, labelled as claims

[Pindrop](https://www.pindrop.com/) states on its homepage that its detection is independently validated at 99% accuracy, and a footnote on its [audio deepfake page](https://www.pindrop.com/use-cases/audio-deepfake-detection) says its Pulse accuracy was computed on the ASVspoof 5 dataset of 32 engines. A benchmark result, not a field result. [Resemble AI](https://www.resemble.ai/detect/) advertises up to 99.5% accuracy for its multimodal detector, from its own testing. The [ElevenLabs AI Speech Classifier](https://elevenlabs.io/ai-speech-classifier) makes a narrower promise: it reports the probability that a clip was made with ElevenLabs, analyses the first minute only, and its page states that it does not reliably classify audio from the Eleven v3 model. None publish per-condition error rates in the ASVspoof style. The structure mirrors the text side, and the [text detector accuracy thread](/q/how-accurate-are-ai-text-detectors-really/) walks through the same distinction.

**Reported voice detection results, and the conditions behind them**

| Detector or study | Reported result | Condition | Caveat |
| --- | --- | --- | --- |
| ASVspoof 2021, logical access | 1.32% EER (best team) | Lab-generated synthetic speech with channel and codec variation | No matched training data, but lab conditions throughout |
| ASVspoof 2021, deepfake task | 15.64% EER (best team); 22.38% (best baseline) | Social-media style audio, unknown compression | Same era of detectors, ten times the error |
| ASVspoof 5, closed condition | 8.61% EER (best); 29.12% (AASIST baseline) | Crowdsourced speech, adversarial attacks | Training limited to challenge data |
| ASVspoof 5, open condition | 2.59% EER (best) | Same data, pre-trained speech models allowed | Best case for a well-resourced lab system |
| Müller et al. 2022, in-the-wild | Up to 1000% performance degradation | 37.9 hours of found recordings, 17.2 hours fake | Detectors tuned to the benchmark |
| Pindrop (vendor claim) | 99% accuracy | Computed on ASVspoof 5 data, 32 engines | Vendor's own computation, benchmark not field |
| Resemble AI (vendor claim) | Up to 99.5% accuracy | Vendor tests across 250+ models | No per-condition error rates published |
| ElevenLabs classifier | Probability the clip is ElevenLabs audio | First minute of upload | Not reliable on Eleven v3; silent on other tools |

## What that means in practice

- **Ask which condition the number came from.** A 1% EER on clean benchmark audio and a 15% EER on social media audio can describe the same detector.
- **Expect the worst on new generators.** Every benchmark shows the largest errors on attacks the detector has not seen.
- **Use detectors to add friction, not to acquit.** A low score on a phone call passed through a codec is weak evidence; provenance such as a watermark is stronger where it exists, as the [watermarking thread](/q/does-watermarking-ai-generated-text-actually-work/) explains for text.

## Reply 1 by Tomas Reyes (staff), 2026-09-22

When you read a vendor accuracy figure, ask three questions before you believe it: what was the false alarm rate on genuine speech at that setting, was the test audio passed through a phone or meeting codec, and how many of the generators in the test were released after the detector was trained. A 99% figure with no answer to the first question is close to meaningless, because a detector that flags everything also catches every fake.

---
Cite as: Global 100 Forum, "How accurate are AI voice detectors in real conditions?", https://forum.global100.org/q/how-accurate-are-ai-voice-detectors-in-real-conditions/, accessed 2026-10-11.
