
Synthetic speech has become realistic enough to fool both humans and machines. Most deepfake audio detectors are trained to spot known synthesis artifacts — but they routinely fail when faced with generation tools they've never seen. A new detection approach flips the problem: instead of hunting for manipulation traces, it verifies whether a voice actually belongs to the claimed speaker, using off-the-shelf speaker verification technology.
Most detection systems are built as two-class classifiers, trained on examples of both real and fake audio. This works well in lab conditions but breaks down in the real world.
Rather than learning what fake audio sounds like, this method learns what a specific person's real voice sounds like — and flags anything that doesn't match.
Both the audio under test and the reference recordings are converted into numerical "embeddings" — compact representations of vocal identity — using a speaker embedding extractor. Two strategies are then used to compare them:
If the similarity score falls below a threshold, the audio is flagged as fake.
Several established speaker verification architectures were adapted for this task:
The models were tested across three datasets: ASVSpoof2019, FakeAVCeleb, and the In-The-Wild Audio Deepfake dataset (IWA).
Real audio circulating online is rarely clean. The approach was also tested against everyday background noise — footsteps, traffic, wind, and more.
_____________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________
How does speaker verification help detect deepfake audio? It checks whether a voice recording actually matches a reference set of genuine recordings from the claimed speaker, rather than searching for signs of synthetic manipulation.
Why does this method generalize better than traditional deepfake detectors? It is trained only on real audio, so it doesn't depend on having seen a specific synthesis or voice-cloning tool in advance — any mismatch with the real voice is treated as suspicious.
What is the difference between Centroid-Based and Maximum-Similarity testing? Centroid-Based testing compares test audio to one averaged reference voice profile, while Maximum-Similarity compares it individually against every reference recording and keeps the best match.
Does this approach work with noisy or low-quality audio? Yes, though performance decreases as noise increases. The Maximum-Similarity strategy and noise-augmented training both improve robustness in real-world conditions.
Which model performed best overall? POI-Forensics delivered the strongest overall results, with the H/ASP speaker verification model close behind despite not being originally designed for deepfake detection.
Discover more about the paper: identifai.net/scientific-publications/identifai-publications-on-deepfake-audio-detection
Note: This article was prepared with the support of artificial intelligence
talk to a human expert
No sales pitch. Just a conversation.
We stand for truth