Deepfake Audio Detection via Speaker Verification | identifAI

Blogs

Can Speaker Verification Detect Deepfake Audio Better Than Traditional Methods?

Synthetic speech has become realistic enough to fool both humans and machines. Most deepfake audio detectors are trained to spot known synthesis artifacts — but they routinely fail when faced with generation tools they've never seen. A new detection approach flips the problem: instead of hunting for manipulation traces, it verifies whether a voice actually belongs to the claimed speaker, using off-the-shelf speaker verification technology.

Why Do Traditional Deepfake Audio Detectors Struggle?

Most detection systems are built as two-class classifiers, trained on examples of both real and fake audio. This works well in lab conditions but breaks down in the real world.

  • They perform near-perfectly on their training dataset, since they were built specifically for it
  • Performance collapses to coin-toss level on audio generated by tools not seen during training
  • New text-to-speech and voice-cloning tools appear constantly, so static training data quickly becomes outdated
  • Adding more known attacks to training only pushes the problem further — it never fully solves it

What Is the "Person-of-Interest" Approach?

Rather than learning what fake audio sounds like, this method learns what a specific person's real voice sounds like — and flags anything that doesn't match.

  • The detector is trained exclusively on real audio, so it never needs examples of fake speech
  • Because no specific manipulation type is learned, generalization to new attacks is built in by design
  • The system checks whether the voice in a test recording matches a reference set of authentic recordings from the claimed speaker
  • This reframes deepfake detection as a speaker verification problem

How Does the Detection Process Actually Work?

Both the audio under test and the reference recordings are converted into numerical "embeddings" — compact representations of vocal identity — using a speaker embedding extractor. Two strategies are then used to compare them:

  • Centroid-Based (CB) testing — averages all reference embeddings into a single centroid, then compares the test audio against that one average voice profile
  • Maximum-Similarity (MS) testing — compares the test audio against each reference recording individually and keeps the highest similarity score

If the similarity score falls below a threshold, the audio is flagged as fake.

Which Speaker Verification Models Were Tested?

Several established speaker verification architectures were adapted for this task:

  • ClovaAI — a lightweight ResNet-34 variant optimized for open-set speaker recognition
  • H/ASP — an improved version using attentive statistics pooling and noise/reverberation augmentation
  • ECAPA-TDNN — a Time Delay Neural Network with channel attention for stronger feature extraction
  • POI-Forensics — originally designed for audio-visual deepfake detection, here used with audio only

How Well Does This Approach Generalize to Unseen Attacks?

The models were tested across three datasets: ASVSpoof2019, FakeAVCeleb, and the In-The-Wild Audio Deepfake dataset (IWA).

  • Traditional supervised detectors scored near-perfectly on their own training dataset but dropped to near-random performance on the other two
  • The speaker-verification-based approach delivered consistently strong results across all three datasets
  • POI-Forensics achieved the best overall accuracy, closely followed by H/ASP — remarkable given it wasn't originally built for this task
  • The Maximum-Similarity strategy scaled better with larger reference sets, especially on more challenging, "in the wild" audio

Is This Method Robust to Real-World Noise?

Real audio circulating online is rarely clean. The approach was also tested against everyday background noise — footsteps, traffic, wind, and more.

  • Performance degrades as noise increases, as expected
  • The Maximum-Similarity strategy holds up better than Centroid-Based testing under noisy conditions
  • Training with noise augmentation significantly reduces performance loss, making the system more deployment-ready

Who Can Benefit From This Approach?

  • Forensic analysts investigating suspected audio manipulation
  • Newsrooms and fact-checkers verifying statements attributed to public figures
  • Platforms and moderators screening impersonation attempts of known individuals
  • Security teams protecting voice-based authentication systems

_____________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________

Frequently Asked Questions

How does speaker verification help detect deepfake audio? It checks whether a voice recording actually matches a reference set of genuine recordings from the claimed speaker, rather than searching for signs of synthetic manipulation.

Why does this method generalize better than traditional deepfake detectors? It is trained only on real audio, so it doesn't depend on having seen a specific synthesis or voice-cloning tool in advance — any mismatch with the real voice is treated as suspicious.

What is the difference between Centroid-Based and Maximum-Similarity testing? Centroid-Based testing compares test audio to one averaged reference voice profile, while Maximum-Similarity compares it individually against every reference recording and keeps the best match.

Does this approach work with noisy or low-quality audio? Yes, though performance decreases as noise increases. The Maximum-Similarity strategy and noise-augmented training both improve robustness in real-world conditions.

Which model performed best overall? POI-Forensics delivered the strongest overall results, with the H/ASP speaker verification model close behind despite not being originally designed for deepfake detection.

Discover more about the  paper:  identifai.net/scientific-publications/identifai-publications-on-deepfake-audio-detection

Note: This article was prepared with the support of artificial intelligence 

Recent Blogs
See all blog articles

talk to a human expert

Tell us about your business. We'll come back to you within one business day.

Thank you!
Your submission has been successfully sent to our team
Oops! Something went wrong while submitting the form.

No sales pitch. Just a conversation.

We stand for truth