What Is POI-Forensics and How Does It Detect Deepfakes Using Voice and Face Identity?
Deepfakes no longer come in just one flavor. A single video can have a swapped face, a cloned voice, manipulated facial expressions, or any combination of the three — and new manipulation methods appear constantly. POI-Forensics, developed by researchers at the University of Naples Federico II and the Technical University of Munich, tackles this problem with a simple but powerful idea: instead of asking "is this fake?", it asks "does this voice and face really belong to the person it claims to be?"
Why Do Traditional Deepfake Detectors Struggle With So Many Manipulation Types?
Most detectors are trained on large datasets of real and fake videos, learning to spot artifacts left behind by specific manipulation methods.
- They perform well only on manipulation types included in training
- New generation methods appear faster than detectors can be retrained on them
- Fine-tuning on new attacks risks degrading performance on previously learned ones
- Performance often drops sharply on low-quality or compressed video, which is the norm on social media
- Many detectors are also vulnerable to adversarial attacks designed to fool them
What Is the Person-of-Interest (POI) Approach?
Rather than learning what manipulated content looks like, a POI detector learns what a specific person's real face and voice look and sound like — and flags anything inconsistent.
- The model is trained exclusively on real talking-face videos, never on fakes
- Since it never depends on a specific manipulation method, it achieves strong generalization to unseen attacks by design
- At test time, it compares the video under analysis to a reference set of authentic videos of the claimed identity
What Makes POI-Forensics Different From Earlier POI Methods?
Previous person-of-interest detectors relied mostly on facial and head-movement patterns, largely ignoring the voice.
- POI-Forensics is fully multimodal, combining both audio and video identity cues
- It uses contrastive learning to keep embeddings of the same person close together and embeddings of different people far apart
- Adding the audio signal alone produces a large accuracy gain — 15 to 20 percentage points on average — over video-only approaches
- It can detect single-modality attacks (video-only or audio-only) as well as combined audio-video manipulations
How Does the Detection System Actually Work?
- Short video and audio segments are converted into embeddings using dedicated neural networks
- These embeddings are compared against a reference set of pristine videos (around 10 clips, roughly 30 seconds each) of the claimed identity
- Separate similarity scores are computed for video, audio, and combined audio-video signals
- The scores are normalized and fused into a single decision statistic
- A statistical threshold, set to achieve a target false-alarm rate, determines whether the video is flagged as manipulated
What Types of Manipulation Can It Detect?
- Face swapping — replacing one person's face with another's
- Facial reenactment — puppeteering someone's real face with another person's expressions or speech
- Voice cloning / audio-only manipulation — replacing the audio track while keeping the original video
- Combined audio-video attacks — manipulating both modalities together, where performance is highest since more inconsistencies are exposed
How Well Does POI-Forensics Perform Compared to Other Methods?
- On high-quality video, it outperforms the best reference methods by roughly 2.3% AUC on average
- On low-quality, compressed video — the most realistic real-world scenario — the advantage grows to an 11% AUC gain over the next-best method
- Under adversarial attacks, most competing detectors collapse, while POI-Forensics retains meaningful performance
- Detection improves further with more reference footage, though variety of clips matters more than raw video length
Who Can Use POI-Forensics?
- Newsrooms and fact-checkers verifying footage of public figures
- Platforms and content moderators screening impersonation attempts
- Security and fraud-prevention teams protecting executives and high-profile individuals from voice- or face-based scams
- Forensic analysts conducting in-depth authenticity investigations
Frequently Asked Questions
What is POI-Forensics?
POI-Forensics is an audio-visual deepfake detection method that verifies whether the face and voice in a video genuinely belong to the person they claim to represent, rather than searching for manipulation artifacts.
Why does POI-Forensics generalize better than traditional deepfake detectors?
It is trained only on real audio-visual data, so it has no dependency on any specific manipulation method — any mismatch with the real person's identity is treated as suspicious, regardless of how the fake was made.
How much reference footage does it need for a given person?
Around 10 short pristine video clips, roughly 30 seconds each, are enough to build a reliable reference set.
Can it detect audio-only or video-only manipulations?
Yes. Because it evaluates video and audio identity separately as well as jointly, it can flag single-modality attacks like voice cloning alone or face swapping alone.
Does it work on low-quality or compressed video?
Yes, and this is where it shows its biggest advantage over other methods, since it doesn't rely on subtle pixel-level artifacts that degrade under compression.
Note: This article was prepared with the support of artificial intelligence
Discover more about the paper: click here