Imperceptible Voiceprint Protection via Human-Machine Perception Discrepancy Feature Disentanglement
Stage 1: Disentanglement-Reconstruction Training
Goal: Decompose speech into independent content and speaker representations through an information bottleneck.
The AutoVC-based network learns to separate linguistic content from speaker identity.
Speaker
Original Speech
Reconstructed Speech
Voice Conversion (Different Speaker)
Male 1
Female 1
Male 2
Female 2
Stage 2: Defensive Perturbation Generation
Goal: Generate adversarial perturbations in the speaker embedding space to protect against voice cloning attacks.
Protected speech sounds identical to humans but prevents machines from extracting speaker identity.
Speaker
Original Speech
Protected Speech
Cloned from Original
Cloned from Protected (Attack Failed)
Male 1
Female 1
Male 2
Female 2
Comparison with Baseline Methods
Comparison of audio quality after protection. Our method maintains higher perceptual quality (MOS: 4.18)
while achieving effective defense (DSR: 87.2%).