Imperceptible Voiceprint Protection via Human-Machine Perception Discrepancy Feature Disentanglement

Stage 1: Disentanglement-Reconstruction Training

Goal: Decompose speech into independent content and speaker representations through an information bottleneck. The AutoVC-based network learns to separate linguistic content from speaker identity.
Speaker Original Speech Reconstructed Speech Voice Conversion
(Different Speaker)
Male 1
Female 1
Male 2
Female 2

Stage 2: Defensive Perturbation Generation

Goal: Generate adversarial perturbations in the speaker embedding space to protect against voice cloning attacks. Protected speech sounds identical to humans but prevents machines from extracting speaker identity.
Speaker Original Speech Protected Speech Cloned from Original Cloned from Protected
(Attack Failed)
Male 1
Female 1
Male 2
Female 2

Comparison with Baseline Methods

Comparison of audio quality after protection. Our method maintains higher perceptual quality (MOS: 4.18) while achieving effective defense (DSR: 87.2%).
Method Original VoiceGuard RoVo Ours
Example 1
Example 2
Example 3