Critical Evaluation: Real ElevenLabs vs. YouTube Human SpeechAnti-Sycophancy Protocol (ArXiv 2602.23971)
Following the Ask Don't Tell protocol, we downloaded 60 real audio clips from garystafford/deepfake-audio-detection (Hugging Face Hub) to test whether synthetic 100% AUROC holds up in the wild.
Real Benchmark AUROC
0.2844
Heuristic breakdown
Real-World Accuracy
60.00%
36/60 Correct
Real Recall (Detection)
30.00%
Missed 21/30 Clones
False Accusation Rate
10.00%
3 Humans Falsely Accused
The Real-World Forensic Finding: Real ElevenLabs voice clones retain room ambient noise from reference audio (noise floor ~ -45 dBFS) and use modern flow-matching rather than transposed-conv combs. This proves that simple noise-floor and comb heuristics cannot be relied upon in isolation without full multi-resolution RVQ neural embeddings.