The Study
Investigation of Deepfake Voice Detection Using Speech Pause Patterns: Algorithm Development and Validation
This study found that real people’s voices have natural pauses when they breathe or think, but fake voices don’t pause the same way. It used a computer to spot these differences and guess which voices were real or fake — and it got it right about 80% of the time in its test. But it doesn’t prove fake voices are always like this — just that in this test, they were different.
Analysis score
Maximum 0 for a computational/algorithm study.
Where the score came from
Real people pause when they breathe or think, but AI voices don’t—so scientists taught a computer to spot fake voices by how long they pause between words.
Where does this study sit?
Reviews of RCTs (Meta-analyses)
Max 100Randomized Trials
Max 90Reviews of Cohort Studies
Max 85Cohort Studies
Max 72Reviews of Case-Control Studies
Max 63Case-Control Studies
Max 58Cross-Sectional & Case Series
Max 50Expert Opinion
Max 50 / 100
Quality score
Based on clinical experience or non-systematic literature reviews. The lowest level of evidence as they are most susceptible to bias and personal perspective.
Key takeaways
Summary
Based on the study abstract and findings.
- 1Yes—this means even if fake voices get better, this method might still work because it looks at natural human biology, not just digital tricks.
- 2The computer got 81% of the voices right in tests, and still got 79% right when it saw new voices it had never seen before.
Score breakdown, methodology, conflicts of interest, evidence analysis & raw study data
Publication
Journal
JMIR Biomedical Engineering
Year
2024
Authors
Nikhil Valsan Kulangareth, Jaycee M. Kaufman, Jessica Oreskovic, Yan Fossat
Related Content
Claims (6)
AI-generated deepfakes can create realistic video and audio simulations of real people to spread false information.
Machine learning models that analyze pauses in speech can correctly identify deepfakes 79% of the time when tested on new speakers, new text, and new cloning tools.
Cloned voices have less variation in how long each speech segment lasts and spend more time speaking than real human voices, because they lack natural pauses like breathing and thinking.
Analyzing pauses in speech can identify synthetic audio generated by deepfake tools, even when those tools were not used during training, and this method is more reliable than techniques that look for digital fingerprints.
Cloned voices have more uniform speech patterns with longer pauses between phrases, less variation in how long each word or segment lasts, more time spent speaking, and fewer brief or long pauses than human speech, allowing machine learning systems to identify them as synthetic with up to 81% accuracy in controlled tests.
A machine learning model using pauses in speech can correctly identify cloned voices 81% of the time in a controlled test with 49 people and three voice cloning tools, and it outperforms other common machine learning methods.
Not medical advice. For informational purposes only. Always consult a qualified healthcare professional before making health decisions.