Understanding Cepstral Peak Prominence (CPP): The Anchor Measure of Modern Voice Assessment
🎯 Key Takeaways
- The 2018 ASHA instrumental protocol chose a cepstral measure as the acoustic correlate of voice quality, because it works across the entire severity range and in connected speech (Patel et al., 2018)
- CPP works where perturbation measures fail: severely dysphonic voices and continuous speech, with no cycle-by-cycle pitch tracking needed
- Published Praat cutoffs: below 14.45 dB (sustained /a/) or 9.33 dB (read speech) suggests a voice disorder, with up to 94.5% accuracy (Murton et al., 2020)
- But cutoffs are software-specific: the same study's ADSV cutoffs are ~3 dB lower, and a third implementation puts its running-speech cutoff at 4.0 dB. A CPP value without a software name is not interpretable
- Meta-analytic backing: smoothed CPP was identified as the most robust acoustic correlate of overall dysphonia severity (Maryn et al., 2009)
- Telehealth caveat: live videoconferencing depresses CPPS by 1.4–2.3 dB toward the dysphonic side on every platform tested. Uploaded local recordings, yes; live platform audio against a cutoff, never
In recent years, cepstral peak prominence (CPP) has become the anchor acoustic measure for assessing voice quality. The 2018 ASHA expert panel on instrumental voice assessment selected a cepstral measure as the recommended acoustic correlate of voice quality—“a general measure of dysphonia”—in place of traditional perturbation measures like jitter and shimmer (Patel et al., 2018). The panel itself noted this was the most contentious choice in the entire protocol, and the recommendation remains current as of 2026. Yet many clinicians still reach first for the traditional measures—measures that fail precisely when we need them most: with severely dysphonic voices.
If you were trained on jitter and shimmer, you're not alone. Most speech-language pathology programs still emphasize them. But the field has shifted, and understanding CPP is now essential for evidence-based voice assessment.
Let me explain what CPP actually measures, where it came from, why the evidence favors it, and—most importantly—how to use it correctly in clinical practice: which cutoffs apply to which software, and what changes when the recording comes from a phone or a telehealth session.
The Problem with Traditional Measures
For decades, jitter (cycle-to-cycle pitch variability) and shimmer (cycle-to-cycle amplitude variability) were the go-to acoustic measures. They made intuitive sense: a smooth, periodic voice signal should have minimal perturbation, while a rough or hoarse voice would show increased variability.
But there's a fundamental problem: both measures require the software to identify individual glottal cycles.
The Catch-22 of Perturbation Measures
Jitter and shimmer work for mildly dysphonic voices, but fail for severely dysphonic voices—where you need objective measures most. When the voice becomes buried in noise (severe breathiness, roughness, or hoarseness), there isn't enough periodic information for the algorithm to identify individual vocal fold cycles, and the numbers it returns are artifacts.
Additional limitations of jitter and shimmer:
- Valid only for steady sustained vowels: they can't assess voice quality during continuous speech, which is how patients actually communicate (Patel et al., 2018)
- Inconsistent correlation with perception: meta-analytic correlations with roughness and breathiness vary widely across studies (Barsties v. Latoszek et al., 2018)
- Software-dependent values: Praat and MDVP disagree by design, especially on noisy signals
- Fragile under noise and loudness variation: ambient noise and vocal intensity both shift the values substantially
We cover these limits, the signal-typing rules, and the software problem in depth in our jitter and shimmer guide.
Enter Cepstral Peak Prominence (CPP)
CPP came out of signal processing research on breathiness. Hillenbrand, Cleveland, and Erickson introduced the measure in 1994; Hillenbrand and Houde extended it in 1996 with the smoothed variant (CPPS) and showed it worked on continuous speech, not just vowels. Clinical adoption followed (Heman-Ackah et al., 2002), and a 2009 meta-analysis of acoustic correlates of overall voice quality concluded that smoothed CPP was the most robust acoustic correlate of dysphonia severity across studies (Maryn et al., 2009). That meta-analysis is the evidence base the ASHA panel cited when it selected a cepstral measure in 2018.
What is CPP?
The cepstrum is, loosely, a spectrum of the log spectrum: it re-analyzes the spectrum to ask “how regularly spaced are the harmonics?” A voice with clear, evenly spaced harmonics produces a sharp peak in the cepstrum at the quefrency corresponding to the vocal fold vibration period. CPP measures how far that peak rises above a regression line through the rest of the cepstrum, in dB. No individual cycles need to be found—the measure looks at the signal as a whole, which is exactly why it keeps working when cycle identification breaks down.
Conceptual Framework
Think of the voice spectrum as a city skyline. In a healthy voice, harmonic peaks (the buildings) stand tall and distinct above the noise floor (the horizon). In dysphonia, increased aperiodic energy (fog) obscures these harmonics. CPP quantifies how prominently the harmonic structure rises above this noise floor.
Why CPP Works Better
| Feature | Jitter/Shimmer | CPP |
|---|---|---|
| Cycle-by-cycle tracking | ✗ Required (fails for severe dysphonia) | ✓ Not required |
| Valid for continuous speech | ✗ No (sustained vowels only) | ✓ Yes (Hillenbrand & Houde, 1996) |
| Correlates with perceived severity | ⚠ Varies widely across studies (Barsties v. Latoszek et al., 2018) | ✓ Most robust correlate in meta-analysis (Maryn et al., 2009); r² up to .74 for estimating severity (Murton et al., 2020) |
| Valid across severity range | ✗ Mild-to-moderate, Type 1 signals only | ✓ Entire range (Patel et al., 2018) |
| Software-independent values | ✗ No | ✗ Also no—see the cutoffs below |
Clinical Cutoffs (Software- and Task-Specific)
CPP vs. CPPS – and Why the Software Name Is Part of the Number
Praat computes CPPS (smoothed CPP), which averages across time and quefrency before measuring the peak (Hillenbrand & Houde, 1996; Maryn & Weenink, 2015). Other programs implement CPP or CPPS with different smoothing, windowing, and voicing-detection choices. The result: the same recording yields systematically different values in different software, and cutoffs do not transport. In Murton et al. (2020), the Praat cutoffs sit about 3 dB above the ADSV cutoffs on the same recordings.
| Implementation & source | Sustained /a/ | Connected speech |
|---|---|---|
| Praat / PhonaLab CPPS (Murton et al., 2020) | < 14.45 dB | < 9.33 dB (Rainbow Passage) |
| ADSV CPP (Murton et al., 2020) | < 11.46 dB | < 6.11 dB (Rainbow Passage) |
| CPPS, running speech (Heman-Ackah et al., 2014) | — | < 4.0 dB (sensitivity 92.4%, specificity 79%) |
Values below the cutoff suggest a voice disorder; higher values indicate clearer voice quality. The Murton cutoffs classified disordered vs. healthy voices with up to 94.5% accuracy (295 patients, 50 controls). The Heman-Ackah study (835 patients, 50 controls) used a third implementation on running speech—its 4.0 dB cutoff is on a different scale entirely, which is the point: a CPP value means nothing without the software name and the task attached.
Both programs work—they just don't interchange
Sauder, Bretl, and Eadie (2017) compared CPPS from Praat and ADSV head to head for predicting voice disorder status: Praat reached an AUC of 0.91 (82% accuracy) and ADSV 0.81 (75% accuracy), with the two programs' values highly correlated (r = .88) yet numerically different. Either can support clinical screening. Mixing one program's values with the other's cutoffs cannot.
Two practical corollaries. First, report the task: vowel and connected-speech CPP sit on different scales (compare the columns above), and connected-speech analysis is also sensitive to how much unvoiced material the passage contains and how the program handles it. Second, report the settings: Praat's CPPS depends on its analysis parameters (Maryn & Weenink, 2015), so a defensible report names software, version, task, and settings—then compares only against cutoffs derived the same way.
CPP in the Real World: Phones, Telehealth, and Recording Quality
CPP is more robust to recording conditions than perturbation measures, but robust is not immune. Three situations deserve explicit rules:
Uploaded smartphone recordings: yes, with a headset
In the Bridge2AI-Voice validation, CPP from smartphones and tablets—with and without a low-cost headset microphone—correlated at r ≈ .99 with a research-grade chain (Awan et al., 2024). Offsets between devices remain, so pipeline-matched references still apply, but ranking and tracking are preserved. Protocol details in our smartphone recording guide.
Live telehealth audio: never against a cutoff
CPPS decreased on all six videoconferencing platforms tested, with large effect sizes—by 1.4–2.3 dB in connected speech, always downward, toward the dysphonic side of any published threshold (Weerathunge et al., 2021). A 2025 home-recording replication found the same: CPP from live Zoom capture failed validity criteria (Hwang et al., 2025). Live capture systematically over-calls dysphonia. Use the asynchronous workflow in our telehealth guide.
File requirements: lossless, adequate sample rate
Cepstral analysis needs the spectral bandwidth the algorithm expects. In PhonaLab, CPP requires an original sample rate of at least 10 kHz; files below that (common for messaging-app voice notes) return a flagged null rather than a plausible-looking number. Prefer WAV or FLAC at 44.1 or 48 kHz, recorded in a quiet room.
Common Questions About CPP
Q: Can I calculate CPP with free software?
Yes! Praat is free (though it has a learning curve). Our PhonaLab Voice Analyzer also provides free browser-based CPPS calculation with cited normative comparisons and PDF reports—no installation required, and audio is never stored.
Q: Why are my Praat/PhonaLab values higher than values I see in papers?
This is expected. Praat and PhonaLab calculate CPPS, and on the same recordings the Praat cutoffs sit about 3 dB above ADSV's (Murton et al., 2020). Papers using different software report different ranges; a running-speech study with a third implementation works on a scale where 4.0 dB is the cutoff (Heman-Ackah et al., 2014). Always check which software and task a paper used, and apply software-matched values only.
Q: Is CPP valid for telehealth recordings?
Split the question in two. Locally recorded files uploaded for analysis: yes—smartphone CPP correlates at r ≈ .99 with research-grade recording when a headset and a quiet room are used (Awan et al., 2024). Live videoconferencing audio: no—every platform tested depressed CPPS by 1.4–2.3 dB toward the dysphonic side (Weerathunge et al., 2021), so a live-capture value read against a normative cutoff will over-call dysphonia. Record locally, upload, analyze the file.
Q: Should I report CPP to referring physicians?
Absolutely. Many ENTs are familiar with CPP from the literature. Include a brief interpretation with the software named: “CPPS (Praat) of 8.2 dB on sustained /a/ is below the published cutoff of 14.45 dB for distinguishing disordered from healthy voices (Murton et al., 2020).” The software name is part of the number.
Bottom Line: Why CPP Matters for Your Practice
- 1CPP is the acoustic voice-quality measure of the 2018 ASHA protocol—a choice grounded in meta-analytic evidence and still current in 2026
- 2It works where jitter and shimmer fail: moderate-to-severe dysphonia and continuous speech
- 3The software name is part of the number: Praat, ADSV, and other implementations sit on different scales, and cutoffs do not transport
- 4Track trends over time with a constant recording chain, task, and software, rather than fixating on single values
- 5Use alongside perceptual assessment and laryngoscopy for comprehensive evaluation—CPP quantifies severity, it does not diagnose
- 6Mind the recording chain: uploaded lossless files yes; live videoconferencing audio against a cutoff, never
🎤 Calculate CPP on Your Patient Recordings
Upload any voice recording and get instant CPPS analysis computed with Praat's algorithm, compared against published software-matched cutoffs with the citation shown, plus professional PDF reports. Files that cannot support cepstral analysis return a flagged null with the reason. Audio is processed in memory and never stored.
Try Free Voice Analyzer →Includes CPP/CPPS, F0, jitter, shimmer, HNR, and AVQI multiparametric assessment
⚠️ Clinical Documentation Tool
The information in this article is provided for educational purposes and clinical documentation support. Acoustic measures like CPP are intended to supplement—not replace—comprehensive voice evaluation including perceptual assessment, patient history, and laryngoscopic examination when indicated. All clinical decisions should be made by qualified healthcare professionals based on the complete clinical picture. PhonaLab tools are designed for professional use and do not provide medical diagnoses.
References & Further Reading
- Patel RR, Awan SN, Barkmeier-Kraemer J, Courey M, Deliyski D, Eadie T, Paul D, Švec JG, Hillman R. (2018). Recommended protocols for instrumental assessment of voice: American Speech-Language-Hearing Association expert panel to develop a protocol for instrumental assessment of vocal function. American Journal of Speech-Language Pathology, 27(3), 887–905. doi:10.1044/2018_AJSLP-17-0009
- Hillenbrand J, Cleveland RA, Erickson RL. (1994). Acoustic correlates of breathy vocal quality. Journal of Speech and Hearing Research, 37(4), 769–778.
- Hillenbrand J, Houde RA. (1996). Acoustic correlates of breathy vocal quality: Dysphonic voices and continuous speech. Journal of Speech and Hearing Research, 39(2), 311–321.
- Maryn Y, Roy N, De Bodt M, Van Cauwenberge P, Corthals P. (2009). Acoustic measurement of overall voice quality: A meta-analysis. The Journal of the Acoustical Society of America, 126(5), 2619–2634.
- Maryn Y, Corthals P, Van Cauwenberge P, Roy N, De Bodt M. (2010). Toward improved ecological validity in the acoustic measurement of overall voice quality: Combining continuous speech and sustained vowels. Journal of Voice, 24(5), 540–555.
- Maryn Y, Weenink D. (2015). Objective dysphonia measures in the program Praat: Smoothed cepstral peak prominence and Acoustic Voice Quality Index. Journal of Voice, 29(1), 35–43. doi:10.1016/j.jvoice.2014.06.015
- Murton O, Hillman R, Mehta D. (2020). Cepstral peak prominence values for clinical voice evaluation. American Journal of Speech-Language Pathology, 29(3), 1596–1607. doi:10.1044/2020_AJSLP-20-00001
- Heman-Ackah YD, Michael DD, Goding GS Jr. (2002). The relationship between cepstral peak prominence and selected parameters of dysphonia. Journal of Voice, 16(1), 20–27. doi:10.1016/S0892-1997(02)00067-X
- Heman-Ackah YD, Sataloff RT, Laureyns G, Lurie D, Michael DD, Heuer R, Rubin A, Eller R, Chandran S, Abaza M, Lyons K, Divi V, Lott J, Johnson J, Hillenbrand J. (2014). Quantifying the cepstral peak prominence, a measure of dysphonia. Journal of Voice, 28(6), 783–788. doi:10.1016/j.jvoice.2014.05.005
- Sauder C, Bretl M, Eadie T. (2017). Predicting voice disorder status from smoothed measures of cepstral peak prominence using Praat and Analysis of Dysphonia in Speech and Voice (ADSV). Journal of Voice, 31(5), 557–566. doi:10.1016/j.jvoice.2017.01.006
- Barsties v. Latoszek B, Maryn Y, Gerrits E, De Bodt M. (2018). A meta-analysis: Acoustic measurement of roughness and breathiness. Journal of Speech, Language, and Hearing Research, 61(2), 298–323.
- Weerathunge HR, Segina RK, Tracy L, Stepp CE. (2021). Accuracy of acoustic measures of voice via telepractice videoconferencing platforms. Journal of Speech, Language, and Hearing Research, 64(7), 2586–2599. doi:10.1044/2021_JSLHR-20-00625
- Awan SN, Bahr R, Watts S, Boyer M, Budinsky R, Bridge2AI Voice Consortium, Bensoussan Y. (2024). Validity of acoustic measures obtained using various recording methods including smartphones with and without headset microphones. Journal of Speech, Language, and Hearing Research, 67(6), 1712–1730. doi:10.1044/2024_JSLHR-23-00759
- Hwang K, van Brenk F, McAuliffe MJ, Choi J, Švec JG, Chang YHM, Keller B, Levy ES. (2025). Validity of acoustic speech measures obtained through videoconferencing with children with dysarthria. International Journal of Speech-Language Pathology. Advance online publication. doi:10.1080/17549507.2025.2563845
- ASHA Practice Portal: Voice Disorders – Assessment and Treatment