Jitter and Shimmer: What They Really Tell You (And What They Don't)
🎯 Key Takeaways
- Jitter and shimmer aren't obsolete—but the 2018 ASHA expert panel chose CPP over them as the acoustic measure of voice quality, because perturbation is only valid for mild-to-moderate dysphonia and steady sustained vowels (Patel et al., 2018)
- Validity depends on the signal, not the patient. Perturbation measures are meaningful only for nearly periodic (Type 1) signals; as a practical guideline, values below about 5% have been found reliable (Titze, 1995)
- They fail when you need them most—in severe dysphonia the algorithm cannot identify cycles, and the numbers it returns are artifacts
- Values are software-specific. With 1% added noise on an otherwise perfectly periodic signal, Praat reports 0.02% jitter and MDVP reports 0.6%—from the same sound (Boersma, 2009). Never mix Praat values with MDVP norms
- Loudness is a hidden confound. Voice SPL was the single largest influence on jitter and shimmer in healthy speakers (Brockmann et al., 2011), and the effect holds in patients (Brockmann-Bauser et al., 2018)
- Never compute them from live telehealth audio or lossy files. Perturbation fails below about 30 dB SNR (Deliyski et al., 2005), and platform processing adds exactly that kind of noise
If you trained as an SLP anytime before 2018, you probably learned that jitter and shimmer were the standard for objective voice assessment. These measures appeared in every textbook, every voice lab report, every research paper. They felt scientific—precise numbers that could quantify what we heard perceptually.
Then the 2018 ASHA Expert Panel published its recommended protocol for instrumental voice assessment, and the acoustic measure of voice quality it selected was CPP (cepstral peak prominence), not jitter or shimmer. The panel was explicit about why: cepstral measures are viable across the entire range of dysphonia severity, in vowels and in connected speech, whereas jitter and shimmer “are only valid for mild-to-moderate dysphonia” and for relatively long, steady sustained vowels (Patel et al., 2018). The panel also noted that the choice of the acoustic voice-quality measure was the most contentious part of the entire protocol. Many clinicians were left confused: were all those years of jitter and shimmer measurements meaningless?
The answer is nuanced. Jitter and shimmer remain useful tools—but only when you understand their limitations. This article explains exactly when these traditional measures work, when they fail, and how to use them appropriately in modern clinical practice.
What Jitter and Shimmer Actually Measure
Let's start with the basics, because the underlying concept matters for understanding the limitations.
Jitter
Cycle-to-cycle variation in pitch period.
Measures how consistent the time between vocal fold vibrations is. A perfectly periodic voice would have zero jitter. Higher jitter = more irregular period duration from one cycle to the next.
Perceptual association: irregularity percepts such as roughness—but the correlations vary widely across studies and are not one-to-one
Shimmer
Cycle-to-cycle variation in amplitude.
Measures how consistent the amplitude is from one vocal fold vibration to the next. Higher shimmer = more irregular amplitude, associated with breathiness and reduced glottal efficiency.
Perceptual association: breathiness and overall dysphonia—again variable across studies, not a one-to-one mapping
Here's the critical insight: both measures require the software to identify individual vocal fold cycles. The algorithm must locate each glottal period to calculate how much consecutive periods vary. This works well for relatively regular voices—and fails for severely disordered ones, where there are no clear cycles to find. For how each acoustic measure relates to the CAPE-V and GRBAS dimensions, see our perceptual-acoustic reconciliation guide.
The Fundamental Problem: When Algorithms Fail
The 1995 National Center for Voice and Speech workshop, summarized by Ingo Titze, established a signal classification that still governs when perturbation measures are valid. Titze defined Types 1–3; a fourth type for predominantly stochastic signals was added later by Sprecher and colleagues (2010):
Clear harmonic structure. The algorithm can identify cycles reliably. Jitter and shimmer are VALID.
Strong subharmonics or modulations approaching the energy of f0. Titze's summary statement is direct: for Type 2 signals, perturbation measures “are unreliable and contain little pattern information”—visual displays such as spectrograms are the appropriate tool. Do not report jitter/shimmer.
No reliable periodicity. The algorithm cannot identify cycles; Titze recommends perceptual ratings of roughness for these voices. Jitter and shimmer are INVALID.
Stochastic noise dominates, as in severe breathiness. There is no periodic structure to analyze. Jitter and shimmer are MEANINGLESS.
Signal typing is done by looking at a narrowband spectrogram, not by looking at the patient's diagnosis. If you are not comfortable classifying signals visually, our spectrogram reading guide walks through Types 1–4 with examples.
The Clinical Catch-22
Here's the frustrating reality: jitter and shimmer become invalid precisely when you need objective measures most—in severe dysphonia. Your patient with the most obviously disordered voice? The algorithm can't reliably measure it. Your patient with a mild voice change that's hard to characterize? The algorithm works fine.
The 1995 NCVS summary statement offers two practical anchors. First: “perturbation measures less than about 5% have been found to be reliable”—above that, cycle identification is suspect and the number may be an artifact of algorithmic failure rather than a measure of the voice. Second: the analysis window should be on the order of 100 cycles for a stable perturbation estimate, which is why a 3–5 s sustained vowel is the standard task.
Why the 2018 ASHA Protocol Chose CPP
The 2018 ASHA Expert Panel (Patel et al.) selected a cepstral measure as the recommended acoustic correlate of voice quality. Three limitations of perturbation measures explain the choice:
- 1
Only valid for steady sustained vowels
The panel's wording: valid “for relatively long-duration sustained vowel contexts in which the client is attempting a relatively steady pitch and loudness production.” Connected speech—how patients actually communicate—cannot be analyzed with perturbation measures. A voice may sound disordered in conversation and look “normal” on a sustained /a/.
- 2
Only valid for mild-to-moderate dysphonia
Perturbation requires reliable cycle identification, which fails as severity increases. The more disordered the voice, the less reliable the measurement—exactly backwards from what we need clinically.
- 3
Inconsistent correlation with perception
Meta-analytic work on acoustic correlates of roughness and breathiness shows widely varying correlations for jitter and shimmer across studies (Barsties v. Latoszek et al., 2018). The numbers do not reliably track what clinicians hear.
Why CPP Works Better
CPP doesn't require cycle identification—it measures the prominence of the cepstral peak, a global index of how dominant the harmonic structure is relative to noise. That is why it works across the severity spectrum and in connected speech. Murton, Hillman, and Mehta (2020), analyzing 295 patients and 50 controls, identified Praat CPP cutoffs of 14.45 dB for sustained /a/ and 9.33 dB for the Rainbow Passage, with values below threshold indicating a voice disorder with up to 94.5% accuracy. Note the fine print: the ADSV cutoffs from the same paper are ~3 dB lower (11.46 and 6.11 dB)—CPP is as software-dependent as jitter, just better behaved. See our CPP guide for interpretation.
When Jitter and Shimmer ARE Still Valid
Despite their limitations, jitter and shimmer retain genuine value in specific contexts:
Type 1 signals (mild-to-moderate dysphonia)
When the spectrogram shows clear harmonic structure, perturbation measures are valid and can provide useful supplementary information about cycle-to-cycle stability (Titze, 1995; Ma & Yiu, 2005).
Sustained vowels at consistent, documented loudness
Voice SPL was the largest single influence on jitter and shimmer in healthy speakers (Brockmann et al., 2011). Comfortable-to-loud phonation, held consistent across sessions, is what makes values comparable.
Values below ~5%
Per the NCVS guideline, perturbation values under about 5% have been found reliable. If jitter is 1.2%, that is likely a real measurement. If jitter is 9%, be skeptical—that is probably the algorithm failing, not the voice.
As components within multiparametric indices
AVQI includes shimmer (Maryn, De Bodt, & Roy, 2010); DSI includes jitter (Wuyts et al., 2000). Within these validated composites, perturbation parameters contribute alongside cepstral and spectral measures. See our AVQI/ABI guide and DSI guide.
Group-level research comparisons
Many clinical populations show elevated group means relative to controls, which is why perturbation persists in research batteries. But group-level separation does not license individual-level diagnosis: distributions overlap substantially, and the measures say nothing about which disorder is present.
Praat vs. MDVP: Why Software Matters
Here's something many clinicians don't realize: Praat and MDVP produce systematically different values for the same voice sample (Maryn et al., 2009; Amir et al., 2009). This isn't a bug—it's a fundamental difference in how the algorithms find periods.
| Feature | Praat | MDVP |
|---|---|---|
| Period detection method | Waveform matching (cross-correlation) | Peak picking |
| Sensitivity to additive noise | Low (noise largely averaged out) | High (noise counted as jitter) |
| Typical jitter values | Lower, especially in noisy signals | Higher |
| Availability | Free, open source | Commercial (KayPENTAX CSL) |
The Dramatic Difference
Boersma (2009) demonstrated this with a synthesized test. On a constant-period glottal signal, both programs correctly report near-zero jitter, and on a signal with 1% genuine period variation, both correctly report 1%. But add 1% white noise to the constant-period signal—a quite ordinary amount of recording noise—and Praat reports jitter of 0.02% while MDVP reports 0.6%, more than half of MDVP's own pathological threshold, from the exact same sound. Praat's waveform matching separates period variation from noise; MDVP's peak picking lumps them together.
Bottom line: never compare Praat values to MDVP norms, or vice versa. Never combine results from different software in the same report.
What About Normative Values?
MDVP manual thresholds (peak picking)
Source: MDVP manual (KayPENTAX). Valid only for MDVP's own algorithm.
Praat-specific norms: mostly absent
There is no equivalently established set of Praat perturbation cutoffs. Comparative studies show Praat runs lower than MDVP on the same voices, with the gap widening as signals get noisier (Maryn et al., 2009; Boersma, 2009), so importing MDVP thresholds into Praat systematically misses pathology.
The defensible uses of Praat jitter and shimmer are therefore relative: within a patient across sessions (same task, same loudness, same software) or within a validated composite index. If you need an absolute threshold for disorder detection, use a measure with published Praat-specific cutoffs—CPP has them (Murton et al., 2020); Praat perturbation largely does not.
The Hidden Variable: Vocal Intensity
Perhaps the most underappreciated confound in perturbation measurement is sound pressure level (SPL). In 57 healthy adults, Brockmann and colleagues (2011) found that voice SPL was the most important factor affecting jitter and shimmer—larger than vowel choice, gender, or f0. Brockmann-Bauser, Bohlender, and Mehta (2018) then showed the same pattern in 58 female patients (nodules, polyps, muscle tension dysphonia) and 58 matched controls phonating at soft, comfortable, and loud levels: perturbation values improve as voices get louder, in patients and controls alike.
Why This Matters Clinically
Without loudness control, a patient's pathology can be masked by louder phonation—they sound disordered but the numbers look normal because they phonated loudly. Conversely, a healthy voice can look pathological at soft intensities. An apparent “improvement” between baseline and discharge may be nothing more than the patient phonating louder at discharge.
Clinical recommendation: elicit comfortable (not soft) phonation, measure and record SPL if you can, and at minimum document the loudness condition and keep it consistent across sessions. A jitter comparison across sessions at different loudness levels is not a comparison of the voice.
Which Jitter and Shimmer Variants Are Most Reliable?
If you've ever been confused by the alphabet soup of jitter variants (local, RAP, PPQ5) and shimmer variants (local, APQ3, APQ5, APQ11), here's what each one does. All are defined over consecutive cycles; they differ in how many periods the comparison is smoothed over:
| Variant | What It Does | Smoothing |
|---|---|---|
| Local Jitter | Compares adjacent cycles only | None (most noise-sensitive) |
| RAP (Relative Average Perturbation) | Smooths over 3 periods | 3 periods |
| PPQ5 (5-point Period Perturbation Quotient) | Smooths over 5 periods | 5 periods |
| Local Shimmer | Compares adjacent cycles only | None (most noise-sensitive) |
| APQ5 | Smooths over 5 periods | 5 periods |
| APQ11 | Smooths over 11 periods (MDVP default) | 11 periods (most smoothed) |
Practical Recommendation
Smoothed variants (RAP and PPQ5 for jitter; APQ5 and APQ11 for shimmer) average across several periods, which damps the effect of single mis-detected cycles; local variants react to everything, signal and artifact alike. Whichever you report, the essential disciplines are the same: report the variant by name, report the software, and compare only like with like. A “jitter” without a variant and a software name is not an interpretable number.
Recording Quality: The Precondition Nobody Mentions
Everything above assumes the recording itself is clean. Perturbation measures are the most noise-fragile numbers in the acoustic toolbox:
- Noise floor. Jitter and shimmer lose accuracy and reliability when the signal-to-noise ratio drops below about 30 dB (Deliyski et al., 2005)—a level that ordinary rooms with HVAC, traffic, or a computer fan routinely violate. And as the Boersma demonstration shows, what noise does to the reported value depends on which software you use.
- Live telehealth audio: never. Videoconferencing platforms add noise, compression, and noise-suppression processing. The foundational telepractice validity study did not even test jitter and shimmer, citing the SNR limit as disqualifying (Weerathunge et al., 2021). If you assess remotely, have the patient record locally and upload the file—see our telehealth acoustic assessment guide.
- Smartphone recordings: extra caution. A 2025 meta-analysis of smartphone versus clinical recording systems (10 studies, 379 participants) concluded that for some parameters—perturbation measures prominent among them—phone recordings do not yet match the precision of a clinical chain (Barsties v. Latoszek et al., 2025). Cepstral measures transfer better. Details in our smartphone recording guide.
Evidence-Based Protocol for Valid Measurements
If you're going to use jitter and shimmer, follow this protocol to maximize validity:
- 1
Record cleanly: quiet room, close microphone, lossless format
Target an SNR of at least 30 dB (Deliyski et al., 2005). No live-platform audio, no lossy voice notes, no Bluetooth microphones.
- 2
Check signal type FIRST
View a narrowband spectrogram. Clear harmonic structure (Type 1)? Proceed. Subharmonics, chaos, or noise domination (Types 2–4)? Stop—report CPP and perceptual ratings instead.
- 3
Use a sustained vowel of adequate length, at consistent loudness
3–5 seconds of steady /a/ at comfortable pitch and loudness gives the ~100 cycles the NCVS guideline calls for. Document the loudness condition, and keep it constant across sessions (Brockmann et al., 2011).
- 4
Exclude problem segments
Remove voice breaks, diplophonic segments, and onset/offset portions. Analyze the stable middle portion of the vowel.
- 5
Be skeptical of high values
Above roughly 5%, the NCVS guideline says the number is no longer reliable. Report such values with explicit caveats, or fall back on CPP and perceptual assessment.
- 6
Name the software, the version, and the variant
Never compare Praat values to MDVP cutoffs. Prefer within-patient, same-chain comparisons over absolute thresholds.
- 7
Always report CPP alongside perturbation measures
CPP is valid across the severity spectrum and in connected speech, and it has published Praat-specific cutoffs (Murton et al., 2020). Jitter and shimmer are the supplement, not the anchor.
Common Questions
Q: Should I stop using jitter and shimmer entirely?
No. The evidence supports continued use for Type 1 signals at controlled loudness, and as components of validated indices like AVQI and DSI. The key is knowing when they're valid and when they're not. Use CPP as your primary measure, and add jitter/shimmer as supplementary information when the signal type permits.
Q: My patient's voice sounds terrible but jitter/shimmer are normal. What's happening?
Several possibilities: (1) The voice is too disordered for valid measurement—the algorithm is returning artifacts, and severely aperiodic signals can produce paradoxically ordinary-looking numbers. (2) The patient phonated loudly, which lowers perturbation values (Brockmann-Bauser et al., 2018). (3) The pathology shows up in connected speech, which perturbation cannot analyze. (4) The disorder involves qualities perturbation doesn't capture—strain, tremor, breathiness. This is exactly why CPP on connected speech, plus your ears, outrank a sustained-vowel jitter value.
Q: I've been using MDVP cutoffs with Praat. Is that a problem?
Yes, unfortunately. Praat typically produces lower values than MDVP for the same voice, so MDVP thresholds applied to Praat data systematically miss pathology (false negatives). Use Praat values as relative, within-patient measures, or anchor disorder detection on a measure with Praat-specific published cutoffs, such as CPP (Murton et al., 2020).
Q: How do I explain this to referring physicians who expect jitter/shimmer?
Frame it as evolution, not abandonment. You might say: “Per the 2018 ASHA instrumental assessment protocol, we use CPP as our primary acoustic measure because it's valid across the severity spectrum and in connected speech. Jitter and shimmer are reported as supplementary values when the signal permits. CPP of 8.2 dB on sustained /a/ is below the published Praat cutoff of 14.45 dB for distinguishing disordered from healthy voices (Murton et al., 2020).”
Q: Does PhonaLab report jitter and shimmer?
Yes, along with CPP/CPPS and the standard acoustic parameters. PhonaLab computes them with Praat algorithms (via Parselmouth), so values sit on the Praat side of the software divide and must not be read against MDVP norms. Every threshold shown in the interface carries its published source, and analyses that a file cannot support are returned as flagged nulls rather than plausible-looking numbers.
Bottom Line: A Balanced Perspective
- 1Jitter and shimmer aren't obsolete—but they're supplementary, not primary (Patel et al., 2018)
- 2CPP is the anchor measure—valid across the severity spectrum, in connected speech, and with published software-specific cutoffs
- 3Validity is a property of the signal—Type 1 only; check the spectrogram before trusting the number
- 4Values above ~5% are suspect—likely algorithmic failure, not measurement (Titze, 1995)
- 5Software matters—Praat and MDVP disagree by design; never mix values and norms across programs
- 6Control loudness—SPL is the largest single confound; louder phonation lowers perturbation values in patients and controls alike
- 7Recording quality is a precondition—at least 30 dB SNR, lossless files, never live-platform audio, extra caution with phone-derived perturbation
- 8Multiparametric indices rehabilitate these measures—shimmer within AVQI and jitter within DSI contribute meaningfully inside validated composites
🔬 Get the Complete Acoustic Picture
PhonaLab computes CPP/CPPS, jitter, shimmer, HNR, f0, and AVQI from your uploaded recording in one analysis, using Praat algorithms, with every threshold traced to its published source. Recordings that cannot support a given measure return a flagged null with the reason—never a plausible-looking artifact. Audio is processed in memory and never stored.
Try Free Voice Analyzer →Praat-based algorithms • CPP + perturbation measures • AVQI multiparametric index
⚠️ Clinical Documentation Tool
The information in this article is provided for educational purposes and clinical workflow support. Acoustic measures should be interpreted within the context of comprehensive voice evaluation including perceptual assessment and patient history. Parameter validity depends on signal characteristics and recording conditions. All clinical decisions should be made by qualified healthcare professionals.
References & Further Reading
- Patel RR, Awan SN, Barkmeier-Kraemer J, Courey M, Deliyski D, Eadie T, Paul D, Švec JG, Hillman R. (2018). Recommended protocols for instrumental assessment of voice: American Speech-Language-Hearing Association expert panel to develop a protocol for instrumental assessment of vocal function. American Journal of Speech-Language Pathology, 27(3), 887–905. doi:10.1044/2018_AJSLP-17-0009
- Titze IR. (1995). Workshop on acoustic voice analysis: Summary statement. National Center for Voice and Speech. Available at https://ncvs.org/archive/freebooks/summary-statement.pdf
- Sprecher A, Olszewski A, Jiang JJ, Zhang Y. (2010). Updating signal typing in voice: Addition of type 4 signals. The Journal of the Acoustical Society of America, 127(6), 3710–3716.
- Murton O, Hillman R, Mehta D. (2020). Cepstral peak prominence values for clinical voice evaluation. American Journal of Speech-Language Pathology, 29(3), 1596–1607. doi:10.1044/2020_AJSLP-20-00001
- Boersma P. (2009). Should jitter be measured by peak picking or by waveform matching? Folia Phoniatrica et Logopaedica, 61(5), 305–308. doi:10.1159/000245159
- Maryn Y, Corthals P, De Bodt M, Van Cauwenberge P, Deliyski D. (2009). Perturbation measures of voice: A comparative study between Multi-Dimensional Voice Program and Praat. Folia Phoniatrica et Logopaedica, 61(4), 217–226. doi:10.1159/000227999
- Amir O, Wolf M, Amir N. (2009). A clinical comparison between two acoustic analysis softwares: MDVP and Praat. Biomedical Signal Processing and Control, 4(3), 202–205. doi:10.1016/j.bspc.2008.11.002
- Brockmann M, Drinnan MJ, Storck C, Carding PN. (2011). Reliable jitter and shimmer measurements in voice clinics: The relevance of vowel, gender, vocal intensity, and fundamental frequency effects in a typical clinical task. Journal of Voice, 25(1), 44–53. doi:10.1016/j.jvoice.2009.07.002
- Brockmann-Bauser M, Bohlender JE, Mehta DD. (2018). Acoustic perturbation measures improve with increasing vocal intensity in individuals with and without voice disorders. Journal of Voice, 32(2), 162–168.
- Deliyski DD, Shaw HS, Evans MK. (2005). Adverse effects of environmental noise on acoustic voice quality measurements. Journal of Voice, 19(1), 15–28. doi:10.1016/j.jvoice.2004.07.003
- Ma EPM, Yiu EML. (2005). Suitability of acoustic perturbation measures in analysing periodic and nearly periodic voice signals. Folia Phoniatrica et Logopaedica, 57(1), 38–47.
- Barsties v. Latoszek B, Maryn Y, Gerrits E, De Bodt M. (2018). A meta-analysis: Acoustic measurement of roughness and breathiness. Journal of Speech, Language, and Hearing Research, 61(2), 298–323.
- Barsties v. Latoszek B, Lammertz CZ, Awan SN, Binkofski F, Hetjens S. (2025). The accuracy of smartphone recordings for clinical voice diagnostics in acoustic voice quality assessments: A systematic review and meta-analysis. American Journal of Speech-Language Pathology, 34(6), 3531–3548. doi:10.1044/2025_AJSLP-25-00140
- Maryn Y, De Bodt M, Roy N. (2010). The Acoustic Voice Quality Index: Toward improved treatment outcomes assessment in voice disorders. Journal of Communication Disorders, 43(3), 161–174.
- Wuyts FL, De Bodt MS, Molenberghs G, Remacle M, Heylen L, Millet B, Van Lierde K, Raes J, Van de Heyning PH. (2000). The Dysphonia Severity Index: An objective measure of vocal quality based on a multiparameter approach. Journal of Speech, Language, and Hearing Research, 43(3), 796–809.
- Weerathunge HR, Segina RK, Tracy L, Stepp CE. (2021). Accuracy of acoustic measures of voice via telepractice videoconferencing platforms. Journal of Speech, Language, and Hearing Research, 64(7), 2586–2599. doi:10.1044/2021_JSLHR-20-00625