Why One Number Isn't Enough: Dimension-Specific Voice Quality Indices
🎯 Key Takeaways
- A voice can be breathy without being rough, and strained without being breathy—but most acoustic indices collapse everything into a single severity number
- Four dimension-specific scores from one sustained /a/—severity, breathiness, roughness, and strain, each on the CAPE-V 100-point scale
- The roughness and strain indices are the first of their kind—no composite acoustic index had been published for either dimension before
- Selection was data-driven, not hand-picked—QR factorization removed redundant parameters, then orthogonal matching pursuit ranked the survivors per perceptual dimension
- These are hypothesis-generating tools, not diagnostic verdicts—derived from a single English-language database and awaiting external validation before clinical adoption
Listen to two patients with the same overall dysphonia severity. One sounds like air escaping through an incomplete glottal closure—breathy, weak, effortless in the wrong way. The other sounds like gravel— irregular, popping, rough. A trained ear separates these instantly. The CAPE-V and GRBAS protocols exist precisely because clinicians need to document which kind of deviant a voice is, not just how deviant.
Acoustic analysis has lagged behind that insight. The most successful multiparametric indices—AVQI for overall severity, ABI for breathiness—compress the acoustic signal into one number each. They do that job well. But if your patient's roughness matters to the treatment plan, or their strain is the target of therapy, no published composite acoustic index would give you a number for it.
That was the gap a recent study set out to close (Lucero, 2026). This guide explains what came out of it: four dimension-specific indices—severity, breathiness, roughness, and strain—computable from a single sustained /a/ vowel, now available in PhonaLab. It also explains, with equal care, what these indices are not.
One Number Was Never Enough
Perceptual voice assessment is explicitly multidimensional. The CAPE-V asks raters to score six qualities on separate 100 mm visual analog scales; GRBAS separates grade, roughness, breathiness, asthenia, and strain. The clinical rationale is straightforward: different qualities point toward different mechanisms, and different mechanisms call for different interventions. Breathiness suggests incomplete glottal closure. Roughness suggests irregular vocal fold vibration. Strain suggests hyperfunction.
On the acoustic side, the landscape has been lopsided. AVQI (Maryn et al., 2010) combines six parameters into a single severity score and has accumulated validation across more than a dozen languages. ABI (Barsties von Latoszek et al., 2017) targets breathiness with nine parameters. CSID (Awan et al., 2016) estimates overall severity from spectral and cepstral measures. All three are valuable. None of them tells you how rough or how strained a voice is.
This is not because roughness and strain lack acoustic correlates—individual parameters like jitter, shimmer, and spectral tilt have long been associated with them. It is because nobody had published a data-driven composite index for either dimension. Single parameters are noisy and unreliable in isolation; the whole point of the multiparametric approach is that combinations outperform any single measure. Roughness and strain simply never got their combination.
Four Indices, One Vowel
The four indices are linear combinations of acoustic parameters extracted from a sustained /a/ vowel of at least three seconds. Each index outputs a score on the CAPE-V 100-point scale—the same scale clinicians already use perceptually—where higher means more deviant. Each index uses a different, deliberately small set of parameters:
Severity Index — 6 parameters
Shimmer (dB, log), GNE, high-frequency noise (Hno-6000), period standard deviation (log), F0, and HNR-D. The omnibus index: responds to overall dysphonia regardless of type.
Breathiness Index — 3 parameters
CPPS, GNE, and Hno-6000. The leanest index—three noise-sensitive measures suffice to track turbulent airflow through an incompletely closed glottis.
Roughness Index — 5 parameters
Shimmer (dB, log), H1-H2, GNE, period standard deviation (log), and jitter (log). Weighted toward cycle-to-cycle irregularity. First published composite acoustic index for this dimension.
Strain Index — 6 parameters
HNR, F0, period standard deviation (log), spectral tilt, H1-H2, and GNE. Sensitive to the spectral reshaping of pressed, hyperfunctional phonation. First published composite acoustic index for this dimension.
Notice that GNE appears in all four indices—consistent with earlier findings that glottal-to-noise excitation carries information no other common parameter does (see our GNE guide). Period standard deviation appears in three. This sharing is not sloppiness; it reflects the real correlational structure of dysphonic voices. Perceptual dimensions overlap, and so do their acoustic footprints. The indices differ in weighting and in the parameters that are unique to each.
How the Indices Were Built
The methodology matters, because it is what separates these indices from an arbitrary weighting of someone's favorite parameters. The derivation used the Perceptual Voice Qualities Database (PVQD; Walden, 2022)—296 speakers with sustained vowels and CAPE-V ratings from multiple expert listeners—reduced to 284 after removing recordings with incomplete data.
Fourteen candidate parameters were extracted from each vowel: measures of perturbation (jitter, shimmer in % and dB, period standard deviation), noise (HNR, GNE, CPPS, high-frequency noise, HNR-D), spectral shape (slope, tilt, alpha ratio, H1-H2), and pitch (F0). Four right-skewed parameters were log-transformed. Then selection proceeded in two stages:
Redundancy analysis via QR factorization
An unsupervised step that identifies which parameters carry independent information and which merely duplicate others. Parameters that add nothing beyond what the rest already encode are dropped before any perceptual ratings are consulted.
Supervised ranking via orthogonal matching pursuit
For each perceptual dimension separately, OMP greedily selects the parameter that most improves prediction of the CAPE-V ratings, then the next, then the next—each choice accounting for what the already-selected parameters explain. The Bayesian information criterion determines where to stop, balancing fit against complexity.
Stability checks via bootstrap resampling
The selection was repeated across resampled versions of the dataset to confirm the chosen parameters were not artifacts of particular speakers. The final regressions were then evaluated with repeated 10-fold cross-validation.
The result is that each index's parameter set has a defensible answer to the question "why these and not others?"—the answer being that the data, not the author, made the selection. The full regression coefficients, analysis code, and derivation details are in the open-access paper and its public repository (Lucero, 2026).
How Accurate Are They?
Classification accuracy was assessed against the standard criterion of CAPE-V ≥ 10 on each dimension—the threshold for clinically relevant deviation—using area under the ROC curve (AUC) with repeated cross-validation:
| Index | Parameters | AUC vs CAPE-V ≥ 10 |
|---|---|---|
| Severity | 6 | 0.870 |
| Breathiness | 3 | 0.865 |
| Roughness | 5 | 0.830 |
| Strain | 6 | 0.790 |
For context, AVQI and CSID rescored on the same 284 speakers achieved AUCs of 0.827 and 0.787 against overall severity. The new Severity Index's 0.870 outperformed both on this sample—but that comparison deserves an honest asterisk, which brings us to the next section.
What These Indices Are Not
Every measurement tool earns trust by being clear about its boundaries. Four boundaries matter here:
1. The AVQI/CSID comparison is task-asymmetric
AVQI and CSID were designed to score connected speech plus a vowel; the new indices use the vowel alone. Comparing them on vowel-heavy material favors the vowel-only indices. The fair reading: the new Severity Index is a strong vowel-only alternative, not a demonstrated replacement for AVQI in speech material.
2. Roughness and strain are hypothesis-generating
Being first of their kind cuts both ways: there is no prior literature to compare against, and no independent sample has yet confirmed the derivation. The paper itself states that external validation is essential before clinical adoption. Treat these two scores as structured, reproducible observations—not established clinical instruments.
3. The derivation sample is English-language
The PVQD comprises American English speakers. Vowel acoustics transfer across languages better than connected speech does, but reference behavior in Brazilian Portuguese, Spanish, or other populations has not been established. Cross-language validation studies are in planning.
4. Agreement with perception, not diagnosis
The indices quantify agreement with expert auditory-perceptual judgment on the CAPE-V scale. They do not establish diagnostic validity against laryngoscopic or medical criteria. A high roughness score describes how the voice sounds, not what the larynx is doing.
Why lead with the limitations?
Because that is how measurement science works. AVQI needed years of independent validation studies before its current standing; these indices are at the start of that road, not the end. Using them now, with clear eyes about what they are, contributes to exactly the validation record they need. Overselling them would not.
A Different Kind of Check: Simulated Voices
External validation on human voices takes time—new samples, new raters, new studies. But there is a complementary check available now: physics-based simulation. Using a computational model of vocal fold vibration, we can synthesize voices in which a single mechanism is perturbed at a time—aspiration noise alone, jitter alone, vibratory asymmetry alone, hyperfunctional strain alone—and ask whether each dimension index responds selectively to its matching mechanism.
The logic is simple: if the Breathiness Index is really tracking breathiness, it should respond disproportionately when aspiration noise is dialed up and other mechanisms are held fixed—something impossible to arrange with human voices, where mechanisms co-occur. Preliminary results from this mechanism-selectivity analysis, using our physics-based simulator across controlled perturbation series in male and female voice configurations, show the predicted selectivity pattern in the large majority of test conditions, with one intriguing sex-specific exception under asymmetric vibration that we are investigating further. A full report is in preparation.
Why simulation-based checks matter
Human validation answers "does the index agree with listeners?" Simulation answers a different question: "does the index respond to the physical mechanism it claims to track?" The two are complementary, and to our knowledge no other composite voice quality index has been subjected to a controlled mechanism-selectivity test of this kind.
Using the Tool in Practice
The PhonaLab implementation is deliberately simple. Upload or record a sustained /a/ of at least three seconds at a sampling rate of 22,050 Hz or higher (44,100 Hz recommended). The tool returns the four scores, each with:
- A threshold interpretation against the same CAPE-V ≥ 10 criterion used in the derivation—below the threshold, or at/above it. No invented severity brackets: the paper defines one threshold, so the tool reports one threshold.
- The contributing parameters—the raw acoustic values that entered each score, so you can see why an index landed where it did. A high Breathiness score with a collapsed GNE tells a different story than one driven by low CPPS.
- Full methods and citation—every number traces to the source paper, and the analysis pipeline (Praat 6.1.38 via Parselmouth) is documented in the report footer.
One practical note: because the tool reproduces the paper's parameter extraction exactly, the raw parameter values shown in the expanders may differ slightly from those in PhonaLab's AVQI/ABI tool, which follows the Maryn & Latoszek protocol. Each tool is faithful to its own source publication—that is a feature, not a discrepancy.
Bottom Line: Four Numbers, Honestly Labeled
- 1Perceptual assessment is multidimensional; acoustic indices mostly weren't—roughness and strain had no composite index until now
- 2Four scores from one vowel—severity (AUC 0.870), breathiness (0.865), roughness (0.830), strain (0.790) against CAPE-V ≥ 10
- 3The selection was data-driven—QR redundancy analysis plus OMP ranking, with public code and open-access methods
- 4Read the scores as structured observations—hypothesis-generating for roughness and strain, vowel-only for severity, English-derived for all four
- 5Use the contributing parameters—the value is not just the score, but seeing which acoustic measures drove it
📐 Try the Dimension Indices Tool
Upload a sustained /a/ vowel and get all four dimension-specific scores—severity, breathiness, roughness, and strain—with contributing parameters, threshold interpretation, and full methods traceability. Computed through the same pipeline validated in the Journal of Voice.
Open Dimension Indices →Free to use with a PhonaLab account • PDF reports with Contributor
⚠️ Clinical Documentation Tool
The information in this article is provided for educational purposes and clinical documentation support. The dimension-specific indices described here quantify agreement with auditory-perceptual judgment and are intended to supplement—not replace—comprehensive voice evaluation including perceptual assessment, patient history, and laryngoscopic examination when indicated. The roughness and strain indices in particular await external validation and should be treated as hypothesis-generating. All clinical decisions should be made by qualified healthcare professionals. PhonaLab tools do not provide medical diagnoses.
References & Further Reading
- Lucero JC. (2026). Optimal Acoustic Parameter Subsets for Dimension-Specific Voice Quality Prediction. Journal of Voice. doi: 10.1016/j.jvoice.2026.07.051
- Lucero JC. (2026). Algorithm verification and concurrent validity of a web-based platform for multiparametric acoustic voice quality indices. Journal of Voice.doi: 10.1016/j.jvoice.2026.04.009
- Maryn Y, Corthals P, Van Cauwenberge P, Roy N, De Bodt M. (2010). Toward improved ecological validity in the acoustic measurement of overall voice quality: combining continuous speech and sustained vowels. Journal of Voice, 24(5), 540-555. doi:10.1016/j.jvoice.2008.12.014.
- Barsties von Latoszek B, Maryn Y, Gerrits E, De Bodt M. (2017). The Acoustic Breathiness Index (ABI): A multivariate acoustic model for breathiness. Journal of Voice, 31(4), 511.e11–511.e27. doi:10.1016/j.jvoice.2016.11.017.
- Awan SN, Roy N, Zhang D, Cohen SM. (2016). Validation of the Cepstral Spectral Index of Dysphonia (CSID) as a screening tool for voice disorders: Development of clinical cutoff scores. Journal of Voice, 30(2), 130–144. doi:10.1016/j.jvoice.2015.04.009.
- Walden PR. (2022). Perceptual Voice Qualities Database (PVQD): Database characteristics. Journal of Voice, 36(6), 875.e15–875.e23. doi:10.1016/j.jvoice.2020.10.001.
- Kempster GB, Gerratt BR, Verdolini Abbott K, Barkmeier-Kraemer J, Hillman RE. (2009). Consensus auditory-perceptual evaluation of voice: Development of a standardized clinical protocol. American Journal of Speech-Language Pathology, 18(2), 124–132. doi:10.1044/1058-0360(2008/08-0017).