Back to Guides

Acoustic Voice Assessment via Telehealth: What Survives the Compression

May 10, 2026 (updated September 1, 2026)18 min readJorge C. Lucero

🎯 Key Takeaways

  • Mean f0 survives transmission. It was the only measure not significantly altered across all six platforms tested by Weerathunge et al. (2021), and it met validity criteria again in a 2025 home-recording replication (Hwang et al., 2025).
  • Noise-based measures do not. HNR dropped 5–9 dB and L/H ratio up to 9 dB after transmission, several times larger than the published difference between dysphonic and typical voices. Jitter and shimmer were not even tested because they fail below 30 dB SNR.
  • CPPS is the interesting case. It fell on every platform (1.4–2.3 dB in connected speech), but the shift was smaller than the dysphonic–typical gap. It is biased, not destroyed: never compare a live-capture CPPS to a normative cutoff.
  • Perceptual ratings survive. Experienced clinicians rated CAPE-V dimensions on transmitted samples with reliability and accuracy that were not substantially reduced (Dahl et al., 2021).
  • Platform rankings are perishable. Teams was least disruptive and Zoom-with-enhancements most disruptive in 2021, but every major platform has since replaced its audio pipeline with neural-network enhancement. Treat rankings as time-stamped.
  • Asynchronous recording is the solution. A phone or tablet with a wired headset microphone at a fixed distance produces CPP values correlating at r ≈ .99 with a research-grade chain (Awan et al., 2024). Record locally, upload the file, analyze the file.

Voice therapy over telehealth is no longer experimental. In the United States, Congress extended Medicare telehealth authority for speech-language pathologists through December 31, 2027, and the SLP services that had been covered under temporary flexibilities since 2020 became permanently authorized telehealth services on January 1, 2026 (ASHA, 2026). Outcome studies have supported telepractice delivery for Parkinson's disease (Constantinescu et al., 2011), muscle tension dysphonia (Rangarathnam et al., 2015), and vocal fold nodules (Fu et al., 2015). The therapy works.

The harder question is the assessment. Voice therapy progress is documented quantitatively: pre- and post-treatment acoustic measures, baseline-to-discharge comparisons, outcome scores. If those measurements are collected over Zoom or Microsoft Teams during a live session, can clinicians trust them? Or does the platform itself alter the signal in ways that invalidate the analysis?

This guide answers that question with the published evidence and outlines a practical protocol that works around the technical limitations of live videoconferencing. This September 2026 revision adds a home-based replication with real speakers (Hwang et al., 2025), a meta-analysis of smartphone recording accuracy (Barsties v. Latoszek et al., 2025), the Bridge2AI-Voice recording validation work (Awan et al., 2024), and a note on why the audio pipelines inside the platforms have changed since the foundational study was run. The short version has not changed: f0 and perceptual ratings survive transmission, noise-based measures do not, CPPS is systematically biased, and the right solution is asynchronous recording rather than live capture.

Why Videoconferencing Distorts Voice Measurements

Videoconferencing platforms are optimized for intelligible conversation, not for high-fidelity acoustic capture. To deliver clear speech under variable bandwidth, every major platform applies a chain of processing before the audio reaches the other end:

  • Lossy, bandwidth-limited coding. Voice-over-IP codecs discard information judged perceptually unimportant and reduce spectral bandwidth when the connection is constrained. Weerathunge et al. (2021) found that both uplink and downlink speeds predicted the size of measurement errors for spectral and cepstral measures.
  • Noise suppression. Algorithms attenuate what they classify as non-speech. Sustained vowels look, to a suppressor, a lot like steady-state noise: Weerathunge et al. observed cases where a platform suppressed sustained-vowel amplitude almost entirely, and other cases where the envelope was flattened to a constant peak.
  • Automatic gain control and level normalization. Level-dependent gain rewrites the amplitude envelope. This does not simply “compress dynamics”: in the 2021 data, Cisco WebEx and Zoom-with-enhancements inflated the SPL range of connected speech, while Zoom with enhancements turned off reduced it because of added noise.
  • Echo cancellation and de-reverberation. Nonlinear processing that changes spectral content, increasingly performed jointly with noise suppression by a single neural network.

Each of these processes is helpful for conversation and harmful for measurement. The result is a signal that sounds like the patient's voice but is acoustically a different object than what the microphone captured.

Noise suppression cannot tell breathiness from HVAC

Noise suppressors are trained to remove background noise. They have no way to distinguish ambient noise (fan hum, traffic) from voice-source noise (turbulent airflow through an incompletely closed glottis). For a patient with breathiness, the platform may be actively removing the acoustic signature the clinician is trying to measure. This is why HNR, L/H ratio, and CPPS, all of which quantify the noise-to-harmonic balance, are the measures that move the most after transmission.

The pipelines have changed since the evidence was collected

The foundational platform comparison was run in 2020–2021. Since then, the major platforms have moved from classical signal-processing enhancement to deep-learning models that perform echo cancellation, noise suppression, and de-reverberation jointly; Microsoft, for example, has published the model deployed in Teams (Ristea et al., 2023). Neural suppressors are more aggressive and less predictable than their predecessors, and they are updated silently. Any ranking of platforms by acoustic fidelity is therefore a snapshot, not a durable property of the product. The pattern of which measures are vulnerable is durable; the platform league table is not.

What the Evidence Shows

The most systematic study of this question remains Weerathunge, Segina, Tracy, and Stepp (2021). They took recordings from 29 speakers with dysphonia (Parkinson's disease, muscle tension dysphonia, adductor laryngeal dystonia, nodules, polyp, scar, paralysis; CAPE-V overall severity 0–64 mm), originally captured in a sound booth with a headset microphone at 44.1 kHz, and transmitted each one through six HIPAA-compliant configurations: Zoom with enhancements, Zoom without enhancements (“original sound”), Cisco WebEx, Microsoft Teams, Doxy.me, and VSee Messenger. The received audio was recorded at the far end and re-analyzed in Praat.

Two design details matter for interpretation. First, samples were played through a loudspeaker 58 cm from a laptop's built-in microphone, in three real home environments, with the computers' own audio enhancements switched off. That isolates the platform's effect but also represents a patient sitting at a laptop without a headset, which is a common but not optimal setup. Second, the authors set significance at p < .005 to control for multiple comparisons, so “not significant” is a conservative statement.

The measures were mean f0, f0 standard deviation, f0 range, SPL range (connected speech), HNR (sustained vowel), L/H spectral ratio and CPPS (both sustained vowel and connected speech), and relative fundamental frequency. Here is what happened to each:

MeasureEffect of transmissionSize of the shiftLive-capture verdict
Mean f0 (speech)Not significant on any platformMean shifts of 7–25 Hz were still present across platformsUsable, with same-method comparisons
f0 SD (speech)Significant, medium effect; increased on every platform except Teams, with large post hoc effects6–28 HzInflated; do not interpret
f0 range (speech)Significant, small effect; increased on Doxy.me, VSee, and Zoom-with-enhancements14–28 HzCaution
SPL range (speech)Significant, large effect. Inflated by WebEx (d = 3.5) and Zoom-with-enhancements; reduced by Zoom-original-sound; unchanged on Teams, Doxy.me, VSeeDirection depends on platformPlatform-dependent; never compare across platforms
HNR (vowel)Decreased on all six platforms, large effects (d = −0.96 to −1.48)5.4–8.8 dBInvalid
CPPS (vowel and speech)Decreased on all six platforms, large effects (d = −1.08 to −1.79)1.4–2.3 dB (speech)Systematically biased; never against a cutoff
L/H ratioDecreased on Doxy.me, VSee, Zoom-with-enhancements (speech) and WebEx (vowel)1.0–8.9 dB (vowel)Invalid
RFFCould not be analyzed: the automated algorithm rejected 7–75% of transmitted tokens as unusableNot measurable
Jitter, shimmerNot tested. Perturbation measures lose accuracy below 30 dB SNR (Deliyski et al., 2005); the original booth recordings averaged 30.7 dBDo not compute
AVQI, ABI, CSID, GNENot tested. These indices are built from CPPS, HNR, shimmer, spectral slope and tilt, and noise-band ratios, all of which are affected aboveNot defensible from live capture (inference, not a tested result)

How big is big? Weerathunge et al. put the shifts next to published between-group differences. The 5.4–8.8 dB drop in HNR dwarfs the roughly 1 dB difference reported between typical and dysphonic voices; the same is true for L/H ratio. CPPS in connected speech moved by 1.4–2.3 dB, which is less than the 2.62 dB separation between speakers with and without voice disorders reported by Sauder, Bretl, and Eadie (2017). The authors' own conclusion is that f0 measures and CPPS are the acoustic measures most likely to carry clinically relevant information over telepractice. Note the direction: every platform pushed CPPS down, toward the dysphonic side of any cutoff. A live-capture CPPS compared against a normative threshold will over-call dysphonia.

Platforms and settings. Microsoft Teams had the fewest effects: f0 measures and SPL range were not significantly altered, although CPPS and HNR still fell with large effect sizes. Zoom with enhancements had the most pronounced effects overall. Turning Zoom's enhancements off (“original sound”) did not rescue the noise-based measures: CPPS and HNR still dropped with large effects, and the un-suppressed signal carried enough ambient noise to reduce the SPL range. WebEx, VSee, and Doxy.me artificially sustained the amplitude of vowels, so connected-speech measures were less affected than sustained-vowel measures on those platforms.

Predictors of error. Ambient noise at the transmitting end was a significant predictor of the difference between transmitted and original values for every measure except f0. Internet speed mattered too: uplink speed predicted errors in L/H ratio and vowel CPPS, and downlink speed predicted errors in f0 mean and SD, SPL range, and CPPS. Voice severity also predicted error for CPPS, HNR, and f0 range, meaning the patients whose voices you most need to measure are the ones whose signals are most degraded. An earlier version of this guide stated that the room mattered more than the connection; the data support “both matter,” with the room the more controllable of the two.

A home-based replication with real speakers (2025)

The 2021 study used replayed recordings. Hwang and colleagues (2025) recorded 17 children with dysarthria due to cerebral palsy in their own homes, simultaneously via Zoom (laptop microphone at 30 cm) and via an offline recorder (8 cm), and tested nine measures against predefined validity criteria. Mean f0, SPL range, second-formant range of diphthongs, and duration-based measures met the criteria. f0 range, signal-to-noise ratio, and cepstral peak prominence did not. A different population, a different platform build four years later, and the same pattern: timing and f0 survive, noise-based measures do not.

Perceptual ratings survive transmission

In a companion study, Dahl and colleagues (2021) had 20 experienced clinicians (10 SLPs, 10 laryngologists) rate overall severity, roughness, breathiness, and strain on 20 voice samples after transmission through the same platforms, using a modified CAPE-V. Transmission produced statistically but not clinically significant differences in roughness ratings only. Overall severity had the highest inter-rater agreement and strain the lowest, as in in-person work. The authors concluded that telepractice transmission does not substantially reduce the reliability or accuracy of expert auditory-perceptual evaluation. Your ears are more robust than your algorithms.

A cautionary note on cross-session comparison

Even for a measure that is “preserved” on average, platform-specific biases mean that comparing a baseline collected in person with a follow-up collected over telepractice can produce spurious “improvements” or “deteriorations.” Whether a single platform's bias is stable enough over time to support within-patient trend monitoring is an open question that Weerathunge et al. explicitly flagged as untested, and platform updates since then make it less likely, not more. Use one collection method consistently, or use the asynchronous workflow below.

Why Asynchronous Recording Solves the Problem

The fundamental issue with live videoconferencing is that the audio is processed in real time by the platform before it reaches the clinician. The signal that arrives is not the signal the patient produced.

The workaround is to bypass the platform entirely for the recording itself. The patient records the voice sample locally on their own device, using an app that saves an uncompressed or losslessly compressed file (WAV, FLAC, or Apple Lossless), then uploads the file separately through a secure portal or compliant file transfer. The clinician analyzes the uploaded file rather than the live transmission. Weerathunge et al. recommended exactly this in their clinical implications.

The evidence that local recordings on consumer devices are good enough has firmed up considerably:

  • Bridge2AI-Voice validation (Awan et al., 2024). Across a diverse corpus of typical and disordered voices, CPP obtained from smartphones and tablets, with and without a low-cost headset microphone, correlated with a research-standard chain (GRAS 40AF microphone at 2.5 cm in a sound-treated booth) at a mean r of .99 (range .98–.99). Recording method still produced statistically significant offsets, and the L/H ratio was strongly affected by microphone type, but the relationships were linear and therefore correctable. Correlation is not identity: a phone CPPS is not a booth CPPS, and normative comparisons still require device-aware, pipeline-matched references.
  • Smartphones versus the gold standard (Awan et al., 2025). iPhone 12 and SE and Samsung S21 and S9, at 15 and 30 cm, in a booth and in an ordinary quiet office, produced cepstral and spectral measures comparable to a flat-response reference microphone across 24 speakers spanning sex, age, f0, and voice-quality types.
  • Meta-analysis (Barsties v. Latoszek et al., 2025). Ten studies, 379 participants, Apple and Samsung devices against clinical recording systems for jitter, shimmer, HNR, CPPS, and AVQI. The pooled picture is favorable but not uniform: the authors caution that for some parameters, current smartphone recordings do not yet match the precision of a clinical recording system. Treat perturbation measures from phones with more suspicion than cepstral ones.
  • Clinical implementation (Schneider et al., 2024). The UCSF Voice and Swallowing Center attempted remote recordings on 108 patients over six months with a phone app, and documented the process questions that matter in practice: which normative data to compare against, how to train clinicians, and which patient-side problems actually occur.
  • Formants (Zhang et al., 2021). In simultaneous recordings, f0 was tracked accurately by a field recorder, by lossless phone apps, and by Zoom, but formant estimates from phone apps were much closer to the reference than those from Zoom. If you use vowel-space or formant measures, they belong in the uploaded file, not the live session.

Asynchronous capture also has practical advantages that are independent of the fidelity argument: the patient can wear a headset at a fixed distance, the clinician can listen to the file before computing anything and request a re-recording if there is clipping or background noise, and each session uses the same device and protocol, which removes platform variability from longitudinal comparisons.

Recording specification that makes the evidence apply

  • Format: WAV, FLAC, or Apple Lossless. Not MP3, not a messaging-app voice note.
  • Sample rate: 44.1 kHz or 48 kHz. Cepstral and spectral-band measures have hard minimum sample rates (in PhonaLab: CPP and GNE require at least 10 kHz, CSID 16 kHz, AVQI and ABI 22.05 kHz); low-rate files from voice-note apps silently fall below them.
  • App processing off: Voice-memo apps now ship their own enhancement. In iOS Voice Memos, set Audio Quality to Lossless and keep “Enhance Recording” off. Check the equivalent on Android.
  • Wired, not Bluetooth: Bluetooth headsets route the microphone through a narrowband speech codec and their own processing, which reintroduces the problem you are avoiding.
  • Fixed distance: Headset boom at a constant off-axis distance, following the ASHA instrumental protocol (Patel et al., 2018); the Bridge2AI work used 2.5 cm for its reference chain. Consistency matters more than the exact number.
  • Room: Quiet, soft furnishings, away from windows and air handling. Ambient noise was the strongest patient-side predictor of error in every study cited here (Deliyski et al., 2005; Lebacq et al., 2017; Maryn et al., 2017; Marsano-Cornejo et al., 2021; Weerathunge et al., 2021).

A Practical Telehealth Workflow

For a clinician integrating acoustic assessment into a telepractice voice caseload, the following workflow balances measurement quality against patient burden:

1. Give the patient a one-page recording protocol

Screenshots, not prose: which app, the lossless setting, enhancement off, microphone placement, room setup, and the exact prompts. Patients will follow one page. They will not follow four.

2. Send a wired headset with a boom microphone

A consumer USB or 3.5 mm headset costing under US$30, mailed at intake or handed over at the first in-person visit. Low-cost electret headsets have been shown to capture the full range of voice quality despite non-flat frequency responses (Awan et al., 2022, cited in Awan et al., 2024). A fixed short distance is what makes the correlation evidence apply; it is the single highest-leverage investment in measurement quality.

3. Have the patient record before the session

Sustained /a/ at comfortable pitch and loudness (3–5 s, three trials), the CAPE-V sentences or the first paragraph of the Rainbow Passage (or the standard passage for the patient's language), and any task-specific items such as maximum phonation time. Files are uploaded a few hours before the live session. The clinician listens first, runs the analysis, and brings results to the session.

4. Use the live videoconference for everything else

Case history, perceptual evaluation (the evidence supports this), therapy practice, patient education, and outcome discussion all work over Zoom or Teams. The video session captures the clinical interaction; the uploaded file captures the measurement.

5. Document the recording chain in the chart

Device, app and settings, microphone, distance, room, and for any live-capture value the platform name and version. This is what makes a session-to-session comparison defensible.

When Live Telepractice Capture Is Acceptable

Asynchronous recording is the right default, but live capture has legitimate uses when the limitations are understood:

  • Mean f0 for biofeedback, including gender-affirming voice work. Mean f0 was preserved across platforms and populations; within-session pitch tracking during practice is valuable and does not require absolute accuracy. Do not extend this to f0 variability, which was inflated on most platforms.
  • Auditory-perceptual evaluation. Expert CAPE-V ratings on transmitted audio were not substantially degraded (Dahl et al., 2021). Rate live; measure from the file.
  • Initial triage when no asynchronous option exists. A live impression is better than none, provided the chart records that the assessment was platform-mediated and that acoustic indices were not collected.
  • Relative change within one session. Showing a patient their pitch or loudness trace in real time during a session is informative as a relative display, even though the absolute values are not.

The line to draw is between using the platform's audio for clinical interaction and quantifying the platform's audio as if it were a clean recording. The first is fine; the second is not.

Documentation and Defensibility

Telepractice acoustic data needs to be defensible if it ends up in an outcome report, a peer-reviewed publication, or a third-party payer audit. Three documentation practices help:

  • Describe the recording chain explicitly. “Patient recorded sustained /a/ and CAPE-V sentences on an iPhone with a wired boom headset at approximately 5 cm, Voice Memos at lossless quality with enhancement off, in a quiet home office, uploaded via [secure portal].”
  • Use the same chain for baseline and follow-up. If the device or microphone changes, document it and treat the first recording on the new chain as a new reference point.
  • Note any compromises. If a value was taken from live videoconferencing, say so and say what was not collected: “Mean f0 collected via Teams (version noted); CPPS, HNR, and AVQI not collected this session due to platform limitations.”

These practices make the methodology auditable and protect the clinician from later questions about whether an observed change reflects the voice or the measurement.

Summary

  1. 1Videoconferencing platforms alter acoustic measurements. Coding, noise suppression, and gain control all change the signal, and the neural pipelines now in use are more aggressive than the ones that were tested.
  2. 2Mean f0 and perceptual ratings survive; HNR, L/H ratio, f0 variability, and SPL range do not. CPPS is biased downward on every platform and must never be read against a cutoff.
  3. 3Composite indices are off the table for live capture. AVQI, ABI, CSID, and GNE inherit the errors of their components.
  4. 4Asynchronous recording is the practical solution. A phone or tablet with a wired headset at a fixed distance, lossless format, enhancement off, quiet room, file uploaded for analysis.
  5. 5Ambient noise is the most controllable source of error. Every study in this guide found it; the patient's room is the protocol element with the best return.
  6. 6Live videoconferencing is for clinical interaction, not quantification. Use the platform for the session and the uploaded file for the numbers.

📊 Analyze Uploaded Recordings in PhonaLab

PhonaLab accepts WAV, MP3, M4A, and other common audio formats directly in the browser, and computes f0, CPPS, HNR, jitter, shimmer, AVQI, ABI, and more from the uploaded file with published citations for every threshold. Every spectral-band measure checks the file's original sample rate first and returns a flagged null rather than a plausible-looking number if the file cannot support it. Audio is processed in memory and never stored on PhonaLab's servers.

Open Voice Analyzer →

For device-level detail on phones and tablets, see the companion guide: Smartphone Recording for Voice Assessment.

⚠️ Educational Information

This article summarizes published research on telepractice acoustic assessment for educational purposes. It does not constitute clinical advice, regulatory guidance, or billing recommendations. Payer and licensure rules vary by jurisdiction and change over time; the Medicare status described above reflects ASHA's published summary as of September 2026. Clinical decisions regarding voice assessment and telepractice protocols should be made by qualified, licensed professionals based on individual patient circumstances. PhonaLab provides acoustic measurement tools; it does not provide clinical interpretations or medical diagnoses.

References & Further Reading

  • Weerathunge HR, Segina RK, Tracy L, Stepp CE. (2021). Accuracy of acoustic measures of voice via telepractice videoconferencing platforms. Journal of Speech, Language, and Hearing Research, 64(7), 2586–2599. doi:10.1044/2021_JSLHR-20-00625
  • Dahl KL, Weerathunge HR, Buckley DP, Dolling AS, Díaz-Cádiz M, Tracy LF, Stepp CE. (2021). Reliability and accuracy of expert auditory-perceptual evaluation of voice via telepractice platforms. American Journal of Speech-Language Pathology, 30(6), 2446–2455. doi:10.1044/2021_AJSLP-21-00091
  • Hwang K, van Brenk F, McAuliffe MJ, Choi J, Švec JG, Chang YHM, Keller B, Levy ES. (2025). Validity of acoustic speech measures obtained through videoconferencing with children with dysarthria. International Journal of Speech-Language Pathology. Advance online publication. doi:10.1080/17549507.2025.2563845
  • Awan SN, Bahr R, Watts S, Boyer M, Budinsky R, Bridge2AI Voice Consortium, Bensoussan Y. (2024). Validity of acoustic measures obtained using various recording methods including smartphones with and without headset microphones. Journal of Speech, Language, and Hearing Research, 67(6), 1712–1730. doi:10.1044/2024_JSLHR-23-00759
  • Awan SN, Shaikh MA, Awan JA, Abdalla I, Lim KO, Misono S. (2025). Smartphone recordings are comparable to “gold standard” recordings for acoustic measurements of voice. Journal of Voice, 39(4), 1019–1032. doi:10.1016/j.jvoice.2023.01.031
  • Barsties v. Latoszek B, Lammertz CZ, Awan SN, Binkofski F, Hetjens S. (2025). The accuracy of smartphone recordings for clinical voice diagnostics in acoustic voice quality assessments: A systematic review and meta-analysis. American Journal of Speech-Language Pathology, 34(6), 3531–3548. doi:10.1044/2025_AJSLP-25-00140
  • Schneider SL, Habich L, Weston ZM, Rosen CA. (2024). Observations and considerations for implementing remote acoustic voice recording and analysis in clinical practice. Journal of Voice, 38(1), 69–76. doi:10.1016/j.jvoice.2021.06.011
  • Zhang C, Jepson K, Lohfink G, Arvaniti A. (2021). Comparing acoustic analyses of speech data collected remotely. The Journal of the Acoustical Society of America, 149(6), 3910–3916. doi:10.1121/10.0005132
  • Ristea NC, Indenbom E, Saabas A, Pärnamaa T, Guzhvin J, Cutler R. (2023). DeepVQE: Real time deep voice quality enhancement for joint acoustic echo cancellation, noise suppression and dereverberation. Proceedings of Interspeech 2023, 3819–3823. doi:10.21437/Interspeech.2023-1028
  • Sauder C, Bretl M, Eadie T. (2017). Predicting voice disorder status from smoothed measures of cepstral peak prominence using Praat and Analysis of Dysphonia in Speech and Voice (ADSV). Journal of Voice, 31(5), 557–566. doi:10.1016/j.jvoice.2017.01.006
  • Deliyski DD, Shaw HS, Evans MK. (2005). Adverse effects of environmental noise on acoustic voice quality measurements. Journal of Voice, 19(1), 15–28. doi:10.1016/j.jvoice.2004.07.003
  • Lebacq J, Schoentgen J, Cantarella G, Bruss FT, Manfredi C, DeJonckere P. (2017). Maximal ambient noise levels and type of voice material required for valid use of smartphones in clinical voice research. Journal of Voice, 31(5), 550–556. doi:10.1016/j.jvoice.2017.02.017
  • Maryn Y, Ysenbaert F, Zarowski A, Vanspauwen R. (2017). Mobile communication devices, ambient noise, and acoustic voice measures. Journal of Voice, 31(2), 248.e11–248.e23. doi:10.1016/j.jvoice.2016.07.023
  • Marsano-Cornejo MJ, Roco-Videla Á, Capona-Corbalán D, Silva-Harthey C. (2021). Variación del parámetro acústico harmonic-to-noise ratio en relación con distintos niveles de ruido de fondo. Acta Otorrinolaringológica Española, 72(3), 177–181. doi:10.1016/j.otorri.2020.04.007
  • Patel RR, Awan SN, Barkmeier-Kraemer J, Courey M, Deliyski D, Eadie T, Paul D, Švec JG, Hillman R. (2018). Recommended protocols for instrumental assessment of voice: American Speech-Language-Hearing Association expert panel to develop a protocol for instrumental assessment of vocal function. American Journal of Speech-Language Pathology, 27(3), 887–905. doi:10.1044/2018_AJSLP-17-0009
  • Constantinescu G, Theodoros D, Russell T, Ward E, Wilson S, Wootton R. (2011). Treating disordered speech and voice in Parkinson's disease online: A randomized controlled non-inferiority trial. International Journal of Language & Communication Disorders, 46(1), 1–16. doi:10.3109/13682822.2010.484848
  • Rangarathnam B, McCullough GH, Pickett H, Zraick RI, Tulunay-Ugur O, McCullough KC. (2015). Telepractice versus in-person delivery of voice therapy for primary muscle tension dysphonia. American Journal of Speech-Language Pathology, 24(3), 386–399. doi:10.1044/2015_AJSLP-14-0017
  • Fu S, Theodoros DG, Ward EC. (2015). Delivery of intensive voice therapy for vocal fold nodules via telepractice: A pilot feasibility and efficacy study. Journal of Voice, 29(6), 696–706. doi:10.1016/j.jvoice.2014.12.003
  • Grillo EU. (2019). Building a successful voice telepractice program. Perspectives of the ASHA Special Interest Groups, 4(1), 100–110. doi:10.1044/2018_PERS-SIG3-2018-0014
  • American Speech-Language-Hearing Association. (2026). Providing audiology and speech-language pathology telehealth services under Medicare. Retrieved September 1, 2026, from https://www.asha.org/practice/reimbursement/medicare/providing-telehealth-services-under-medicare/