Skip to main content

Research

Prosodically-conditioned temporal variation in whispered vs. normal speech

Whispered speech prototypically lacks vocal fold vibration and thus exhibits substantial spectral differences from speech with normal vocalization, including the absence of F0-related prosodic cues [1].

 

Furthermore, whispered speech is aerodynamically disadvantaged: subglottal pressure is exhausted more rapidly during whispered speech, due to higher airflow [2].

 

These spectral and aerodynamic factors make competing predictions regarding temporal differences between whispered and normal speech: on one hand, speakers may compensate for impoverished prosodic cues by extending duration in prosodically strong positions; on the other hand, they may address aerodynamic needs by lengthening stop closures and shortening high airflow segments like vowels and fricatives.

 

We tested these hypotheses by examining temporal differences between whispered and normal productions of The North Wind and the Sun passage and analyzing their segmental, prosodic, and global temporal differences.

 

We found that speech rate was slower in whispered speech, but most speakers paused less; furthermore, durational increases primarily occurred in phrase-final and phrase-initial prosodic positions. Overall, the results indicate that the main prosodic adjustment of whispered speech is an exaggeration of phrase-final lengthening, suggesting compensation for the absence of F0-related prosodic cues. 

Date

2026

Links