Opens in a new tab

AES E-Library

← Back to search

Convention Paper

Mapping voice gender and emotion to acoustic properties of natural speech

Authors: Oh, Eunmi; Lee, Jaeeun; Lee, Dayoung

AES Convention 150 · Paper 10461 · May 2021

Abstract

This study is concerned with listener’s natural ability to identify an anonymous speaker’s gender and emotion from voice alone. We attempt to map psychological characteristics of the speaker, such as gender image and emotion, to acoustical properties. The acoustical parameters of voice samples were pitch (mean, maximum, and minimum), pitch variation over time, jitter, shimmer, and Harmonics-to-Noise Ratio (HNR). Participants listened to 2-second voice clips and were asked to rate each voice’s gender image and emotion using a 7-point scale. Emotional responses were obtained for 7 opposite pairs of affective attributes (Goble and Ni Chasaide, 2003). The pairs of affective attributes were relaxed/stressed, content/angry, friendly/hostile, sad/happy, bored/interested, intimate/formal, and timid/confident. Experimental results show that listeners were able to identify voice gender and assess emotional status from short utterances. Statistical analyses revealed that these acoustic parameters were related to listeners’ perception of a voice’s gender image and its affective attributes. For voice gender perception, there were significant correlations with jitter, shimmer, and HNR parameters in addition to pitch parameters. For perception of affective attributes, acoustic parameters were analyzed with respect to the valence-arousal dimension. Voices perceived as positive tended to have higher variance in pitch and higher maximum pitch than those perceived as negative. Voices perceived as strongly active tended to have higher number of voice breaks, jitter, shimmer, and lower HNR than those perceived as passive. We expect that our experimental results on mapping acoustical parameters with voice gender and emotion perception could be applied to the field of Artificial Intelligence (AI) when assigning specific tone or quality to voice agents. Moreover, such psycho-acoustical mapping can improve the naturalness of synthesized speech, especially neural TTS (Text-To-Speech), because it can assist in selecting the appropriate speech database for voice interaction and for situations where certain voice gender and affective expressions are needed.

Details

Published in
AES Convention 150
AES Convention
150
Paper number
10461
Publication date
May 6, 2021
Session subject
Psychology
Affiliation
Yonsei University, Seoul, Korea (See document for exact affiliation information.)
Type
Convention Paper