Opens in a new tab

AES E-Library

← Back to search

Journal Article

Audiovisual Congruence and Localization Performance in Virtual Reality: 3D Loudspeaker Model vs. Human Avatar

Authors: Hofmann, Anja; Meyer-Kahlen, Nils; Schlecht, Sebastian J.; Lokki, Tapio

Journal of the Audio Engineering Society · Volume 72 · Issue 10 · pp. 679–690 · October 2024

Abstract

This paper investigates audiovisual congruence in virtual reality with both horizontal and vertical offsets between audio and visual rendering. Audiovisual congruence and localization errors are assessed using loudspeaker playback and nonindividualized headphone rendering. To account for the influence of different types of visual information on congruence, presentations of a loudspeaker model and 3D human avatar were compared. Therefore, a new dataset of audiovisual speech was recorded. Results show that human avatar rendering increases perceived congruence, and experienced listeners have an increased tendency to respond with “incongruent” when a loudspeaker model is shown but not when the human avatar is presented. Moreover, a correlation is found between localization precision and audiovisual congruence for horizontally offset stimuli and avatar presentation. For vertical offsets, the angular range of congruence is generally large, and localization errors are high, so no correlation can be observed between the two. The paper contributes congruence ranges for audiovisual speech in virtual reality, which also has implications for augmented reality telepresence use.

Details

Publication
Journal of the Audio Engineering Society
Volume
72
Issue
10
Pages
679–690
Publication date
October 15, 2024
Affiliation
Aalto Acoustics Lab, Department of Information and Communications Engineering, Aalto University, and Media Lab (See document for exact affiliation information.)
Type
Journal Article