Opens in a new tab

AES E-Library

← Back to search

Conference Paper

DeepEarNet: Individualizing Spatial Audio with Photography, Ear Shape Modeling, and Neural Networks

Authors: Kaneko, Shoken; Suenaga, Tsukasa; Sekine, Satoshi

AES Conference: 2016 AES International Conference on Audio for Virtual and Augmented Reality · Paper 6-3 · September 2016

Abstract

Individualizing spatial audio is of crucial importance for high-quality virtual and augmented reality audio. In this paper we propose a method for individualizing spatial audio by combining the recently proposed ear shape modeling technique with computer vision and machine learning. We use a convolutional neural network to obtain estimates of the ear shape model parameters from stereo photographs of the user ear. The individualized ear shape and its associated individualized head-related transfer function (HRTF) can be calculated from the obtained parameters based on the ear shape model and numerical acoustic simulations. Preliminary experiments, evaluating the shapes of the estimated individual ears, proved the effect of individualization.

Details

Published in
AES Conference: 2016 AES International Conference on Audio for Virtual and Augmented Reality
Paper number
6-3
Publication date
September 6, 2016
Session subject
Perceptual Consideration for VR/AR
Affiliation
Yamaha Corporation, Iwata-shi, Japan (See document for exact affiliation information.)
Type
Conference Paper