Opens in a new tab

AES E-Library

← Back to search

Convention Paper

Joint Neural Translation and Classification of Videos for Audio Processing

Authors: Bharitkar, Sunil; Cajica, Alejandro

Convention Paper · Paper 10261 · May 2026

Abstract

A low-parameter-count machine-learning model for classifying streaming video can enable content-aware audio/video processing on consumer edge devices with latency, computational, and battery constraints. In this paper, we propose a low-compute classification technique that uses only text metadata from the streaming file header, enabling near-instantaneous inference without decoding and analyzing audio or video signals as is traditionally done. In particular, to support multilingual platforms such as YouTube, we first apply neural machine translation as
a pre-processing step for the text metadata and optimize a lightweight neural classifier for a three-class audio-centric classification taxonomy (movie, music, dialog/other). Experiments on a mixed-language YouTube dataset achieve 90% classification accuracy on a test set using a combined translation and a classification model (with only 22K parameters), demonstrating a globally-scalable approach for robust classification on the edge.

Details

AES Convention
160
Paper number
10261
Publication date
May 28, 2026
Session subject
AI and Machine Learning in Audio, Audio Processing
Affiliation
Digital Media Solutions, Audio Lab, Samsung Research America; Samsung Research, Mexico (See document for exact affiliation information.)
Type
Convention Paper