Conference Paper
Open Access
A Scalable AI Architecture for Audio and Multimodal Analysis on Mobile Devices: A Case of Environmental Monitoring
2025 AES International Conference on Artificial Intelligence and Machine Learning for Audio · Paper 13 · September 2025
Abstract
The increasing need for real-time environmental monitoring has made Artificial Intelligence (AI) and Machine Learning (ML) models for audio classification and multimodal sensing essential for detecting and analyzing pollution-related sounds. In cases where mobile devices are used for capturing and processing audiovisual content, such models offer significant potential for research, storytelling, and public awareness. This study proposes a scalable and modular architecture that enables direct and flexible access to AI models for multimodal and audio processing. A CNN-LSTM hybrid model is trained for real-time environmental sound classification and deployed as a service. Building on prior 1D CNN-based approaches and incorporating temporal dependencies, the new model achieves an AUC of 0.91, demonstrating improved accuracy and generalization. The system leverages a REST API and Docker-based containerization, allowing deployment of independent AI services and supporting mobile and IoT use cases. The architecture accommodates both pre-trained and custom models, accessible from any device via a unified interface, confirming that a hybrid CNN-LSTM topology can support effective real-time sound classification, within a modular, containerized framework.
