Specific Loudness Features for Sound Event Classification and Localization in Machine Learning

relationships.isAuthorOf

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

This thesis investigates specific loudness representations as input features for machine learning systems that classify and localize sound events. Mel spectrograms are the de factostandard in audio machine learning, yet they are an incomplete perceptual transform: they warp the frequency axis to match human hearing but leave the magnitude axis in raw logpower, omitting the equal-loudness weighting, spectral masking, temporal integration, and compressive loudness growth of the auditory system. This work completes the transform by employing the ISO 532-1 (Zwicker) and ISO 532-3 (Moore-Glasberg-Schlittenlacher) specific loudness models, which map both the frequency and magnitude axes into perceptual units. The thesis is primarily a study of sound event classification. On the ESC-50 benchmark, specific loudness features significantly outperform both linear STFT and mel spectrograms under an identical CNN, with accuracies of 39.7% for STFT, 53.1% for mel, 59.5% for ISO 532-1, and 63.0% for ISO 532-3. A magnitude-axis ablation localizes the cause: applying a perceptual transform to the magnitude of an ordinary mel spectrogram—a simple power-law compression, optionally combined with a fixed equal-loudness weighting, on the same mel filterbank—recovers most of the gain at essentially mel-level cost and without computing the full ISO 532 loudness model. Because these variants hold the mel frequency scale fixed and vary only the magnitude mapping, they isolate the perceptual transform ofthe magnitude axis, rather than the auditory frequency scale, as the primary driver. The advantage replicates across two further datasets: UrbanSound8K, a second environmentalsound-event benchmark, and the TAU Urban Acoustic Scenes corpus, where it holds at roughly 14–16 percentage points over mel and extends even to acoustic scene classification, a task outside the sound-event framing of the thesis statement. A second study extends the investigation to spatial audio as an exploratory probe. In a Sound Event Localization and Detection (SELD) experiment on STARSS23, 2-channel binaural specific loudness does not match the published 4-channel audio-only first-order ambisonics baseline (ER = 0.97). Instead, the ISO 532-3 model behaves as a more conservative detector: it obtains a lower error rate than our own FOA reproduction, but with lower F-score and localization recall. A multi-threshold analysis indicates that this lower error rate reflects conservative detection—fewer false alarms—rather than better localization, and that the F-score deficit persists even when direction of arrival is ignored. Classification, not localization, is the decisive result of the thesis. Together, these results support the thesis that for sound event classes defined by human perception, perceptual input representations that include level-dependent spectral weighting align more strongly with classification targets than representations based solely on physical signal properties or partial perceptual transforms.

Description

Thesis (Master's)--University of Washington, 2026

Citation

DOI