Specific Loudness Features for Sound Event Classification and Localization in Machine Learning
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
This thesis investigates specific loudness representations as input features for machine learning systems that classify and localize sound events. Mel spectrograms are the de factostandard in audio machine learning, yet they are an incomplete perceptual transform: they
warp the frequency axis to match human hearing but leave the magnitude axis in raw logpower, omitting the equal-loudness weighting, spectral masking, temporal integration, and
compressive loudness growth of the auditory system. This work completes the transform by
employing the ISO 532-1 (Zwicker) and ISO 532-3 (Moore-Glasberg-Schlittenlacher) specific
loudness models, which map both the frequency and magnitude axes into perceptual units.
The thesis is primarily a study of sound event classification. On the ESC-50 benchmark,
specific loudness features significantly outperform both linear STFT and mel spectrograms
under an identical CNN, with accuracies of 39.7% for STFT, 53.1% for mel, 59.5% for
ISO 532-1, and 63.0% for ISO 532-3. A magnitude-axis ablation localizes the cause: applying a perceptual transform to the magnitude of an ordinary mel spectrogram—a simple
power-law compression, optionally combined with a fixed equal-loudness weighting, on the
same mel filterbank—recovers most of the gain at essentially mel-level cost and without
computing the full ISO 532 loudness model. Because these variants hold the mel frequency
scale fixed and vary only the magnitude mapping, they isolate the perceptual transform ofthe magnitude axis, rather than the auditory frequency scale, as the primary driver. The
advantage replicates across two further datasets: UrbanSound8K, a second environmentalsound-event benchmark, and the TAU Urban Acoustic Scenes corpus, where it holds at
roughly 14–16 percentage points over mel and extends even to acoustic scene classification,
a task outside the sound-event framing of the thesis statement.
A second study extends the investigation to spatial audio as an exploratory probe. In
a Sound Event Localization and Detection (SELD) experiment on STARSS23, 2-channel
binaural specific loudness does not match the published 4-channel audio-only first-order ambisonics baseline (ER = 0.97). Instead, the ISO 532-3 model behaves as a more conservative
detector: it obtains a lower error rate than our own FOA reproduction, but with lower
F-score and localization recall. A multi-threshold analysis indicates that this lower error
rate reflects conservative detection—fewer false alarms—rather than better localization, and
that the F-score deficit persists even when direction of arrival is ignored. Classification, not
localization, is the decisive result of the thesis.
Together, these results support the thesis that for sound event classes defined by human
perception, perceptual input representations that include level-dependent spectral weighting
align more strongly with classification targets than representations based solely on physical
signal properties or partial perceptual transforms.
Description
Thesis (Master's)--University of Washington, 2026
