2021

Audio Transformers

Verma, Prateek, Berger, Jonathan

Understand

Over the past two decades, CNN architectures have produced compelling models of sound perception and cognition, learning hierarchical organizations of features.

  • Analogous to successes in computer vision, audio feature classification can be optimized for a particular task of interest, over a wide variety of datasets and labels.
  • In fact similar architectures designed for image understanding have proven effective for acoustic scene analysis.
  • Here we propose applying Transformer based architectures without convolutional layers to raw audio signals.

Reading the bibliography…