Fetching the paper…
Reading the bibliography…
What audio embedding approach generalizes best to a wide range of downstream tasks across a variety of everyday domains without fine-tuning? The aim of the HEAR benchmark is to develop a general-purpose audio representation that provides a strong basis for learning in a wide variety of tasks and scenarios.
A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Neil Houlsby · 1910
Earlier work this paper cites.
Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
Steven Davis and Paul Mermelstein · 1980
Earlier work this paper cites.
Mel frequency cepstral coefficients for music modeling
Beth Logan · 2000
Earlier work this paper cites.
A Neural Probabilistic Language Model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2001
Earlier work this paper cites.
Musical genre classification of audio signals
George Tzanetakis and Perry Cook · 2002
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Sumit Chopra, Raia Hadsell, and Yann LeCun · 2005
Earlier work this paper cites.
Interpretable Meta-Measure for Model Performance
Alicja Gosiewska, Katarzyna Woznica, and Przemyslaw Biecek · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
FSD50K: an Open Dataset of Human-Labeled Sound Events
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Constant-Q transform toolbox for music processing
Christian Schörkhuber and Anssi Klapuri · 2010
Earlier work this paper cites.
Word Representations: A Simple and General Method for Semi-Supervised Learning
Joseph Turian, Lev-Arie Ratinov, and Yoshua Bengio · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa · 2011
Earlier work this paper cites.
On Random Weights and Unsupervised Feature Learning
Andrew M Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y Ng · 2011
Earlier work this paper cites.
Modal analysis and transcription of strokes of the Mridangam using non-negative matrix factorization
Akshay Anantapadmanabhan, Ashwin Bellur, and Hema A Murthy · 2013
Earlier work this paper cites.
The INTERSPEECH 2013 Computational Paralinguistics Challenge – A Brief Review
Björn Schuller, Stefan Steidl, and Anton Batliner · 2013
Earlier work this paper cites.
The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use
Bob L. Sturm · 2013
Earlier work this paper cites.
Deep Scattering Spectrum
Joakim Andén and Stéphane Mallat · 2014
Earlier work this paper cites.
CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Ten years of MIREX: reflections, challenges and opportunities
J Stephen Downie, Xiao Hu, Jin Ha Lee, Kahyun Choi, Sally Jo Cunningham, and Yun Hao · 2014
Earlier work this paper cites.
A study of instrument-wise onset detection in Beijing Opera percussion ensembles
Mi Tian, Ajay Srinivasamurthy, Mark Sandler, and Xavier Serra · 2014
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
Karol J Piczak · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
On the Potential of Simple Framewise Approaches to Piano Transcription
Rainer Kelz, Matthias Dorfer, Filip Korzeniowski, Sebastian Böck, Andreas Arzt, and Gerhard Widmer · 2016
Earlier work this paper cites.
Metrics for polyphonic sound event detection
Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Adieu Features? End-to-End Speech Emotion Recognition using a Deep Convolutional Recurrent Network
George Trigeorgis, Fabien Ringeval, Raymond Brückner, Erik Marchi, Mihalis Nicolaou, Björn Schuller, and Stefanos Zafeiriou · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Snore Sound Classification Using Image-based Deep Spectrum Features
Shahin Amiriparian, Maurice Gerczuk, Sandra Ottl, Nicholas Cummins, Michael Freitag, Sergey Pugachevskiy, and Björn Schuller · 2017
Cited alongside, same era.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
Jesse H. Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan · 2017
Cited alongside, same era.
Audio Set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel P W Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
CNN Architectures for Large-Scale Audio Classification
ERASER: A Benchmark to Evaluate Rationalized NLP Models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace · 2020
Later among the works it cites.
DDSP: Differentiable Digital Signal Processing
Jesse Engel, Lamtharn (Hanoi) Hantrakul, Chenjie Gu, and Adam Roberts · 2020
Later among the works it cites.
Bootstrap Your Own Latent – A New Approach to Self-Supervised Learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko · 2020
Later among the works it cites.
PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley · 2020
Later among the works it cites.
Towards Learning a Universal Non-Semantic Representation of Speech
Joel Shor, Aren Jansen, Ronnie Maor, Oran Lang, Omry Tuval, Felix de Chaumont Quitry, Marco Tagliasacchi, Ira Shavitt, Dotan Emanuel, and Yinnon Haviv · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shawn Hershey, Sourish Chaudhuri, Daniel P W Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, Malcolm Slaney, Ron J Weiss, and Kevin Wilson · 2017
Cited alongside, same era.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
Soroush Merhi, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron C. Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
Detection and classification of acoustic scenes and events: Outcome of the dcase 2016 challenge
Annamaria Mesaros, Toni Heittola, Emmanouil Benetos, Peter Foster, Mathieu Lagrange, Tuomas Virtanen, and Mark D Plumbley · 2017
Cited alongside, same era.
Deep convolutional neural networks and data augmentation for environmental sound classification
Justin Salamon and Juan Pablo Bello · 2017
Cited alongside, same era.
Onsets and Frames: Dual-Objective Piano Transcription
Curtis Hawthorne, Erich Elsen, Jialin Song, Adam Roberts, Ian Simon, Colin Raffel, Jesse Engel, Sageev Oore, and Douglas Eck · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
What Makes for Good Views for Contrastive Learning?
Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola · 2020
Later among the works it cites.
Self-Supervised Learning of Audio Representations From Permutations With Differentiable Ranking
Andrew N Carr, Quentin Berthet, Mathieu Blondel, Olivier Teboul, and Neil Zeghidour · 2021
Later among the works it cites.
Exploring Simple Siamese Representation Learning
Xinlei Chen and Kaiming He · 2021
Later among the works it cites.
Revisiting the Onsets and Frames Model with Additive Attention
Kin Wai Cheuk, Yin-Jyun Luo, Emmanouil Benetos, and Dorien Herremans · 2021
Later among the works it cites.
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexander Kolesnikov, Alexey Dosovitskiy, Dirk Weissenborn, Georg Heigold, Jakob Uszkoreit, Lucas Beyer, Matthias Minderer, Mostafa Dehghani, Neil Houlsby, Sylvain Gelly, Thomas Unterthiner, and Xiaohua Zhai · 2021
Later among the works it cites.
Efficient Training of Audio Transformers with Patchout
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer · 2021
Later among the works it cites.
On Generative Spoken Language Modeling from Raw Audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Later among the works it cites.
Pay Attention to MLPs
Hanxiao Liu, Zihang Dai, David So, and Quoc Le · 2021
Later among the works it cites.
Learning Audio Representations With MLPs
Mashrur M Morshed, Ahmad Omar Ahsan, Hasan Mahmud, and Md. Kamrul Hasan · 2021
Later among the works it cites.
BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Contrastive learning of general-purpose audio representations
Aaqib Saeed, David Grangier, and Neil Zeghidour · 2021
Later among the works it cites.
Pay Attention to MLPs
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Peter Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
VOXLINGUA107: A Dataset for Spoken Language Recognition
Jörgen Valk and Tanel Alumäe · 2021
Later among the works it cites.
Multi-Format Contrastive Learning of Audio Representations
Luyu Wang and Aaron van den Oord · 2021
Later among the works it cites.
SUPERB: Speech Processing Universal PERformance Benchmark
Shu-Wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-Tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung-Yi Lee · 2021
Later among the works it cites.
Barlow Twins: Self-Supervised Learning via Redundancy Reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny · 2021
Later among the works it cites.
LEAF: A Learnable Frontend for Audio Classification
Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, and Marco Tagliasacchi · 2021
Later among the works it cites.
VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
Adrien Bardes, Jean Ponce, and Yann LeCun · 2022
Closest in time.
Byol-s: Learning self-supervised speech representations by bootstrapping
Gasser Elbanna, Neil Scheidwasser-Clow, Mikolaj Kegler, Pierre Beckmann, and Milos Cernak · 2022
Closest in time.
Vision models are more robust and fair when pretrained on uncurated images without supervision
Priya Goyal, Quentin Duval, Isaac Seessel, Mathilde Caron, Ishan Misra, Levent Sagun, Armand Joulin, and Piotr Bojanowski · 2022
Closest in time.