Fetching the paper…
Reading the bibliography…
Limited diversity in standardized benchmarks for evaluating audio representation learning (ARL) methods may hinder systematic comparison of current methods' capabilities.
“Evaluation of algorithms using games: The case of music tagging.,”
Edith Law, Kris West, Michael I Mandel, Mert Bay, and J Stephen Downie, · 2009
Earlier work this paper cites.
“A comparison of sound segregation techniques for predominant instrument recognition in musical audio signals.,”
Juan J Bosch, Jordi Janer, Ferdinand Fuhrmann, and Perfecto Herrera, · 2012
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research,”
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello, · 2014
Earlier work this paper cites.
“EMOVO corpus: an Italian emotional speech database,”
Giovanni Costantini, Iacopo Iaderola, Andrea Paoloni, and Massimiliano Todisco, · 2014
Earlier work this paper cites.
“Esc: Dataset for environmental sound classification,”
Karol J. Piczak, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Deep convolutional networks on the pitch spiral for musical instrument recognition,”
Vincent Lostanlen and Carmine-Emanuele Cella, · 2016
Earlier work this paper cites.
“Fma: A dataset for music analysis,”
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, and Xavier Bresson, · 2017
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F. Gemmeke and et al., · 2017
Earlier work this paper cites.
“The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R. Livingstone and Frank A. Russo, · 2018
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2018
Earlier work this paper cites.
“Learning Problem-Agnostic Speech Representations from Multiple Self-Supervised Tasks,”
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio, · 2019
Earlier work this paper cites.
“wav2vec 2.0: a framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“SLURP: A spoken language understanding resource package,”
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn and et al., · 2020
Cited alongside, same era.
“Codified audio language modeling learns useful representations for music information retrieval,”
Rodrigo Castellon, Chris Donahue, and Percy Liang, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu and et al., · 2021
Cited alongside, same era.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
“Towards learning universal audio representations,”
Luyu Wang and et a., · 2022
Later among the works it cites.
“Decorrelating feature spaces for learning general-purpose audio representations,”
Sreyan Ghosh, Ashish Seth, and S Umesh, · 2022
Later among the works it cites.
“Hear: Holistic evaluation of audio representations,”
Joseph Turian and et al., · 2022
Later among the works it cites.
“Fsd50k: An open dataset of human-labeled sound events,”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra, · 2022
Later among the works it cites.
“The variably intense vocalizations of affect and emotion (vivae) corpus prompts new perspective on nonspeech perception.,”
Natalie Holz, Pauline Larrouy-Maestri, and David Poeppel, · 2022
Later among the works it cites.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shu wen Yang and et al., · 2021
Cited alongside, same era.
“Lebenchmark: A reproducible framework for assessing self-supervised representation learning from speech,”
Solène Evain and et al., · 2021
Cited alongside, same era.
“GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio,”
Guoguo Chen and et al., · 2021
Cited alongside, same era.
“Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”
Changhan et al. Wang, · 2021
Cited alongside, same era.
“ACAV100M: Automatic curation of large-scale datasets for audio-visual video representation learning,”
Sangho Lee and et al., · 2021
Cited alongside, same era.
“The efficacy of self-supervised speech models for audio representations,”
Tung-Yu Wu, Tsu-Yuan Hsu, Chen-An Li, Tzu-Han Lin, and Hung-yi Lee, · 2022
Cited alongside, same era.
“Pretext tasks selection for multitask self-supervised audio representation learning,”
Salah Zaiem, Titouan Parcollet, Slim Essid, and Abdelwahab Heba, · 2022
Cited alongside, same era.
Chen Sanyuan and et al., · 2022
Later among the works it cites.
“data2vec: A general framework for self-supervised learning in speech, vision and language,”
Alexei Baevski and et al., · 2022
Later among the works it cites.
“XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,”
Arun Babu and et al., · 2022
Later among the works it cites.
“Efficient training of audio transformers with patchout,”
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer, · 2022
Later among the works it cites.
“Ssast: Self-supervised audio spectrogram transformer,”
Yuan Gong, Cheng-I Lai, Yu-An Chung, and James Glass, · 2022
Later among the works it cites.
“BYOL for Audio: Exploring pre-trained general-purpose audio representations,”
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino, · 2023
Later among the works it cites.
“AudioMNIST: Exploring explainable artificial intelligence for audio analysis on a simple benchmark,”
Sören Becker, Johanna Vielhaben, Marcel Ackermann, Klaus-Robert Müller, Sebastian Lapuschkin, and Wojciech Samek, · 2024
Closest in time.