Fetching the paper…
Reading the bibliography…
Audio classification is an important task of mapping audio samples into their corresponding labels.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, and Richard Socher et al., · 2009
Earlier work this paper cites.
“Automatic classification of musical instrument sounds,”
Perfecto Herrera, Geoffroy Peeters, and Shlomo Dubnov, · 2010
Earlier work this paper cites.
“ESC: dataset for environmental sound classification,”
Karol J. Piczak, · 2015
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F. Gemmeke, Daniel P. W. Ellis, and Dylan Freedman et al., · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, and Niki Parmar et al., · 2017
Earlier work this paper cites.
“Speech commands: A dataset for limited-vocabulary speech recognition,”
Pete Warden, · 2018
Earlier work this paper cites.
“mixup: Beyond empirical risk minimization,”
Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz, · 2018
Earlier work this paper cites.
“A deep residual network for large-scale acoustic scene analysis,”
Logan Ford, Hao Tang, and François Grondin et al., · 2019
Earlier work this paper cites.
“A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling,”
Yun Wang, Juncheng Li, and Florian Metze, · 2019
Earlier work this paper cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park et al., · 2019
Earlier work this paper cites.
“Efficientnet: Rethinking model scaling for convolutional neural networks,”
Mingxing Tan and Quoc V. Le, · 2019
Cited alongside, same era.
“Music sketchnet: Controllable music generation via factorized representations of pitch and rhythm,”
Ke Chen, Cheng-i Wang, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2020
Cited alongside, same era.
“Muspy: A toolkit for symbolic music generation,”
Hao-Wen Dong, Ke Chen, Julian J. McAuley, and Taylor Berg-Kirkpatrick, · 2020
Cited alongside, same era.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, and Turab Iqbal et al., · 2020
Cited alongside, same era.
“Sound event detection in synthetic domestic environments,”
Romain Serizel, Nicolas Turpault, Ankit Parag Shah, and Justin Salamon, · 2020
Cited alongside, same era.
“Continuous melody generation via disentangled short-term representations and structural conditions,”
“Ts-cam: Token semantic coupled attention map for weakly supervised object localization,”
Wei Gao, Fang Wan, and Xingjia Pan et al., · 2021
Later among the works it cites.
“Swin transformer: Hierarchical vision transformer using shifted windows,”
Ze Liu, Yutong Lin, and Yue Cao et al., · 2021
Later among the works it cites.
“Learning efficient representations for keyword spotting with triplet loss,”
Roman Vygon and Nikolay Mikhaylovskiy, · 2021
Later among the works it cites.
“Eranns: Efficient residual audio neural networks for audio pattern recognition,”
Sergey Verbitskiy and Viacheslav Vyshegorodtsev, · 2021
Later among the works it cites.
“Keyword transformer: A self-attention model for keyword spotting,”
Axel Berg, Mark O’Connor, and Miguel Tairum Cruz, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ke Chen, Gus Xia, and Shlomo Dubnov, · 2020
Cited alongside, same era.
“Sound event detection: A tutorial,”
Annamaria Mesaros, Toni Heittola, Tuomas Virtanen, and Mark D. Plumbley, · 2021
Cited alongside, same era.
“Learning audio embeddings with user listening data for content-based music recommendation,”
Ke Chen, Beici Liang, Xiaoshuan Ma, and Minwei Gu, · 2021
Cited alongside, same era.
“Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation,”
Yuan Gong, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“Ast: Audio spectrogram transformer,”
Yuan Gong, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“Dcase 2021 challenge task 4: Sound event detection and separation in domestic environments,” http://dcase.community/challenge2021
2021
Later among the works it cites.
“Training data-efficient image transformers & distillation through attention,”
Hugo Touvron, Matthieu Cord, and Matthijs Douze et al., · 2021
Later among the works it cites.
“The benefit of temporally-strong labels in audio event classification,”
Shawn Hershey, Daniel P. W. Ellis, and Eduardo Fonseca et al., · 2021
Later among the works it cites.
“Zero-shot audio source separation through query-based learning from weakly-labeled data,”
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2022
Closest in time.