Fetching the paper…
Reading the bibliography…
While efficient architectures and a plethora of augmentations for end-to-end image classification tasks have been suggested and heavily investigated, state-of-the-art techniques for audio classifications still rely on numerous representations of the audio signal together with large architectures, fine-tuned from large datasets.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello · 2014
Earlier work this paper cites.
Densenet: Implementing efficient convnet descriptor pyramids
Forrest Iandola, Matt Moskewicz, Sergey Karayev, Ross Girshick, Trevor Darrell, and Kurt Keutzer · 2014
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
End-to-end learning for music audio tagging at scale
Jordi Pons, Oriol Nieto, Matthew Prockup, Erik Schmidt, Andreas Ehmann, and Xavier Serra · 2017
Earlier work this paper cites.
Learning from between-class examples for deep sound recognition
Yuji Tokozume, Yoshitaka Ushiku, and Tatsuya Harada · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms
Jongpil Lee, Jiyoung Park, Keunhyoung Luke Kim, and Juhan Nam · 2017
Earlier work this paper cites.
Muhammad Huzaifah · 2017
Earlier work this paper cites.
Utilizing domain knowledge in end-to-end audio processing
Tycho Max Sylvester Tax, Jose Luis Diez Antich, Hendrik Purwins, and Lars Maaløe · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Cited alongside, same era.
Speech commands: A dataset for limited-vocabulary speech recognition
Pete Warden · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Cited alongside, same era.
Leslie N Smith · 2018
Cited alongside, same era.
Rethinking cnn models for audio classification
Kamalesh Palanisamy, Dipika Singhania, and Angela Yao · 2020
Later among the works it cites.
Efficient end-to-end audio embeddings generation for audio classification on target applications
Paulo Lopez-Meyer, Juan A del Hoyo Ontiveros, Hong Lu, and Georg Stemmer · 2021
Later among the works it cites.
Ast: Audio spectrogram transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Later among the works it cites.
Efficient training of audio transformers with patchout
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer · 2021
Later among the works it cites.
Eranns: Efficient residual audio neural networks for audio pattern recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Cited alongside, same era.
Deep learning for audio signal processing
Hendrik Purwins, Bo Li, Tuomas Virtanen, Jan Schlüter, Shuo-Yiin Chang, and Tara Sainath · 2019
Cited alongside, same era.
Making convolutional networks shift-invariant again
Richard Zhang · 2019
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo · 2019
Cited alongside, same era.
Sergey Verbitskiy, Vladimir Berikov, and Viacheslav Vyshegorodtsev · 2021
Later among the works it cites.
Esresnet: Environmental sound classification based on visual domain models
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2021
Later among the works it cites.
Esresne(x)t-fbsp: Learning robust time-frequency transformation of audio, 2021
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2021
Later among the works it cites.
Multi-format contrastive learning of audio representations
Luyu Wang and Aaron van den Oord · 2021
Later among the works it cites.
Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation
Yuan Gong, Yu-An Chung, and James Glass · 2021
Later among the works it cites.
Tresnet: High performance gpu-dedicated architecture
Tal Ridnik, Hussam Lawen, Asaf Noy, Emanuel Ben Baruch, Gilad Sharir, and Itamar Friedman · 2021
Later among the works it cites.
Axial residual networks for cyclegan-based voice conversion
Jaeseong You, Gyuhyeon Nam, Dalhyun Kim, and Gyeongsu Chae · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Ross Wightman, Hugo Touvron, and Hervé Jégou · 2021
Later among the works it cites.
Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2022
Closest in time.