Fetching the paper…
Reading the bibliography…
Over the past few years, audio classification task on large-scale dataset such as AudioSet has been an important research area.
Hochreiter, Sepp & Schmidhuber, Jürgen. (1997). Long Short-term Memory. Neural computation. 9. 1735-80. 10.1162/neco.1997.9.8.1735
1997
Earlier work this paper cites.
1997
Earlier work this paper cites.
Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urbansound research,” in 22nd ACM International Conference on Multimedia(ACM-MM’14), Orlando, FL, USA, Nov. 2014, pp. 1041–1044
2014
Earlier work this paper cites.
Bahdanau, Dzmitry & Cho, Kyunghyun & Bengio, Y.. (2014). Neural Machine Translation by Jointly Learning to Align and Translate. ArXiv. 1409
2014
Earlier work this paper cites.
Piczak, Karol. (2015). ESC: Dataset for Environmental Sound Classification. 1015-1018. 10.1145/2733373.2806390
2015
Earlier work this paper cites.
Luong, Minh-Thang & Pham, Hieu & Manning, Christopher. (2015). Effective Approaches to Attention-based Neural Machine Translation. 10.18653/v1/D15-1166
2015
Earlier work this paper cites.
K. Choi, G. Fazekas, and M. Sandler, “Automatic tagging using deepconvolutional neural networks,” in Conference of the InternationalSociety for Music Information Retrieval (ISMIR), 2016, pp. 805–811
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning forimage recognition,” in IEEE Conference on Computer Vision and PatternRecognition (CVPR), 2016, pp. 770–778
2016
Earlier work this paper cites.
K. Choi, G. Fazekas, and M. Sandler, “Automatic tagging using deepconvolutional neural networks,” inConference of the InternationalSociety for Music Information Retrieval (ISMIR), 2016, pp. 805–811
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C.Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference onAcoustics, Speech and Signal Processing (ICASSP), 2017, pp. 776–780
2017
Cited alongside, same era.
Gemmeke, Jort & Ellis, Daniel & Freedman, Dylan & Jansen, Aren & Lawrence, Wade & Moore, R. & Plakal, Manoj & Ritter, Marvin. (2017) Audio Set: An ontology and human-labeled dataset for audio events. 776-780. 10.1109/ICASSP.2017.7952261
2017
Cited alongside, same era.
W. Dai, C. Dai, S. Qu, J. Li, and S. Das, “Very deep convolutionalneural networks for raw waveforms,” in IEEE International Conferenceon Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 421–425
2017
Cited alongside, same era.
Bartz, Christian & Herold, Tom & Yang, Haojin & Meinel, Christoph. (2017). Language Identification Using Deep Convolutional Recurrent Neural Networks
2017
Cited alongside, same era.
Sarthak, & Shukla, Shikhar & Mittal, Govind. (2019). Spoken Language Identification using ConvNets
2019
Later among the works it cites.
Q. Kong, C. Yu, Y. Xu, T. Iqbal, W. Wang, and M. D. Plumbley, “Weaklylabelled audioset tagging with attention neural networks, ”IEEE/ACMTransactions on Audio, Speech, and Language Processing, vol. 27, pp.1791–1802, 2019
2019
Later among the works it cites.
Fonseca, Eduardo & Favory, Xavier & Pons, Jordi & Font, Frederic & Serra, Xavier. (2020). FSD50K: an Open Dataset of Human-Labeled Sound Events
2020
Later among the works it cites.
Prateek Verma and Jonathan Berger, Audio Transformers:Transformer Architectures For Large Scale Audio Understanding,2021, arXiv, cs.SD
2021
Later among the works it cites.
Koutini, Khaled & Schlüter, Jan & Eghbal-zadeh, Hamid & Widmer, Gerhard. (2021). Efficient Training of Audio Transformers with Patchout
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vaswani, Ashish & Shazeer, Noam & Parmar, Niki & Uszkoreit, Jakob & Jones, Llion & Gomez, Aidan & Kaiser, Lukasz & Polosukhin, Illia. (2017). Attention Is All You Need
2017
Cited alongside, same era.
A. Mesaros, T. Heittola, and T. Virtanen, “A multi-device dataset forurban acoustic scene classification,” in Workshop on Detection andClassification of Acoustic Scenes and Events (DCASE), 2018, pp. 9–13
2018
Cited alongside, same era.
Yoon, Seunghyun & Byun, Seokhyun & Jung, Kyomin. (2018). Multimodal Speech Emotion Recognition Using Audio and Text. 10.1109/SLT.2018.8639583
2018
Cited alongside, same era.
Sandler, Mark & Howard, Andrew & Zhu, Menglong & Zhmoginov, Andrey & Chen, Liang-Chieh. (2018). MobileNetV2: Inverted Residuals and Linear Bottlenecks. 4510-4520. 10.1109/CVPR.2018.00474
2018
Cited alongside, same era.
Kong, Qiuqiang Cao, Yin Iqbal, Turab Wang, Yuxuan Plumbley, Mark. (2019). PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
2019
Cited alongside, same era.
https://github.com/tensorflow/models/tree/master/research/audioset/vggish : Audioset based Vggish
Cited in the paper.
https://github.com/tensorflow/models/tree/master/research/audioset/YAMNet : AudioSet based YAMNet
Cited in the paper.
https://www.kaggle.com/c/freesound-audio-tagging
Cited in the paper.
2021
Later among the works it cites.
Gong, Yuan & Chung, Yu-An & Glass, James. (2021). PSLA: Improving Audio Tagging With Pretraining, Sampling, Labeling, and Aggregation. IEEE/ACM Transactions on Audio, Speech, and Language Processing
2021
Later among the works it cites.
Niizumi, Daisuke & Takeuchi, Daiki & Ohishi, Yasunori & Harada, Noboru & Kashino, Kunio. (2021). BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation. IJCNN52387.2021.9534474
2021
Later among the works it cites.
Wu, Ho-Hsiang & Seetharaman, Prem& Kumar, Kundan & Bello, Juan. (2022). Wav2CLIP: Learning Robust Audio Representations from Clip
2022
Later among the works it cites.