Fetching the paper…
Reading the bibliography…
Transformer-based models attain excellent results and generalize well when trained on sufficient amounts of data.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur, “Recurrent neural network based language model.” in Interspeech , 2010, pp. 1045–1048
2010
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
K. J. Piczak, “Esc: Dataset for environmental sound classification,” in Proceedings of the 23rd ACM international conference on Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
S. Albawi, T. A. Mohammed, and S. Al-Zawi, “Understanding of a convolutional neural network,” in International Conference on Engineering and Technology (ICET) . IEEE, 2017, pp. 1–6
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
E. Humphrey, S. Durand, and B. McFee, “Openmic-2018: An open data-set for multiple instrument recognition.” in ISMIR , 2018, pp. 438–444
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. Le, and R. Salakhutdinov, “Transformer-xl: Attentive language models beyond a fixed-length context,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 2978–2988
2019
Cited alongside, same era.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang et al. , “A comparative study on transformer vs rnn in speech applications,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 449–456
2019
Cited alongside, same era.
Y. Wang, J. Li, and F. Metze, “A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 31–35
H. Wang, Y. Zou, D. Chong, and W. Wang, “Environmental sound classification with parallel temporal-spectral attention,” Interspeech , pp. 821–825, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
S. d’Ascoli, H. Touvron, M. L. Leavitt, A. S. Morcos, G. Biroli, and L. Sagun, “Convit: Improving vision transformers with soft convolutional inductive biases,” in International Conference on Machine Learning . PMLR, 2021, pp. 2286–2296
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
L. Ford, H. Tang, F. Grondin, and J. R. Glass, “A deep residual network for large-scale acoustic scene analysis.” in Interspeech , 2019, pp. 2568–2572
2019
Cited alongside, same era.
A. Mesaros, T. Heittola, and T. Virtanen, “Acoustic scene classification in dcase 2019 challenge: Closed and open set classification and data mismatch setups,” 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” Interspeech , pp. 5036–5040, 2020
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
H. Wang, Y. Zou, D. Chong, and W. Wang, “Modeling label dependencies for audio tagging with graph convolutional network,” IEEE Signal Processing Letters , vol. 27, pp. 1560–1564, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Sgdr: Stochastic gradient descent with warm restarts.”
Cited in the paper.
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Gong, Y.-A. Chung, and J. Glass, “Psla: Improving audio event classification with pretraining, sampling, labeling, and aggregation,” arXiv e-prints , pp. arXiv–2102, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.