Fetching the paper…
Reading the bibliography…
Audio classification is an active research area with a wide range of applications.
J. P. Woodard, “Modeling and classification of natural sounds by product code hidden markov models,” IEEE Transactions on Signal Processing , vol. 40, no. 7, pp. 1833–1835, 1992
1992
Earlier work this paper cites.
R. S. Goldhor, “Recognition of environmental sounds,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , 1993
1993
Earlier work this paper cites.
Y. LeCun, Y. Bengio et al. , “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, 1995
1995
Earlier work this paper cites.
S. Chachada and C.-C. J. Kuo, “Environmental sound recognition: A survey,” APSIPA Transactions on Signal and Information Processing , vol. 3, 2014
2014
Earlier work this paper cites.
G. Gwardys and D. M. Grzywczak, “Deep image features in music information retrieval,” International Journal of Electronics and Telecommunications , vol. 60, no. 4, pp. 321–326, 2014
2014
Earlier work this paper cites.
K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in IEEE International Workshop on Machine Learning for Signal Processing , 2015
2015
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in NIPS Deep Learning and Representation Learning Workshop , 2015
2015
Earlier work this paper cites.
K. J. Piczak, “Esc: Dataset for environmental sound classification,” in ACM International Conference on Multimedia , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2016
2016
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “Cnn architectures for large-scale audio classification,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , 2017
2017
Earlier work this paper cites.
J. Salamon and J. P. Bello, “Deep convolutional neural networks and data augmentation for environmental sound classification,” IEEE Signal processing letters , vol. 24, no. 3, pp. 279–283, 2017
2017
Earlier work this paper cites.
Y. Tokozume and T. Harada, “Learning environmental sounds with end-to-end convolutional neural network,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , 2017
2017
Earlier work this paper cites.
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein, “Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Mesaros, T. Heittola, E. Benetos, P. Foster, M. Lagrange, T. Virtanen, and M. D. Plumbley, “Detection and classification of acoustic scenes and events: Outcome of the dcase 2016 challenge,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 2, pp. 379–393, 2017
2017
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson, “Averaging weights leads to wider optima and better generalization,” in Conference on Uncertainty in Artificial Intelligence , 2018
2018
Earlier work this paper cites.
K. Koutini, H. Eghbal-zadeh, and G. Widmer, “Iterative knowledge distillation in r-cnns for weakly-labeled semi-supervised sound event detection,” in DCASE Workshop , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Conference of the North American Chapter of the Association for Computational Linguistics , 2019
2019
Cited alongside, same era.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International Conference on Machine Learning , 2019
2019
Cited alongside, same era.
S. Adapa, “Urban sound tagging using convolutional neural networks,” in DCASE Workshop , 2019
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” Interspeech , 2019
2019
Cited alongside, same era.
M. Meire, L. Vuegen, and P. Karsmakers, “The impact of missing labels and overlapping sound events on multi-label multi-instance learning for sound event classification,” in DCASE Workshop , 2019
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning , 2021
2021
Later among the works it cites.
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy, “Do vision transformers see like convolutional neural networks?” Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Gong, Y.-A. Chung, and J. Glass, “Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
L. Ford, H. Tang, F. Grondin, and J. R. Glass, “A deep residual network for large-scale acoustic scene analysis.” in Interspeech , 2019
2019
Cited alongside, same era.
Y. Wang, J. Li, and F. Metze, “A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
E. Fonseca, S. Hershey, M. Plakal, D. P. Ellis, A. Jansen, and R. C. Moore, “Addressing missing labels in large-scale sound event recognition using a teacher-student framework with loss masking,” IEEE Signal Processing Letters , vol. 27, pp. 1235–1239, 2020
2020
Cited alongside, same era.
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le, “Self-training with noisy student improves imagenet classification,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020
2020
Cited alongside, same era.
A. Guzhov, F. Raue, J. Hees, and A. Dengel, “Esresnet: Environmental sound classification based on visual domain models,” in International Conference on Pattern Recognition , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
M. M. Naseer, K. Ranasinghe, S. H. Khan, M. Hayat, F. Shahbaz Khan, and M.-H. Yang, “Intriguing properties of vision transformers,” Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
T. Xiao, P. Dollar, M. Singh, E. Mintun, T. Darrell, and R. Girshick, “Early convolutions help transformers see better,” Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu, “Incorporating convolution designs into visual transformers,” in IEEE/CVF International Conference on Computer Vision , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Dai, H. Liu, Q. Le, and M. Tan, “Coatnet: Marrying convolution and attention for all data sizes,” Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.