Fetching the paper…
Reading the bibliography…
In the past, the rapidly evolving field of sound classification greatly benefited from the application of methods from other domains.
Nesterov, Y.: A method of solving a convex programming problem with convergence rate o (1/kˆ 2) o (1/k2). In: Sov. Math. Dokl. vol. 27 (1983)
1983
Earlier work this paper cites.
Polyak, B.T., Juditsky, A.B.: Acceleration of stochastic approximation by averaging. SIAM journal on control and optimization 30
1992
Earlier work this paper cites.
Teolis, A., Benedetto, J.J.: Computational signal processing with wavelets, vol. 182. Springer (1998)
1998
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition. pp. 248–255. Ieee (2009)
2009
Earlier work this paper cites.
Salamon, J., Jacoby, C., Bello, J.P.: A dataset and taxonomy for urban sound research. In: Proceedings of the 22nd ACM international conference on Multimedia. pp. 1041–1044 (2014)
2014
Earlier work this paper cites.
Piczak, K.J.: Environmental sound classification with convolutional neural networks. In: 2015 IEEE 25th International Workshop on Machine Learning for Signal Processing (MLSP). pp. 1–6. IEEE (2015)
2015
Earlier work this paper cites.
Piczak, K.J.: Esc: Dataset for environmental sound classification. In: Proceedings of the 23rd ACM international conference on Multimedia. pp. 1015–1018 (2015)
2015
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016)
2016
Earlier work this paper cites.
Chollet, F.: Xception: Deep learning with depthwise separable convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1251–1258 (2017)
2017
Earlier work this paper cites.
Gemmeke, J.F., Ellis, D.P., Freedman, D., Jansen, A., Lawrence, W., Moore, R.C., Plakal, M., Ritter, M.: Audio set: An ontology and human-labeled dataset for audio events. In: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 776–780. IEEE (2017)
2017
Earlier work this paper cites.
Hussain, Z., Gimenez, F., Yi, D., Rubin, D.: Differential data augmentation techniques for medical imaging classification tasks. In: AMIA Annual Symposium Proceedings. vol. 2017, p. 979. American Medical Informatics Association (2017)
2017
Earlier work this paper cites.
Sailor, H.B., Agrawal, D.M., Patil, H.A.: Unsupervised filterbank learning using convolutional restricted boltzmann machine for environmental sound classification. In: INTERSPEECH. pp. 3107–3111 (2017)
2017
Earlier work this paper cites.
Salamon, J., Bello, J.P.: Deep convolutional neural networks and data augmentation for environmental sound classification. IEEE Signal Processing Letters 24
2017
Cited alongside, same era.
Tokozume, Y., Harada, T.: Learning environmental sounds with end-to-end convolutional neural network. In: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 2721–2725 (March 2017). https://doi.org/10.1109/ICASSP.2017.7952651
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Palanisamy, K., Singhania, D., Yao, A.: Rethinking cnn models for audio classification (2020)
2020
Later among the works it cites.
2021
Closest in time.
Dzabraev, M., Kalashnikov, M., Komkov, S., Petiushko, A.: Mdmmt: Multidomain multimodal transformer for video retrieval. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3354–3363 (2021)
2021
Closest in time.
Gong, Y., Chung, Y.A., Glass, J.: Ast: Audio spectrogram transformer (2021)
2021
Closest in time.
Guzhov, A., Raue, F., Hees, J., Dengel, A.: Esresnet: Environmental sound classification based on visual domain models. In: 25th International Conference on Pattern Recognition (ICPR). pp. 4933–4940 (January 2021)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Islam, M.T., Nirjon, S.: Soundsemantics: exploiting semantic knowledge in text for embedded acoustic event classification. In: Proceedings of the 18th International Conference on Information Processing in Sensor Networks. pp. 217–228 (2019)
2019
Cited alongside, same era.
Kim, C.D., Kim, B., Lee, H., Kim, G.: Audiocaps: Generating captions for audios in the wild. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). pp. 119–132 (2019)
2019
Cited alongside, same era.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners. OpenAI blog 1
2019
Cited alongside, same era.
Xie, H., Virtanen, T.: Zero-shot audio classification based on class label embeddings. In: 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). pp. 264–267. IEEE (2019)
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Kumar, A., Ithapu, V.: A sequential self teaching approach for improving generalization in sound event recognition. In: International Conference on Machine Learning. pp. 5447–5457. PMLR (2020)
2020
Cited alongside, same era.
2021
Closest in time.
Guzhov, A., Raue, F., Hees, J., Dengel, A.: Esresne(x)t-fbsp: Learning robust time-frequency transformation of audio. In: 2021 International Joint Conference on Neural Networks (IJCNN) (2021)
2021
Closest in time.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision (2021)
2021
Closest in time.
Verbitskiy, S., Vyshegorodtsev, V.: Eranns: Efficient residual audio neural networks for audio pattern recognition (2021)
2021
Closest in time.
2021
Closest in time.
Xie, H., Virtanen, T.: Zero-shot audio classification via semantic embeddings. IEEE/ACM Transactions on Audio, Speech, and Language Processing 29
2021
Closest in time.