Fetching the paper…
Reading the bibliography…
The audio spectrogram is a time-frequency representation that has been widely used for audio classification.
SpecAugment: A simple data augmentation method for automatic speech recognition
Park, D. S.; Chan, W.; Zhang, Y.; Chiu, C.-C.; Zoph, B.; Cubuk, E. D.; and Le, Q. V. 2019 · 1904
Earlier work this paper cites.
A scale for the measurement of the psychological magnitude pitch
Stevens, S. S.; Volkmann, J.; and Newman, E. B. 1937 · 1937
Earlier work this paper cites.
Theory of communication. Part 1: The analysis of information
Gabor, D. 1946 · 1946
Earlier work this paper cites.
The need for biases in learning generalizations
Mitchell, T. M. 1980 · 1980
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
Fukushima, K.; and Miyake, S. 1982 · 1982
Earlier work this paper cites.
A Handbook of Fourier Theorems
Champeney, D. C.; and Champeney, D. 1987 · 1987
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
LeCun, Y.; Boser, B.; Denker, J.; Henderson, D.; Howard, R.; Hubbard, W.; and Jackel, L. 1989 · 1989
Earlier work this paper cites.
The information bottleneck method
Tishby, N.; Pereira, F. C.; and Bialek, W. 2000 · 2000
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E. 2001 · 2001
Earlier work this paper cites.
A tutorial on the cross-entropy method
De Boer, P.-T.; Kroese, D. P.; Mannor, S.; and Rubinstein, R. Y. 2005 · 2005
Earlier work this paper cites.
Online real-time onset detection with recurrent neural networks
Böck, S.; Arzt, A.; Krebs, F.; and Schedl, M. 2012 · 2012
Earlier work this paper cites.
Speaker identification using spectrograms of varying frame sizes
Kekre, H.; Kulkarni, V.; Gaikar, P.; and Gupta, N. 2012 · 2012
Earlier work this paper cites.
Learning filter banks within a deep neural network framework
Sainath, T. N.; Kingsbury, B.; Mohamed, A.; and Ramabhadran, B. 2013 · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
A Dictionary Of Physics
Law, J.; and Rennie, R. 2015 · 2015
Earlier work this paper cites.
Deep Learning
LeCun, Y.; Bengio, Y.; and Hinton, G. 2015 · 2015
Earlier work this paper cites.
Machine Learning with Spark
Pentreath, N. 2015 · 2015
Earlier work this paper cites.
ESC: Dataset for environmental sound classification
Piczak, K. J. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Neural audio synthesis of musical notes with WaveNet autoencoders
Engel, J.; Resnick, C.; Roberts, A.; Dieleman, S.; Norouzi, M.; Eck, D.; and Simonyan, K. 2017 · 2017
Cited alongside, same era.
AudioSet: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Cited alongside, same era.
Comparison of time-frequency representations for environmental sound classification using convolutional neural networks
Huzaifah, M. 2017 · 2017
Cited alongside, same era.
Raw waveform-based audio classification using sample-level CNN architectures
Lee, J.; Kim, T.; Park, J.; and Nam, J. 2017 · 2017
Cited alongside, same era.
Implicit regularization in deep learning
Neyshabur, B. 2017 · 2017
Cited alongside, same era.
Channel-Wise Subband Input for Better Voice and Accompaniment Separation on High Resolution Music
Liu, H.; Xie, L.; Wu, J.; and Yang, G. 2020 · 2020
Later among the works it cites.
Norm-preservation: Why residual networks can become extremely deep?
Zaeemzadeh, A.; Rahnavard, N.; and Shah, M. 2020 · 2020
Later among the works it cites.
Codified audio language modeling learns useful representations for music information retrieval
Castellon, R.; Donahue, C.; and Liang, P. 2021 · 2021
Later among the works it cites.
How low can you go? Reducing frequency and time resolution in current CNN architectures for music auto-tagging
Ferraro, A.; Bogdanov, D.; Jay, X. S.; Jeon, H.; and Yoon, J. 2021 · 2021
Later among the works it cites.
FSD50K: An open dataset of human-labeled sound events
Fonseca, E.; Favory, X.; Pons, J.; Font, F.; and Serra, X. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Opening the black box of deep neural networks via information
Shwartz-Ziv, R.; and Tishby, N. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Trainable frontend for robust and far-field keyword spotting
Wang, Y.; Getreuer, P.; Hughes, T.; Lyon, R. F.; and Saurous, R. A. 2017 · 2017
Cited alongside, same era.
MixUp: Beyond empirical risk minimization
Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2017 · 2017
Cited alongside, same era.
Speech commands: A dataset for limited-vocabulary speech recognition
Warden, P. 2018 · 2018
Cited alongside, same era.
Learning Filterbanks from Raw Speech for Phone Recognition
Zeghidour, N.; Usunier, N.; Kokkinos, I.; Schatz, T.; Synnaeve, G.; and Dupoux, E. 2018 · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Arora, S.; Cohen, N.; Hu, W.; and Luo, Y. 2019 · 2019
Cited alongside, same era.
Slow-fast auditory streams for audio recognition
Kazakos, E.; Nagrani, A.; Zisserman, A.; and Damen, D. 2021 · 2021
Later among the works it cites.
Broadcasted residual learning for efficient keyword spotting
Kim, B.; Chang, S.; Lee, J.; and Sung, D. 2021 · 2021
Later among the works it cites.
Efficient training of audio transformers with Patchout
Koutini, K.; Schlüter, J.; Eghbal-zadeh, H.; and Widmer, G. 2021 · 2021
Later among the works it cites.
Learning Strides in Convolutional Neural Networks
Riad, R.; Teboul, O.; Grangier, D.; and Zeghidour, N. 2021 · 2021
Later among the works it cites.
Shu, X.; Zhu, Y.; Chen, Y.; Chen, L.; Liu, H.; Huang, C.; and Wang, Y. 2021 · 2021
Later among the works it cites.
LEAF: A learnable frontend for audio classification
Zeghidour, N.; Teboul, O.; Quitry, F. d. C.; and Tagliasacchi, M. 2021 · 2021
Later among the works it cites.
Learning fast sample re-weighting without reward data
Zhang, Z.; and Pfister, T. 2021 · 2021
Later among the works it cites.
HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection
Chen, K.; Du, X.; Zhu, B.; Ma, Z.; Berg-Kirkpatrick, T.; and Dubnov, S. 2022 · 2022
Closest in time.
Gazneli, A.; Zimerman, G.; Ridnik, T.; Sharir, G.; and Noy, A. 2022 · 2022
Closest in time.
SSAST: Self-supervised audio spectrogram transformer
Gong, Y.; Lai, C.-I.; Chung, Y.-A.; and Glass, J. 2022 · 2022
Closest in time.
Segment-level Metric Learning for Few-shot Bioacoustic Event Detection
Liu, H.; Liu, X.; Mei, X.; Kong, Q.; Wang, W.; and Plumbley, M. D. 2022 · 2022
Closest in time.
Real time spectrogram inversion on mobile phone
Rybakov, O.; Tagliasacchi, M.; Li, Y.; Jiang, L.; Zhang, X.; and Biadsy, F. 2022 · 2022
Closest in time.
Simple Pooling Front-ends For Efficient Audio Classification
Liu, X.; Liu, H.; Kong, Q.; Mei, X.; Plumbley, M. D.; and Wang, W. 2023 · 2023
Closest in time.