Fetching the paper…
Reading the bibliography…
End-to-end neural network based approaches to audio modelling are generally outperformed by models trained on high-level data representations.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Learning a better representation of speech soundwaves using restricted boltzmann machines
N. Jaitly and G. Hinton · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
End-to-end learning for music audio
S. Dieleman and B. Schrauwen · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
J. Salamon, C. Jacoby, and J. P. Bello · 2014
Earlier work this paper cites.
Acoustic modeling with deep neural networks using raw time signal for lvcsr
Z. Tüske, P. Golik, R. Schlüter, and H. Ney · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Speech acoustic modeling from raw multichannel waveforms
Y. Hoshen, R. J. Weiss, and K. W. Wilson · 2015
Cited alongside, same era.
Environmental sound classification with convolutional neural networks
K. J. Piczak · 2015
Cited alongside, same era.
Esc: Dataset for environmental sound classification
K. J. Piczak · 2015
Cited alongside, same era.
Learning the speech front-end with raw waveform cldnns
T. N. Sainath, R. J. Weiss, A. Senior, K. W. Wilson, and O. Vinyals · 2015
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Later among the works it cites.
Wav2letter: an end-to-end convnet-based speech recognition system
R. Collobert, C. Puhrsch, and G. Synnaeve · 2016
Later among the works it cites.
Acoustic modelling from the signal domain using cnns
P. Ghahremani, V. Manohar, D. Povey, and S. Khudanpur · 2016
Later among the works it cites.
Wavenet: A generative model for raw audio
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Later among the works it cites.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
librosa 0.5.0, 2017
B. McFee, M. McVicar, O. Nieto, S. Balke, C. Thome, D. Liang, E. Battenberg, J. Moore, R. Bittner, R. Yamamoto, D. Ellis, F.-R. Stoter, D. Repetto, S. Waloschek, C. Carr, S. Kranzler, K. Choi, P. Viktorin, J. F. Santos, A. Holovaty, W. Pimenta, and H. Lee · 2017
Closest in time.