Fetching the paper…
Reading the bibliography…
Learning acoustic models directly from the raw waveform data with minimal processing is challenging.
“Understanding the difficulty of training deep feedforward neural networks.,”
Xavier Glorot and Yoshua Bengio, · 2010
Earlier work this paper cites.
“Imagenet classification with deep convolutional neural networks,”
Alex Krizhevsky, Ilya Sutskever, and Geoff Hinton, · 2012
Earlier work this paper cites.
“On rectified linear units for speech processing,”
Matthew D et al. Zeiler, · 2013
Earlier work this paper cites.
“Deep speech: Scaling up end-to-end speech recognition,”
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al., · 2014
Earlier work this paper cites.
“Deepface: Closing the gap to human-level performance in face verification,”
Yaniv Taigman, Ming Yang, Marc’Aurelio Ranzato, and Lior Wolf, · 2014
Earlier work this paper cites.
“Acoustic modeling with deep neural networks using raw time signal for lvcsr.,”
Zoltán Tüske, Pavel Golik, Ralf Schlüter, and Hermann Ney, · 2014
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research,”
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello, · 2014
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Speech acoustic modeling from raw multichannel waveforms,”
Yedid Hoshen, Ron J Weiss, and Kevin W Wilson, · 2015
Cited alongside, same era.
“Learning the speech front-end with raw waveform cldnns,”
Tara N Sainath, Ron J Weiss, Andrew Senior, Kevin W Wilson, and Oriol Vinyals, · 2015
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2015
Cited alongside, same era.
“Environmental sound classification with convolutional neural networks,”
Karol J Piczak, · 2015
Later among the works it cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Later among the works it cites.
“Going deeper with convolutions,”
Christian et al Szegedy, · 2015
Later among the works it cites.
“Fully convolutional networks for semantic segmentation,”
Jonathan Long, Evan Shelhamer, and Trevor Darrell, · 2015
Later among the works it cites.
“CP-JKU submissions for DCASE-2016: a hybrid approach using binaural i-vectors and deep convolutional neural networks,”
Hamid Eghbal-Zadeh, Bernhard Lehner, Matthias Dorfer, and Gerhard Widmer, · 2016
Closest in time.
“Very deep multilingual convolutional neural networks for lvcsr,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Facenet: A unified embedding for face recognition and clustering,”
Florian Schroff, Dmitry Kalenichenko, and James Philbin, · 2015
Cited alongside, same era.
“Convolutional neural networks for acoustic modeling of raw time signal in lvcsr,”
Pavel Golik, Zoltán Tüske, Ralf Schlüter, and Hermann Ney, · 2015
Cited alongside, same era.
Tom Sercu, Christian Puhrsch, Brian Kingsbury, and Yann LeCun, · 2016
Closest in time.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,”
Martın Abadi, Ashish Agarwal, et al., · 2016
Closest in time.