Fetching the paper…
Reading the bibliography…
We train a bank of complex filters that operates on the raw waveform and is fed into a convolutional neural network for end-to-end phone recognition.
“Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,”
Steven Davis and Paul Mermelstein, · 1980
Earlier work this paper cites.
“Parametric coding of speech spectra,”
JL Flanagan, · 1980
Earlier work this paper cites.
“An efficient auditory filterbank based on the gammatone function,”
RD Patterson, Ian Nimmo-Smith, John Holdsworth, and Peter Rice, · 1987
Earlier work this paper cites.
“Timit acoustic-phonetic continuous speech corpus,”
John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, David S Pallett, Nancy L Dahlgren, and Victor Zue, · 1993
Earlier work this paper cites.
“Gradient-based learning applied to document recognition,”
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner, · 1998
Earlier work this paper cites.
“Efficient auditory coding,”
Evan C Smith and Michael S Lewicki, · 2006
Earlier work this paper cites.
“Imagenet classification with deep convolutional neural networks,”
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, · 2012
Earlier work this paper cites.
Dimitri Palaz, Ronan Collobert, and Mathew Magimai Doss, · 2013
Earlier work this paper cites.
Min Lin, Qiang Chen, and Shuicheng Yan, · 2013
Cited alongside, same era.
“End-to-end phoneme sequence recognition using convolutional neural networks,”
Dimitri Palaz, Ronan Collobert, and Mathew Magimai Doss, · 2013
Cited alongside, same era.
“Deep scattering spectrum,”
Joakim Andén and Stéphane Mallat, · 2014
Cited alongside, same era.
“Deep scattering spectrum with deep neural networks,”
Vijayaditya Peddinti, TaraN Sainath, Shay Maymon, Bhuvana Ramabhadran, David Nahamoo, and Vaibhava Goel, · 2014
Cited alongside, same era.
“Dropout: a simple way to prevent neural networks from overfitting.,”
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Cited alongside, same era.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2015
Later among the works it cites.
“Achieving human parity in conversational speech recognition,”
Wayne Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Mike Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig, · 2016
Later among the works it cites.
“A deep scattering spectrum—deep siamese network pipeline for unsupervised acoustic modeling,”
Neil Zeghidour, Gabriel Synnaeve, Maarten Versteegh, and Emmanuel Dupoux, · 2016
Later among the works it cites.
“Wavenet: A generative model for raw audio,”
Aaron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Later among the works it cites.
“Segmental recurrent neural networks for end-to-end speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yedid Hoshen, Ron J Weiss, and Kevin W Wilson, · 2015
Cited alongside, same era.
“Learning the speech front-end with raw waveform cldnns,”
Tara N Sainath, Ron J Weiss, Andrew Senior, Kevin W Wilson, and Oriol Vinyals, · 2015
Cited alongside, same era.
“Phone recognition with hierarchical convolutional deep maxout networks,”
László Tóth, · 2015
Cited alongside, same era.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio,
Cited in the paper.
Liang Lu, Lingpeng Kong, Chris Dyer, Noah A Smith, and Steve Renals, · 2016
Later among the works it cites.
“Wav2letter: an end-to-end convnet-based speech recognition system,”
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve, · 2016
Later among the works it cites.
“Attention-based wav2text with feature transfer learning,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Closest in time.
“Towards end-to-end speech recognition with deep convolutional neural networks,”
Ying Zhang, Mohammad Pezeshki, Philémon Brakel, Saizheng Zhang, Cesar Laurent Yoshua Bengio, and Aaron Courville, · 2017
Closest in time.