Fetching the paper…
Reading the bibliography…
Recent advances in neural network based acoustic modelling have shown significant improvements in automatic speech recognition (ASR) performance.
“DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,”
John S Garofolo, Lori F Lamel, William M Fisher, Jonathon G Fiscus, and David S Pallett, · 1993
Earlier work this paper cites.
“Vocal tract length normalization for large vocabulary continuous speech recognition,”
Puming Zhan and Alex Waibel, · 1997
Earlier work this paper cites.
“Accent issues in large vocabulary continuous speech recognition,”
Chao Huang, Tao Chen, and Eric Chang, · 2004
Earlier work this paper cites.
“Automatic speech recognition and speech variability: A review,”
Mohamed Benzeghiba, Renato De Mori, Olivier Deroo, Stephane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, et al., · 2007
Earlier work this paper cites.
“The application of hidden Markov models in speech recognition,”
Mark Gales and Steve Young, · 2008
Earlier work this paper cites.
“MLLR/MAP adaptation using pronunciation variation for non-native speech recognition,”
Yoo Rhee Oh and Hong Kook Kim, · 2009
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks.,”
Xavier Glorot and Yoshua Bengio, · 2010
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Voxforge,”
Ken MacLean, · 2012
Cited alongside, same era.
“Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks.,”
Dimitri Palaz, Ronan Collobert, and Mathew Magimai-Doss, · 2013
Cited alongside, same era.
“KL-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition,”
Dong Yu, Kaisheng Yao, Hang Su, Gang Li, and Frank Seide, · 2013
Cited alongside, same era.
“On the importance of initialization and momentum in deep learning,”
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton, · 2013
Cited alongside, same era.
“I-vector-based speaker adaptation of deep neural networks for french broadcast audio transcription,”
Vishwa Gupta, Patrick Kenny, Pierre Ouellet, and Themos Stafylakis, · 2014
Cited alongside, same era.
“Learning the speech front-end with raw waveform CLDNNs,”
Tara N Sainath, Ron J Weiss, Andrew Senior, Kevin W Wilson, and Oriol Vinyals, · 2015
Later among the works it cites.
“Domain-adversarial training of neural networks,”
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky, · 2016
Later among the works it cites.
“Acoustic modelling from the signal domain using CNNs.,”
Pegah Ghahremani, Vimal Manohar, Daniel Povey, and Sanjeev Khudanpur, · 2016
Later among the works it cites.
“Adversarial Multi-Task learning of deep neural networks for robust speech recognition.,”
Yusuke Shinohara, · 2016
Later among the works it cites.
“Invariant representations for noisy speech recognition,”
Dmitriy Serdyuk, Kartik Audhkhasi, Philémon Brakel, Bhuvana Ramabhadran, Samuel Thomas, and Yoshua Bengio, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pawel Swietojanski and Steve Renals, · 2014
Cited alongside, same era.
“Analysis of CNN-based speech recognition system using raw speech as input,”
Dimitri Palaz, Ronan Collobert, et al., · 2015
Cited alongside, same era.
“Convolutional neural networks-based continuous speech recognition using raw speech signal,”
Dimitri Palaz, Mathew Magimai Doss, and Ronan Collobert, · 2015
Cited alongside, same era.
“An unsupervised deep domain adaptation approach for robust speech recognition,”
Sining Sun, Binbin Zhang, Lei Xie, and Yanning Zhang, · 2017
Later among the works it cites.
Wei-Ning Hsu, Yu Zhang, and James Glass, · 2017
Later among the works it cites.
“Effects of Talker Dialect, Gender & Race on Accuracy of Bing Speech and YouTube Automatic Captions,”
Rachael Tatman and Conner Kasten, · 2017
Later among the works it cites.