Fetching the paper…
Reading the bibliography…
In this paper we propose a novel data augmentation method for attention-based end-to-end automatic speech recognition (E2E-ASR), utilizing a large amount of text which is not paired with speech signals.
“Continuous speech recognition by statistical methods,”
Frederick Jelinek, · 1976
Earlier work this paper cites.
“Backpropagation through time: what it does and how to do it,”
Paul J Werbos, · 1990
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“ADADELTA: an adaptive learning rate method,”
Matthew D Zeiler, · 2012
Earlier work this paper cites.
“End-to-end continuous speech recognition using attention-based recurrent NN: First results,”
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Learning phrase representations using RNN encoder-decoder for statistical machine translation,”
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Earlier work this paper cites.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling,”
Haşim Sak, Andrew Senior, and Françoise Beaufays, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Cited alongside, same era.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Cited alongside, same era.
“On using monolingual corpora in neural machine translation,”
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2015
Cited alongside, same era.
“Improving neural machine translation models with monolingual data,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2015
“Zoneout: Regularizing rnns by randomly preserving hidden activations,”
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron Courville, and Chris Pal, · 2016
Later among the works it cites.
“Improving end-to-end models for speech recognition,” https://ai.googleblog.com/2017/12/improving-end-to-end-models-for-speech.html
Tara Sainath and Yonghui Wu, · 2017
Later among the works it cites.
“Cold fusion: Training seq2seq models together with language models,”
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates, · 2017
Later among the works it cites.
Takaaki Hori, Shinji Watanabe, Yu Zhang, and William Chan, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Librispeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Effective approaches to attention-based neural machine translation,”
Minh-Thang Luong, Hieu Pham, and Christopher D Manning, · 2015
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al., · 2016
Cited alongside, same era.
“Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition,”
Hagen Soltau, Hank Liao, and Hasim Sak, · 2016
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2016
Cited alongside, same era.
Marcin Junczys-Dowmunt and Roman Grundkiewicz, · 2016
Cited alongside, same era.
Guillaume Lample, Ludovic Denoyer, and Marc’Aurelio Ranzato, · 2017
Later among the works it cites.
“Listening while speaking: Speech chain by deep learning,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Later among the works it cites.
“Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, et al., · 2017
Later among the works it cites.
“Multi-modal data augmentation for end-to-end asr,”
Adithya Renduchintala, Shuoyang Ding, Matthew Wiesner, and Shinji Watanabe, · 2018
Closest in time.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron J Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, et al., · 2018
Closest in time.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al., · 2018
Closest in time.