Fetching the paper…
Reading the bibliography…
Recently, SpecAugment, an augmentation scheme for automatic speech recognition that acts directly on the spectrogram of input utterances, has shown to be highly effective in enhancing the performance of end-to-end networks on public datasets.
“Japanese and korean voice search,”
Mike Schuster and Kaisuke Nakajima, · 2012
Earlier work this paper cites.
“Elastic spectral distortion for low resource speech recognition with deep neural networks,”
Naoyuki Kanda, Ryu Takeda, and Yasunari Obuchi, · 2013
Earlier work this paper cites.
“Vocal Tract Length Perturbation (VTLP) improves speech recognition,”
Navdeep Jaitly and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Data augmentation for low resource languages,”
Anton Ragni, Kate M. Knill, Shakti P. Rath, and Mark J. F. Gales, · 2014
Earlier work this paper cites.
“Deep Speech: Scaling up end-to-end speech recognition,”
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and Andrew Ng, · 2014
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Audio Augmentation for Speech Recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Automatic gain control and multi-style training for robust small-footprint keyword spotting with deep neural networks,”
R. Prabhavalkar, R. Alvarez, C. Parada, P. Nakkiran, and T. N. Sainath, · 2015
Earlier work this paper cites.
“Learning the SpeechFront-end With Raw Waveform CLDNNs,”
Tara Sainath, Ron Weiss, Kevin Wilson, Andrew Senior, and Oriol Vinyals, · 2015
Earlier work this paper cites.
“On using monolingual corpora in neural machine translation,”
Çaglar Gülçehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loïc Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“Novel neural network based fusion for Multistream ASR,”
Sri Harish Mallidi and Hynek Hermansky, · 2016
Cited alongside, same era.
“Layer normalization,”
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Cited alongside, same era.
“Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,”
Chanwoo Kim, Ananya Misra, Kean Chin, Thad Hughes, Arun Narayanan, Tara Sainath, and Michiel Bacchiani, · 2017
Cited alongside, same era.
“Increasing the robustness of cnn acoustic models using autoregressive moving average spectrogram features and channel dropout,”
György Kovács, László Tóth, Dirk Van Compernolle, and Sriram Ganapathy, · 2017
Cited alongside, same era.
“State-of-the-art Speech Recognition With Sequence-to-Sequence Models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani, · 2018
Later among the works it cites.
“Speech recognition for medical conversations,”
Chung-Cheng Chiu, Anshuman Tripathi, Katherine Chou, Chris Co, Navdeep Jaitly, Diana Jaunzeikare, Anjuli Kannan, Patrick Nguyen, Hasim Sak, Ananth Sankar, Justin Tansuwan, Nathan Wan, Yonghui Wu, and Xuedong Zhang, · 2018
Later among the works it cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Closest in time.
“Improved Vocal Tract Length Perturbation for a State-of-the-Art End-to-End Speech Recognition System,”
Chanwoo Kim, Minkyu Shin, Abhinav Garg, and Dhananjaya Gowda, · 2019
Closest in time.
“Examining the Combination of Multi-Band Processing and Channel Dropout for Robust Speech Recognition,”
György Kovács, László Tóth, Dirk Van Compernolle, and Marcus Liwicki, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Acoustic Modeling for Google Home,”
Bo Li, Tara Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean Chin, Khe Chai Sim, Ron Weiss, Kevin Wilson, Ehsan Variani, Chanwoo Kim, Olivier Siohan, Mitchel Weintraub, Erik McDermott, Rick Rose, and Matt Shannon, · 2017
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2017
Cited alongside, same era.
“Toward domain-invariant speech recognition via large scale training,”
Arun Narayanan, Ananya Misra, Khe Chai Sim, Golan Pundak, Anshuman Tripathi, Mohamed Elfeky, Parisa Haghani, Trevor Strohman, and Michiel Bacchiani, · 2018
Cited alongside, same era.
“Data Augmentation for Robust Keyword Spotting under Playback Interference,”
Anirudh Raju, Sankaran Panchapagesan, Xing Liu, Arindam Mandal, and Nikko Strom, · 2018
Cited alongside, same era.
“A perceptually inspired data augmentation method for noise robust cnn acoustic models,”
László Tóth, György Kovács, and Dirk Van Compernolle, · 2018
Cited alongside, same era.
Closest in time.
“RWTH ASR systems for librispeech: Hybrid vs attention - w/o data augmentation,”
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Closest in time.
“A comparative study on transformer vs rnn in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, Shinji Watanabe, Takenori Yoshimura, and Wangyou Zhang, · 2019
Closest in time.
“State-of-the-art speech recognition using multi-stream self-attention with dilated 1d convolutions,”
Kyu J. Han, Ramon Prieto, Kaixing Wu, and Tao Ma, · 2019
Closest in time.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, Qiao Liang, Deepti Bhatia, Yuan Shangguan, Bo Li, Golan Pundak, Khe Chai Sim, Tom Bagby, Shuo yiin Chang, Kanishka Rao, and Alexander Gruenstein, · 2019
Closest in time.