Fetching the paper…
Reading the bibliography…
Previous work has shown that neural encoder-decoder speech recognition can be improved with hierarchical multitask learning, where auxiliary tasks are added at intermediate layers of a deep encoder.
“Switchboard: Telephone speech corpus for research and development,”
John J Godfrey, Edward C Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“Multitask learning,”
Rich Caruana, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Resegmentation of switchboard,”
Neeraj Deshmukh, Aravind Ganapathiraju, Andi Gleeson, Jonathan Hamaker, and Joseph Picone, · 1998
Earlier work this paper cites.
“2000 HUB5 English Evaluation Speech LDC2002S09,”
LD Consortium et al., · 2002
Earlier work this paper cites.
“Bidirectional LSTM networks for improved phoneme classification and recognition,”
Alex Graves, Santiago Fernández, and Jürgen Schmidhuber, · 2005
Earlier work this paper cites.
“Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence Labelling in Structured Domains with Hierarchical Recurrent Neural Networks.,”
Santiago Fernández, Alex Graves, and Jürgen Schmidhuber, · 2007
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Learning phrase representations using rnn encoder-decoder for statistical machine translation,”
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Visualizing and understanding convolutional networks,”
Matthew D Zeiler and Rob Fergus, · 2014
Earlier work this paper cites.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Cited alongside, same era.
“Recurrent Neural Network Regularization,”
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals, · 2014
Cited alongside, same era.
“Learning and transferring mid-level image representations using convolutional neural networks,”
Maxime Oquab, Leon Bottou, Ivan Laptev, and Josef Sivic, · 2014
Cited alongside, same era.
“Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition,”
Haşim Sak, Andrew Senior, Kanishka Rao, and Françoise Beaufays, · 2015
Cited alongside, same era.
“Adam: A Method for Stochastic Optimization,”
Diederik P Kingma and Jimmy Ba, · 2015
Cited alongside, same era.
“Multitask Learning with Low-Level Auxiliary Tasks for Encoder-Decoder Based Speech Recognition,”
Shubham Toshniwal, Hao Tang, Liang Lu, and Karen Livescu, · 2017
Later among the works it cites.
“Joint CTC-Attention based End-to-End Speech Recognition using Multi-task Learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Later among the works it cites.
“Multi-Accent Speech Recognition with Hierarchical Grapheme Based Models,”
Kanishka Rao and Hasim Sak, · 2017
Later among the works it cites.
“Direct Acoustics-to-Word Models for English Conversational Speech Recognition,”
Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon, Michael Picheny, and David Nahamoo, · 2017
Later among the works it cites.
“Subword and Crossword Units for CTC Acoustic Models,”
Thomas Zenkel, Ramon Sanabria, Florian Metze, and Alex Waibel, · 2017
Later among the works it cites.
“Advances in all-neural speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deep multi-task learning with low level tasks supervised at lower layers,”
Anders Søgaard and Yoav Goldberg, · 2016
Cited alongside, same era.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al., · 2016
Cited alongside, same era.
“Neural Machine Translation of Rare Words with Subword Units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Cited alongside, same era.
“Tensorflow: A System for Large-Scale Machine Learning,”
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al., · 2016
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in English and Mandarin,”
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al., · 2016
Cited alongside, same era.
“Analyzing Hidden Representations in End-to-End Automatic Speech Recognition Systems,”
Yonatan Belinkov and James Glass, · 2017
Cited alongside, same era.
Geoffrey Zweig, Chengzhu Yu, Jasha Droppo, and Andreas Stolcke, · 2017
Later among the works it cites.
“Deep contextualized word representations,”
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer., · 2018
Closest in time.
“Hierarchical Multi Task Learning With CTC,”
Ramon Sanabria and Florian Metze, · 2018
Closest in time.
“State-of-the-art Speech Recognition With Sequence-to-Sequence Models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Katya Gonina, et al., · 2018
Closest in time.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Closest in time.
“Building competitive direct acoustics-to-word models for English conversational speech recognition,”
Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, and Michael Picheny, · 2018
Closest in time.