Fetching the paper…
Reading the bibliography…
We introduce a novel segmental-attention model for automatic speech recognition.
“Switchboard: Telephone speech corpus for research and development,”
John J Godfrey, Edward C Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“From HMM’s to segment models: A unified view of stochastic modeling for speech recognition,”
Mari Ostendorf, Vassilios V Digalakis, and Owen A Kimball, · 1996
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Online automatic speech recognition with listen, attend and spell model,”
Roger Hsiao, Dogan Can, Tim Ng, Ruchir Travadi, and Arnab Ghoshal, · 2008
Earlier work this paper cites.
“Rasr - the rwth aachen university open source speech recognition toolkit,”
David Rybach, Stefan Hahn, Patrick Lehnen, David Nolden, Martin Sundermeyer, Zoltán Tüske, Simon Wiesler, Ralf Schlüter, and Hermann Ney, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,” Preprint arXiv:1211.3711, 2012
Alex Graves, · 2012
Earlier work this paper cites.
“Generating sequences with recurrent neural networks,” Preprint arXiv:1308.0850, 2013
Alex Graves, · 2013
Earlier work this paper cites.
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Effective approaches to attention-based neural machine translation,”
Thang Luong, Hieu Pham, and Christopher D. Manning, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Earlier work this paper cites.
“An online sequence-to-sequence model using partial conditioning,”
Navdeep Jaitly, Quoc V Le, Oriol Vinyals, Ilya Sutskever, David Sussillo, and Samy Bengio, · 2016
Earlier work this paper cites.
“Segmental recurrent neural networks,”
Lingpeng Kong, Chris Dyer, and Noah A Smith, · 2016
Earlier work this paper cites.
“Segmental recurrent neural networks for end-to-end speech recognition,”
Liang Lu, Lingpeng Kong, Chris Dyer, Noah A Smith, and Steve Renals, · 2016
Earlier work this paper cites.
“Online segment to segment neural transduction,”
Lei Yu, Jan Buys, and Phil Blunsom, · 2016
Earlier work this paper cites.
“Alignment-based neural machine translation,”
Tamer Alkhouli, Gabriel Bretschner, Jan-Thorsten Peter, Mohammed Hethnawi, Andreas Guta, and Hermann Ney, · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Local monotonic attention mechanism for end-to-end speech and language processing,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Cited alongside, same era.
“Online and linear-time attention by enforcing monotonic alignments,”
Colin Raffel, Thang Luong, Peter J Liu, Ron J Weiss, and Douglas Eck, · 2017
Cited alongside, same era.
“Gaussian prediction based attention for online end-to-end speech recognition,”
Junfeng Hou, Shiliang Zhang, and Lirong Dai, · 2017
Cited alongside, same era.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Cited alongside, same era.
“Joint CTC/attention decoding for end-to-end speech recognition,”
Takaaki Hori, Shinji Watanabe, and John R Hershey, · 2017
Cited alongside, same era.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
“A comparison of end-to-end models for long-form speech recognition,”
Chung-Cheng Chiu, Wei Han, Yu Zhang, Ruoming Pang, Sergey Kishchenko, Patrick Nguyen, Arun Narayanan, Hank Liao, Shuyuan Zhang, Anjuli Kannan, Rohit Prabhavalkar, Zhifeng Chen, Tara N. Sainath, and Yonghui Wu, · 2019
Later among the works it cites.
“An online attention-based model for speech recognition,”
Ruchao Fan, Pan Zhou, Wei Chen, Jia Jia, and Gang Liu, · 2019
Later among the works it cites.
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li, and Colin Raffel, · 2019
Later among the works it cites.
“Windowed attention mechanisms for speech recognition,”
Shucong Zhang, Erfan Loweimi, Peter Bell, and Steve Renals, · 2019
Later among the works it cites.
“An analysis of local monotonic attention variants,”
Andre Merboldt, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Cited alongside, same era.
“The neural noisy channel,”
Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Tomas Kocisky, · 2017
Cited alongside, same era.
“Inverted alignments for end-to-end automatic speech recognition,”
Patrick Doetsch, Mirko Hannemann, Ralf Schlueter, and Hermann Ney, · 2017
Cited alongside, same era.
“Hybrid neural network alignment and lexicon model in direct HMM for statistical machine translation,”
Weiyue Wang, Tamer Alkhouli, Derui Zhu, and Hermann Ney, · 2017
Cited alongside, same era.
“Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping,”
Hasim Sak, Matt Shannon, Kanishka Rao, and Françoise Beaufays, · 2017
Cited alongside, same era.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Cited alongside, same era.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2018
Cited alongside, same era.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Later among the works it cites.
“Streaming end-to-end speech recognition with joint CTC-attention based models,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Later among the works it cites.
“Online hybrid CTC/attention architecture for end-to-end speech recognition,”
Haoran Miao, Gaofeng Cheng, Pengyuan Zhang, Ta Li, and Yonghong Yan, · 2019
Later among the works it cites.
“Two-pass end-to-end speech recognition,”
Tara Sainath, Ruoming Pang, David Rybach, Yanzhang (Ryan) He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, Ian McGraw, and Chung-Cheng Chiu, · 2019
Later among the works it cites.
“A comparison of Transformer and LSTM encoder decoder models for ASR,”
Albert Zeyer, Parnia Bahar, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“Training language models for long-span cross-sentence evaluation,”
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Later among the works it cites.
“Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard,”
Zoltán Tüske, George Saon, Kartik Audhkhasi, and Brian Kingsbury, · 2020
Later among the works it cites.
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
Tara N. Sainath, Yanzhang He, Bo Li, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo-yiin Chang, Wei Li, Raziel Alvarez, Zhifeng Chen, Chung-Cheng Chiu, David Garcia, Alex Gruenstein, Ke Hu, Anjuli Kannan, Qiao Liang, Ian McGraw, Cal Peyser, Rohit Prabhavalkar, Golan Pundak, David Rybach, Yuan Shangguan, Yash Sheth, Trevor Strohman, Mirkó Visontai, Yonghui Wu, Yu Zhang, and Ding Zhao, · 2020
Later among the works it cites.
Alignment-Based Neural Networks for Machine Translation
Tamer Alkhouli, · 2020
Later among the works it cites.
“A new training pipeline for an improved neural transducer,”
Albert Zeyer, André Merboldt, Ralf Schlüter, and Hermann Ney, · 2020
Later among the works it cites.
“A study of latent monotonic attention variants,” Preprint arXiv:2103.16710, Mar. 2021
Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2021
Later among the works it cites.
“Equivalence of segmental and neural transducer modeling: A proof of concept,”
Wei Zhou, Albert Zeyer, André Merboldt, Ralf Schlüter, and Hermann Ney, · 2021
Later among the works it cites.
“Investigating methods to improve language model integration for attention-based encoder-decoder ASR models,”
Mohammad Zeineldeen, Aleksandr Glushko, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2021
Later among the works it cites.
“Phoneme based neural transducer for large vocabulary speech recognition,”
Wei Zhou, Simon Berger, Ralf Schlüter, and Hermann Ney, · 2021
Later among the works it cites.
“On the limit of English conversational speech recognition,”
Zoltán Tüske, George Saon, and Brian Kingsbury, · 2066
Closest in time.