Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel soft and monotonic alignment mechanism used for sequence transduction.
“Recherches quantitatives sur l’excitation electrique des nerfs traitee comme une polarization,”
Louis Lapicque, · 1907
Earlier work this paper cites.
“Networks of spiking neurons: The third generation of neural network models,”
Wolfgang Maass, · 1997
Earlier work this paper cites.
“Lapicque’s introduction of the integrate-and-fire model neuron (1907),”
Larry F Abbott, · 1999
Earlier work this paper cites.
“Hkust/mts: A very large scale mandarin telephone speech corpus,”
Yi Liu, Pascale Fung, Yongsheng Yang, Christopher Cieri, Shudong Huang, and David Graff, · 2006
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,”
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al., · 2016
Earlier work this paper cites.
“Layer normalization,”
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Earlier work this paper cites.
“Purely sequence-trained neural networks for asr based on lattice-free mmi.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“A comparison of sequence-to-sequence models for speech recognition.,”
Rohit Prabhavalkar, Kanishka Rao, Tara N Sainath, Bo Li, Leif Johnson, and Navdeep Jaitly, · 2017
Cited alongside, same era.
“Gaussian prediction based attention for online end-to-end speech recognition.,”
Junfeng Hou, Shiliang Zhang, and Li-Rong Dai, · 2017
Cited alongside, same era.
“Local monotonic attention mechanism for end-to-end speech and language processing,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Cited alongside, same era.
“Monotonic chunkwise attention,”
Chung-Cheng Chiu and Colin Raffel, · 2017
Cited alongside, same era.
“Joint ctc-attention based end-to-end speech recognition using multi-task learning,”
“Extending recurrent neural aligner for streaming end-to-end speech recognition in mandarin,”
Linhao Dong, Shiyu Zhou, Wei Chen, and Bo Xu, · 2018
Later among the works it cites.
“An analysis of local monotonic attention variants,”
André Merboldt, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Closest in time.
“An online attention-based model for speech recognition,”
Ruchao Fan, Pan Zhou, Wei Chen, Jia Jia, and Gang Liu, · 2019
Closest in time.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Closest in time.
“End-to-end speech recognition with adaptive computation steps,”
Mohan Li, Min Liu, and Hattori Masanori, · 2019
Closest in time.
“Self-attention aligner: A latency-control end-to-end model for asr using self-attention network and chunk-hopping,”
Linhao Dong, Feng Wang, and Bo Xu, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2017
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Cited alongside, same era.
“Aishell-2: Transforming mandarin asr research into industrial scale,”
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu, · 2018
Cited alongside, same era.
“wav2letter++: The fastest open-source speech recognition system,”
Vineel Pratap, Awni Hannun, Qiantong Xu, Jeff Cai, Jacob Kahn, Gabriel Synnaeve, Vitaliy Liptchinsky, and Ronan Collobert, · 2018
Cited alongside, same era.
“Improving end-to-end speech recognition with policy learning,”
Yingbo Zhou, Caiming Xiong, and Richard Socher, · 2018
Cited alongside, same era.
“A comparison of modeling units in sequence-to-sequence speech recognition with the transformer on mandarin chinese,”
Shiyu Zhou, Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
Closest in time.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Closest in time.
“Rwth asr systems for librispeech: Hybrid vs attention,”
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Closest in time.
“Jasper: An end-to-end convolutional neural acoustic model,”
Jason Li, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen, Huyen Nguyen, and Ravi Teja Gadde, · 2019
Closest in time.
“Transformers with convolutional context for asr,”
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer, · 2019
Closest in time.
“Self-attention networks for connectionist temporal classification in speech recognition,”
Julian Salazar, Katrin Kirchhoff, and Zhiheng Huang, · 2019
Closest in time.