Fetching the paper…
Reading the bibliography…
Sequence-to-Sequence (seq2seq) tasks transcribe the input sequence to a target sequence.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Offline handwriting recognition with multidimensional recurrent neural networks
Alex Graves and Jürgen Schmidhuber · 2008
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Xsede: Accelerating scientific discovery
J. Towns, T. Cockerill, M. Dahan, I. Foster, K. Gaither, A. Grimshaw, V. Hazlewood, S. Lathrop, D. Lifka, G. D. Peterson, R. Roskies, J. R. Scott, and N. Wilkins-Diehr · 2014
Earlier work this paper cites.
Framewise and ctc training of neural networks for handwriting recognition
Théodore Bluche, Hermann Ney, Jérôme Louradour, and Christopher Kermorvant · 2015
Earlier work this paper cites.
Bridges: a uniquely flexible hpc resource for new communities and data analytics
Nicholas A Nystrom, Michael J Levine, Ralph Z Roskies, and J Ray Scott · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Learning acoustic frame labeling for speech recognition with recurrent neural networks
Haşim Sak, Andrew Senior, Kanishka Rao, Ozan Irsoy, Alex Graves, Françoise Beaufays, and Johan Schalkwyk · 2015
Earlier work this paper cites.
Acoustic modelling with cd-ctc-smbr lstm rnns
Andrew Senior, Haşim Sak, Félix de Chaumont Quitry, Tara Sainath, and Kanishka Rao · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2016
Earlier work this paper cites.
Online detection and classification of dynamic hand gestures with recurrent 3d convolutional neural networks
Pavlo Molchanov, Xiaodong Yang, Shalini Gupta, Kihwan Kim, Stephen Tyree, and Jan Kautz · 2016
Earlier work this paper cites.
Lipnet: End-to-end sentence-level lipreading
Yannis M Assael, Brendan Shillingford, Shimon Whiteson, and Nando de Freitas · 2017
Earlier work this paper cites.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng · 2017
Earlier work this paper cites.
Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping
Haşim Sak, Matt Shannon, Kanishka Rao, and Françoise Beaufays · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Hybrid ctc/attention architecture for end-to-end speech recognition
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi · 2017
Earlier work this paper cites.
End-to-end automatic speech translation of audiobooks
Alexandre Bérard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin · 2018
Earlier work this paper cites.
Aishell-2: Transforming mandarin asr research into industrial scale
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu · 2018
Earlier work this paper cites.
Advancing multi-accented lstm-ctc speech recognition using a domain specific student-teacher learning paradigm
Shahram Ghorbani, Ahmet E. Bulut, and John H.L. Hansen · 2018
Cited alongside, same era.
End-to-end speech recognition using lattice-free mmi
Hossein Hadian, Hossein Sameti, Daniel Povey, and Sanjeev Khudanpur · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Cited alongside, same era.
Taco: Learning task decomposition via temporal alignment for control
Kyriacos Shiarlis, Markus Wulfmeier, Sasha Salter, Shimon Whiteson, and Ingmar Posner · 2018
Cited alongside, same era.
Connectionist temporal fusion for sign language translation
Shuo Wang, Dan Guo, Wen-gang Zhou, Zheng-Jun Zha, and Meng Wang · 2018
Cited alongside, same era.
MuST-C: a Multilingual Speech Translation Corpus
Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi · 2019
CTC-based compression for direct speech translation
Marco Gaido, Mauro Cettolo, Matteo Negri, and Marco Turchi · 2021
Later among the works it cites.
Reducing Streaming ASR Model Delay with Self Alignment
Jaeyoung Kim, Han Lu, Anshuman Tripathi, Qian Zhang, and Hasim Sak · 2021
Later among the works it cites.
Intermediate loss regularization for ctc-based speech recognition
Jaesong Lee and Shinji Watanabe · 2021
Later among the works it cites.
Alignment restricted streaming recurrent neural network transducer
Jay Mahadeokar, Yuan Shangguan, Duc Le, Gil Keren, Hang Su, Thong Le, Ching-Feng Yeh, Christian Fuegen, and Michael L. Seltzer · 2021
Later among the works it cites.
Semi-supervised speech recognition via graph-based temporal classification
Niko Moritz, Takaaki Hori, and Jonathan Le Roux · 2021
Later among the works it cites.
Glancing transformer for non-autoregressive neural machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dense temporal convolution network for sign language translation
Dan Guo, Shuo Wang, Qi Tian, and Meng Wang · 2019
Cited alongside, same era.
Guiding ctc posterior spike timings for improved posterior fusion and knowledge distillation
Gakuto Kurata and Kartik Audhkhasi · 2019
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le · 2019
Cited alongside, same era.
Towards real-time mispronunciation detection in kids’ speech
Peter Plantinga and Eric Fosler-Lussier · 2019
Cited alongside, same era.
Sign language transformers: Joint end-to-end sign language recognition and translation
Necati Cihan Camgoz, Oscar Koller, Simon Hadfield, and Richard Bowden · 2020
Cited alongside, same era.
Fully non-autoregressive neural machine translation: Tricks of the trade
Jiatao Gu and Xiang Kong · 2020
Cited alongside, same era.
Lihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, and Lei Li · 2021
Later among the works it cites.
Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition
Yangyang Shi, Yongqiang Wang, Chunyang Wu, Ching-Feng Yeh, Julian Chan, Frank Zhang, Duc Le, and Mike Seltzer · 2021
Later among the works it cites.
Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit
Zhuoyuan Yao, Di Wu, Xiong Wang, Binbin Zhang, Fan Yu, Chao Yang, Zhendong Peng, Xiaoyu Chen, Lei Xie, and Xin Lei · 2021
Later among the works it cites.
Fastemit: Low-latency streaming asr with sequence-level emission regularization
Jiahui Yu, Chung-Cheng Chiu, Bo Li, Shuo-yiin Chang, Tara N. Sainath, Yanzhang He, Arun Narayanan, Wei Han, Anmol Gulati, Yonghui Wu, and Ruoming Pang · 2021
Later among the works it cites.
Why does ctc result in peaky behavior?
Albert Zeyer, Ralf Schlüter, and Hermann Ney · 2021
Later among the works it cites.
Legonn: Building modular encoder-decoder models
Siddharth Dalmia, Dmytro Okhonko, Mike Lewis, Sergey Edunov, Shinji Watanabe, Florian Metze, Luke Zettlemoyer, and Abdelrahman Mohamed · 2022
Closest in time.
Non-autoregressive translation with layer-wise prediction and deep supervision
Chenyang Huang, Hao Zhou, Osmar R Zaïane, Lili Mou, and Lei Li · 2022
Closest in time.
Pruned RNN-T for fast, memory-efficient ASR training
Fangjun Kuang, Liyong Guo, Wei Kang, Long Lin, Mingshuang Luo, Zengwei Yao, and Daniel Povey · 2022
Closest in time.
CTC Variations Through New WFST Topologies
Aleksandr Laptev, Somshubra Majumdar, and Boris Ginsburg · 2022
Closest in time.
Star temporal classification: Sequence classification with partially labeled data
Vineel Pratap, Awni Hannun, Gabriel Synnaeve, and Ronan Collobert · 2022
Closest in time.
Minimum latency training of sequence transducers for streaming end-to-end speech recognition
Yusuke Shinohara and Shinji Watanabe · 2022
Closest in time.
Ctc alignments improve autoregressive translation
Brian Yan, Siddharth Dalmia, Yosuke Higuchi, Graham Neubig, Florian Metze, Alan W Black, and Shinji Watanabe · 2022
Closest in time.
Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al · 2022
Closest in time.
Investigating sequence-level normalisation for ctc-like end-to-end asr
Zeyu Zhao and Peter Bell · 2022
Closest in time.