Fetching the paper…
Reading the bibliography…
Connectionist Temporal Classification (CTC) is a widely used approach for automatic speech recognition (ASR) that performs conditionally independent monotonic alignment.
Source dependency-aware transformer with supervised self-attention
Chengyi Wang, Shuangzhi Wu, and Shujie Liu. 2019 · 1909
Earlier work this paper cites.
Une Approche théorique de l’Apprentissage Connexionniste: Applications à la Reconnaissance de la Parole
Léon Bottou. 1991 · 1991
Earlier work this paper cites.
Interactive translation of conversational speech
Alex Waibel. 1996 · 1996
Earlier work this paper cites.
Statistical methods for speech recognition
Frederick Jelinek. 1998 · 1998
Earlier work this paper cites.
Statistical machine translation
Yaser Al-Onaizan, Jan Curin, Michael Jahr, Kevin Knight, John Lafferty, Dan Melamed, Franz-Josef Och, David Purdy, Noah A Smith, and David Yarowsky. 1999 · 1999
Earlier work this paper cites.
Speech translation: coupling of recognition and translation
H. Ney. 1999 · 1999
Earlier work this paper cites.
Statistical phrase-based translation
Philipp Koehn, Franz J. Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
Glancing transformer for non-autoregressive neural machine translation
Lihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, and Lei Li. 2021 · 2003
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Speech recognition, machine translation, and speech translation—a unified discriminative learning paradigm [lecture notes]
Xiaodong He and Li Deng. 2011 · 2011
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012 · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves. 2012 · 2012
Earlier work this paper cites.
First-pass large vocabulary continuous speech recognition using bi-directional recurrent dnns
Awni Y Hannun, Andrew L Maas, Daniel Jurafsky, and Andrew Y Ng. 2014 · 2014
Earlier work this paper cites.
A historical perspective of speech recognition
Xuedong Huang, James Baker, and Raj Reddy. 2014 · 2014
Earlier work this paper cites.
Xsede: Accelerating scientific discovery
J. Towns, T. Cockerill, M. Dahan, I. Foster, K. Gaither, A. Grimshaw, V. Hazlewood, S. Lathrop, D. Lifka, G. D. Peterson, R. Roskies, J. R. Scott, and N. Wilkins-Diehr. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Bridges: a uniquely flexible hpc resource for new communities and data analytics
Nicholas A Nystrom, Michael J Levine, Ralph Z Roskies, and J Ray Scott. 2015 · 2015
Earlier work this paper cites.
Learning acoustic frame labeling for speech recognition with recurrent neural networks
Haşim Sak, Andrew Senior, Kanishka Rao, Ozan Irsoy, Alex Graves, Françoise Beaufays, and Johan Schalkwyk. 2015 · 2015
Earlier work this paper cites.
Alignment-based neural machine translation
Tamer Alkhouli, Gabriel Bretschner, Jan-Thorsten Peter, Mohammed Hethnawi, Andreas Guta, and Hermann Ney. 2016 · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
MuST-C: a Multilingual Speech Translation Corpus
Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019 · 2017
Earlier work this paper cites.
Joint ctc-attention based end-to-end speech recognition using multi-task learning
Suyoun Kim, Takaaki Hori, and Shinji Watanabe. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Hybrid ctc/attention architecture for end-to-end speech recognition
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi. 2017 · 2017
Cited alongside, same era.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
Linhao Dong, Shuang Xu, and Bo Xu. 2018 · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
End-to-end non-autoregressive neural machine translation with connectionist temporal classification
Jindřich Libovický and Jindřich Helcl. 2018 · 2018
Cited alongside, same era.
multi-bleu.perl
Moses-SMT. 2018 · 2018
Cited alongside, same era.
Correcting length bias in neural machine translation
Kenton Murray and David Chiang. 2018 · 2018
Best-first beam search
Clara Meister, Tim Vieira, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
On long-tailed phenomena in neural machine translation
Vikas Raunak, Siddharth Dalmia, Vivek Gupta, and Florian Metze. 2020 · 2020
Later among the works it cites.
Non-autoregressive machine translation with latent alignments
Chitwan Saharia, William Chan, Saurabh Saxena, and Mohammad Norouzi. 2020 · 2020
Later among the works it cites.
Alignment-enhanced transformer for constraining nmt with pre-specified translations
Kai Song, Kun Wang, Heng Yu, Yue Zhang, Zhongqiang Huang, Weihua Luo, Xiangyu Duan, and Min Zhang. 2020 · 2020
Later among the works it cites.
Investigating the reordering capability in CTC-based non-autoregressive end-to-end speech translation
Shun-Po Chuang, Yung-Sung Chuang, Chih-Chiang Chang, and Hung-yi Lee. 2021 · 2021
Later among the works it cites.
Searchable hidden intermediates for end-to-end models of decomposable sequence tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Hierarchical multitask learning with CTC
Ramon Sanabria and Florian Metze. 2018 · 2018
Cited alongside, same era.
ESPnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai. 2018 · 2018
Cited alongside, same era.
The Label Bias Problem
Awni Hannun. 2019 · 2019
Cited alongside, same era.
Triggered attention for end-to-end speech recognition
Niko Moritz, Takaaki Hori, and Jonathan Le Roux. 2019a · 2019
Cited alongside, same era.
Triggered attention for end-to-end speech recognition
Niko Moritz, Takaaki Hori, and Jonathan Le Roux. 2019b · 2019
Cited alongside, same era.
Siddharth Dalmia, Brian Yan, Vikas Raunak, Florian Metze, and Shinji Watanabe. 2021 · 2021
Later among the works it cites.
Synchronous syntactic attention for transformer neural machine translation
Hiroyuki Deguchi, Akihiro Tamura, and Takashi Ninomiya. 2021 · 2021
Later among the works it cites.
CTC-based compression for direct speech translation
Marco Gaido, Mauro Cettolo, Matteo Negri, and Marco Turchi. 2021 · 2021
Later among the works it cites.
Fully non-autoregressive neural machine translation: Tricks of the trade
Jiatao Gu and Xiang Kong. 2021 · 2021
Later among the works it cites.
Can latent alignments improve autoregressive machine translation?
Adi Haviv, Lior Vassertail, and Omer Levy. 2021 · 2021
Later among the works it cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Later among the works it cites.
Fast-md: Fast multi-decoder end-to-end speech translation with non-autoregressive hidden intermediates
Hirofumi Inaguma, Siddharth Dalmia, Brian Yan, and Shinji Watanabe. 2021a · 2021
Later among the works it cites.
The multilingual tedx corpus for speech recognition and translation
Elizabeth Salesky, Matthew Wiesner, Jacob Bremerman, Roldano Cattoni, Matteo Negri, Marco Turchi, Douglas W Oard, and Matt Post. 2021 · 2021
Later among the works it cites.
Streaming transformer asr with blockwise synchronous beam search
Emiru Tsunoo, Yosuke Kashiwagi, and Shinji Watanabe. 2021 · 2021
Later among the works it cites.
Neural machine translation with synchronous latent phrase structure
Shintaro Harada Taro Watanabe. 2021 · 2021
Later among the works it cites.
Neural machine translation with explicit phrase alignment
Jiacheng Zhang, Huanbo Luan, Maosong Sun, Feifei Zhai, Jingfang Xu, and Yang Liu. 2021 · 2021
Later among the works it cites.
Composable sparse fine-tuning for cross-lingual transfer
Alan Ansell, Edoardo Ponti, Anna Korhonen, and Ivan Vulić. 2022 · 2022
Closest in time.
LegoNN: Building modular encoder-decoder models
Siddharth Dalmia, Dmytro Okhonko, Mike Lewis, Sergey Edunov, Shinji Watanabe, Florian Metze, Luke Zettlemoyer, and Abdelrahman Mohamed. 2022 · 2022
Closest in time.
Blockwise Streaming Transformer for Spoken Language Understanding and Simultaneous Speech Translation
Keqi Deng, Shinji Watanabe, Jiatong Shi, and Siddhant Arora. 2022 · 2022
Closest in time.
Regularizing end-to-end speech translation with triangular decomposition agreement
Yichao Du, Zhirui Zhang, Weizhi Wang, Boxing Chen, Jun Xie, and Tong Xu. 2022 · 2022
Closest in time.
Hierarchical conditional end-to-end asr with ctc and multi-granular subword units
Yosuke Higuchi, Keita Karube, Tetsuji Ogawa, and Tetsunori Kobayashi. 2022 · 2022
Closest in time.
Non-autoregressive translation with layer-wise prediction and deep supervision
Chenyang Huang, Hao Zhou, Osmar R Zaïane, Lili Mou, and Lei Li. 2022 · 2022
Closest in time.
CMU’s IWSLT 2022 dialect speech translation system
Brian Yan, Patrick Fernandes, Siddharth Dalmia, Jiatong Shi, Yifan Peng, Dan Berrebbi, Xinyi Wang, Graham Neubig, and Shinji Watanabe. 2022 · 2022
Closest in time.
Revisiting end-to-end speech-to-text translation from scratch
Biao Zhang, Barry Haddow, and Rico Sennrich. 2022 · 2022
Closest in time.