Fetching the paper…
Reading the bibliography…
This paper presents BERT-CTC, a novel formulation of end-to-end speech recognition that adapts BERT for connectionist temporal classification (CTC).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
SPLAT: Speech-language joint pre-training for spoken language understanding
Yu-An Chung, Chenguang Zhu, and Michael Zeng. 2021 · 1907
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Fast end-to-end speech recognition via non-autoregressive models and cross-modal knowledge transferring from BERT
Ye Bai, Jiangyan Yi, Jianhua Tao, Zhengkun Tian, Zhengqi Wen, and Shuai Zhang. 2021 · 1911
Earlier work this paper cites.
Align-Refine: Non-autoregressive speech recognition via iterative realignment
Ethan A Chi, Julian Salazar, and Katrin Kirchhoff. 2021 · 1927
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Sequence labelling in structured domains with hierarchical recurrent neural networks
Santiago Fernández, Alex Graves, and Jürgen Schmidhuber. 2007 · 2007
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011 · 2011
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves. 2012 · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdelrahman Mohamed, and Geoffrey Hinton. 2013 · 2013
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Alex Graves and Navdeep Jaitly. 2014 · 2014
Earlier work this paper cites.
Deep Speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al. 2014 · 2014
Earlier work this paper cites.
Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks
Anthony Rousseau, Paul Deléglise, and Yannick Estève. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
XSEDE: Accelerating scientific discovery
J. Towns, T. Cockerill, M. Dahan, I. Foster, K. Gaither, A. Grimshaw, V. Hazlewood, S. Lathrop, D. Lifka, G. D. Peterson, R. Roskies, J. R. Scott, and N. Wilkins-Diehr. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Audio augmentation for speech recognition
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Bridges: A uniquely flexible HPC resource for new communities and data analytics
Nicholas A Nystrom, Michael J Levine, Ralph Z Roskies, and J Ray Scott. 2015 · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng. 2017 · 2017
Earlier work this paper cites.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
State-of-the-art speech recognition with sequence-to-sequence models
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani. 2018 · 2018
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2018 · 2018
Earlier work this paper cites.
An analysis of incorporating an external language model into a sequence-to-sequence model
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, Zhijeng Chen, and Rohit Prabhavalkar. 2018 · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
A time-restricted self-attention layer for ASR
Daniel Povey, Hossein Hadian, Pegah Ghahremani, Ke Li, and Sanjeev Khudanpur. 2018 · 2018
Cited alongside, same era.
Hierarchical multitask learning with CTC
Ramon Sanabria and Florian Metze. 2018 · 2018
Cited alongside, same era.
Cold fusion: Training seq2seq models together with language models
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates. 2018 · 2018
Cited alongside, same era.
ESPnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai. 2018 · 2018
Mask CTC: Non-autoregressive end-to-end ASR with CTC and mask predict
Yosuke Higuchi, Shinji Watanabe, Nanxin Chen, Tetsuji Ogawa, and Tetsunori Kobayashi. 2020 · 2020
Later among the works it cites.
Improving the prosody of RNN-based english text-to-speech synthesis by incorporating a BERT model
Tom Kenter, Manish Sharma, and Rob Clark. 2020 · 2020
Later among the works it cites.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q Nguyen, and Katrin Kirchhoff. 2020 · 2020
Later among the works it cites.
DEJA-VU: Double feature presentation and iterated loss in deep Transformer networks
Andros Tjandra, Chunxi Liu, Frank Zhang, Xiaohui Zhang, Yongqiang Wang, Gabriel Synnaeve, Satoshi Nakamura, and Geoffrey Zweig. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Recent trends in deep learning based natural language processing [review article]
Tom Young, Devamanyu Hazarika, Soujanya Poria, and Erik Cambria. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
Levenshtein Transformer
Jiatao Gu, Changhan Wang, and Junbo Zhao. 2019 · 2019
Cited alongside, same era.
Pre-trained text embeddings for enhanced text-to-speech synthesis
Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Shubham Toshniwal, and Karen Livescu. 2019 · 2019
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Cited alongside, same era.
A study of Transducer based end-to-end ASR with ESPnet: Architecture, auxiliary loss and decoding strategies
Florian Boyer, Yusuke Shinohara, Takaaki Ishii, Hirofumi Inaguma, and Shinji Watanabe. 2021 · 2021
Later among the works it cites.
Align-Denoise: Single-pass non-autoregressive speech recognition
Nanxin Chen, Piotr Żelasko, Laureano Moro-Velázquez, Jesús Villalba, and Najim Dehak. 2021 · 2021
Later among the works it cites.
Innovative BERT-based reranking language models for speech recognition
Shih-Hsuan Chiu and Berlin Chen. 2021 · 2021
Later among the works it cites.
Improving hybrid CTC/attention end-to-end speech recognition with pretrained acoustic and language models
Keqi Deng, Songjun Cao, Yike Zhang, and Long Ma. 2021 · 2021
Later among the works it cites.
ASR rescoring and confidence estimation with ELECTRA
Hayato Futami, Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara. 2021 · 2021
Later among the works it cites.
Recent developments on ESPnet toolkit boosted by Conformer
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi, Shinji Watanabe, Kun Wei, Wangyou Zhang, and Yuekai Zhang. 2021 · 2021
Later among the works it cites.
A comparative study on non-autoregressive modelings for speech-to-text generation
Yosuke Higuchi, Nanxin Chen, Yuya Fujita, Hirofumi Inaguma, Tatsuya Komatsu, Jaesong Lee, Jumon Nozaki, Tianzi Wang, and Shinji Watanabe. 2021a · 2021
Later among the works it cites.
Improved mask-CTC for non-autoregressive end-to-end ASR
Yosuke Higuchi, Hirofumi Inaguma, Shinji Watanabe, Tetsuji Ogawa, and Tetsunori Kobayashi. 2021b · 2021
Later among the works it cites.
Speech recognition by simply fine-tuning BERT
Wen-Chin Huang, Chia-Hua Wu, Shang-Bao Luo, Kuan-Yu Chen, Hsin-Min Wang, and Tomoki Toda. 2021 · 2021
Later among the works it cites.
Multitask learning and joint optimization for Transformer-RNN-Transducer speech recognition
Jae-Jin Jeon and Eesung Kim. 2021 · 2021
Later among the works it cites.
Layer pruning on demand with intermediate CTC
Jaesong Lee, Jingu Kang, and Shinji Watanabe. 2021 · 2021
Later among the works it cites.
Intermediate loss regularization for CTC-based speech recognition
Jaesong Lee and Shinji Watanabe. 2021 · 2021
Later among the works it cites.
Relaxing the conditional independence assumption of CTC-based ASR by conditioning on intermediate predictions
Jumon Nozaki and Tatsuya Komatsu. 2021 · 2021
Later among the works it cites.
Efficiently fusing pretrained acoustic and linguistic encoders for low-resource speech recognition
Cheng Yi, Shiyu Zhou, and Bo Xu. 2021 · 2021
Later among the works it cites.
Wav-BERT: Cooperative acoustic and linguistic representation learning for low-resource speech recognition
Guolin Zheng, Yubei Xiao, Ke Gong, Pan Zhou, Xiaodan Liang, and Liang Lin. 2021 · 2021
Later among the works it cites.
ESPnet-SLU: Advancing spoken language understanding through ESPnet
Siddhant Arora, Siddharth Dalmia, Pavel Denisov, Xuankai Chang, Yushi Ueda, Yifan Peng, Yuekai Zhang, Sujay Kumar, Karthik Ganesan, Brian Yan, Ngoc Thang Vu, Alan W Black, and Shinji Watanabe. 2022 · 2022
Closest in time.
Improving non-autoregressive end-to-end speech recognition with pre-trained acoustic and language models
Keqi Deng, Zehui Yang, Shinji Watanabe, Yosuke Higuchi, Gaofeng Cheng, and Pengyuan Zhang. 2022 · 2022
Closest in time.
Hierarchical conditional end-to-end ASR with CTC and multi-granular subword units
Yosuke Higuchi, Keita Karube, Tetsuji Ogawa, and Tetsunori Kobayashi. 2022 · 2022
Closest in time.
Knowledge transfer from large-scale pretrained language models to end-to-end speech recognizers
Yotaro Kubo, Shigeki Karita, and Michiel Bacchiani. 2022 · 2022
Closest in time.
Memory-efficient training of RNN-Transducer with sampled softmax
Jaesong Lee, Lukas Lee, and Shinji Watanabe. 2022 · 2022
Closest in time.
Integration of pre-trained networks with continuous token interface for end-to-end spoken language understanding
Seunghyun Seo, Donghyun Kwak, and Bowon Lee. 2022 · 2022
Closest in time.
Effect and analysis of large-scale language model rescoring on competitive ASR systems
Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Nobuyasu Itoh, and George Saon. 2022 · 2022
Closest in time.
Non-autoregressive ASR modeling using pre-trained language models for Chinese speech recognition
Fu-Hao Yu, Kuan-Yu Chen, and Ke-Han Lu. 2022 · 2022
Closest in time.