Fetching the paper…
Reading the bibliography…
The recently proposed Conformer architecture has shown state-of-the-art performances in Automatic Speech Recognition by combining convolution with attention to model both local and global dependencies.
“An analysis of noise in recurrent neural networks: convergence and generalization,”
Kam-Chuen Jim, C Lee Giles, and Bill G Horne, · 1996
Earlier work this paper cites.
“Pruning neural networks with distribution estimation algorithms,”
Erick Cantú-Paz, · 2003
Earlier work this paper cites.
“Model compression,”
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil, · 2006
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Kenlm: Faster and smaller language model queries,”
Kenneth Heafield, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Restructuring of deep neural network acoustic models with singular value decomposition.,”
Jian Xue, Jinyu Li, and Yifan Gong, · 2013
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Deep speech: Scaling up end-to-end speech recognition,”
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al., · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“High-performance hardware for machine learning,”
William Dally, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell,”
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Wav2letter: an end-to-end convnet-based speech recognition system,”
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Cited alongside, same era.
“Language modeling with gated convolutional networks,”
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier, · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Mobilenetv2: Inverted residuals and linear bottlenecks,”
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen, · 2018
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Squeeze-and-excitation networks,”
Jie Hu, Li Shen, and Gang Sun, · 2018
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov, · 2019
Later among the works it cites.
“Pytorch: An imperative style, high-performance deep learning library,”
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala, · 2019
Later among the works it cites.
“Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions,”
Samuel Kriman, Stanislav Beliaev, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, and Yang Zhang, · 2020
Later among the works it cites.
“An overview of neural network compression,”
James O’ Neill, · 2020
Later among the works it cites.
“Contextnet: Improving convolutional neural networks for automatic speech recognition with global context,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Image transformer,”
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran, · 2018
Cited alongside, same era.
“Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“Recurrent stacking of layers for compact neural machine translation models,”
Raj Dabre and Atsushi Fujita, · 2019
Cited alongside, same era.
“Efficientnet: Rethinking model scaling for convolutional neural networks,”
Mingxing Tan and Quoc Le, · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Cited alongside, same era.
“Jasper: An end-to-end convolutional neural acoustic model,”
Jason Li, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen, Huyen Nguyen, and Ravi Teja Gadde, · 2019
Cited alongside, same era.
Wei Han, Zhengdong Zhang, Yu Zhang, Jiahui Yu, Chung-Cheng Chiu, James Qin, Anmol Gulati, Ruoming Pang, and Yonghui Wu, · 2020
Later among the works it cites.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Later among the works it cites.
“Specaugment on large scale datasets,”
Daniel S Park, Yu Zhang, Chung-Cheng Chiu, Youzheng Chen, Bo Li, William Chan, Quoc V Le, and Yonghui Wu, · 2020
Later among the works it cites.
Somshubra Majumdar, Jagadeesh Balam, Oleksii Hrinchuk, Vitaly Lavrukhin, Vahid Noroozi, and Boris Ginsburg, · 2021
Closest in time.
“Recent developments on espnet toolkit boosted by conformer,”
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, et al., · 2021
Closest in time.
“Bottleneck transformers for visual recognition,”
Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, and Ashish Vaswani, · 2021
Closest in time.
“Efficient conformer-based speech recognition with linear attention,”
Shengqiang Li, Menglong Xu, and Xiao-Lei Zhang, · 2021
Closest in time.
“Efficient attention: Attention with linear complexities,”
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li, · 2021
Closest in time.
“Efficient conformer with prob-sparse attention mechanism for end-to-endspeech recognition,”
Xiong Wang, Sining Sun, Lei Xie, and Long Ma, · 2021
Closest in time.