Fetching the paper…
Reading the bibliography…
The recently proposed Conformer model has become the de facto backbone model for various downstream speech tasks based on its hybrid attention-convolution architecture that captures both local and global features.
Timit acoustic phonetic continuous speech corpus
John S Garofolo · 1993
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Rigid-motion scattering for texture classification
Laurent Sifre and Stéphane Mallat · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Learning visual representations at scale
Vincent Vanhoucke · 2014
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
U-Net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
MobileNets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Towards end-to-end speech recognition with deep convolutional neural networks
Ying Zhang, Mohammad Pezeshki, Philémon Brakel, Saizheng Zhang, Cesar Laurent Yoshua Bengio, and Aaron Courville · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Hardware-aware softmax approximation for deep neural networks
Xue Geng, Jie Lin, Bin Zhao, Anmin Kong, Mohamed M Sabry Aly, and Vijay Chandrasekhar · 2018
Earlier work this paper cites.
Hardware-aware exponential approximation for deep neural network
Xue Geng, Jie Lin, Bin Zhao, Zhe Wang, Mohamed M Sabry Aly, and Vijay Chandrasekhar · 2018
Earlier work this paper cites.
Squeeze-and-Excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
State-of-the-art speech recognition using multi-stream self-attention with dilated 1d convolutions
Kyu J Han, Ramon Prieto, and Tao Ma · 2019
Earlier work this paper cites.
A comparative study on Transformer vs RNN in speech applications
Shigeki Karita et al · 2019
Cited alongside, same era.
Jasper: An end-to-end convolutional neural acoustic model
Jason Li, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen, Huyen Nguyen, and Ravi Teja Gadde · 2019
Cited alongside, same era.
Understanding and improving Transformer from a multi-particle dynamic system point of view
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2019
Cited alongside, same era.
RWTH ASR Systems for LibriSpeech: Hybrid vs attention–w/o data augmentation
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney · 2019
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
Efficient Conformer: Progressive downsampling and grouped attention for automatic speech recognition
Maxime Burchi and Valentin Vielzeuf · 2021
Later among the works it cites.
WavLM: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2021
Later among the works it cites.
Fast and accurate model scaling
Piotr Dollár, Mannat Singh, and Ross Girshick · 2021
Later among the works it cites.
Multiscale Vision Transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Recent developments on ESPNet toolkit boosted by Conformer
Pengcheng Guo et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Cited alongside, same era.
U-Time: A fully convolutional network for time series segmentation applied to sleep staging
Mathias Perslev, Michael Jensen, Sune Darkner, Poul Jørgen Jennum, and Christian Igel · 2019
Cited alongside, same era.
Wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
End-to-end ASR with adaptive span self-attention
Xuankai Chang, Aswin Shanmugam Subramanian, Pengcheng Guo, Shinji Watanabe, Yuya Fujita, and Motoi Omachi · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Cited alongside, same era.
Wei Han, Zhengdong Zhang, Yu Zhang, Jiahui Yu, Chung-Cheng Chiu, James Qin, Anmol Gulati, Ruoming Pang, and Yonghui Wu · 2020
Cited alongside, same era.
QuartzNet: Deep automatic speech recognition with 1d time-channel separable convolutions
Samuel Kriman, Stanislav Beliaev, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, and Yang Zhang · 2020
Cited alongside, same era.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Later among the works it cites.
I-BERT: Integer-only bert quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2021
Later among the works it cites.
Improved Multiscale Vision Transformers for classification and detection
Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Improving RNN Transducer based ASR with auxiliary tasks
Chunxi Liu, Frank Zhang, Duc Le, Suyoun Kim, Yatharth Saraf, and Geoffrey Zweig · 2021
Later among the works it cites.
Somshubra Majumdar, Jagadeesh Balam, Oleksii Hrinchuk, Vitaly Lavrukhin, Vahid Noroozi, and Boris Ginsburg · 2021
Later among the works it cites.
Pushing the limits of non-autoregressive speech recognition
Edwin G Ng, Chung-Cheng Chiu, Yu Zhang, and William Chan · 2021
Later among the works it cites.
http://nvdla.org/primer.html, 2021
NVDLA Primer · 2021
Later among the works it cites.
Understanding the role of self attention for efficient speech recognition
Kyuhong Shim, Jungwook Choi, and Wonyong Sung · 2021
Later among the works it cites.
Training data-efficient image Transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Unispeech: Unified speech representation learning with labeled and unlabeled data
Chengyi Wang, Yu Wu, Yao Qian, Kenichi Kumatani, Shujie Liu, Furu Wei, Michael Zeng, and Xuedong Huang · 2021
Later among the works it cites.
NN-LUT: Neural approximation of non-linear operations for efficient Transformer inference
Joonsang Yu, Junki Park, Seongmin Park, Minsoo Kim, Sihwa Lee, Dong Hyun Lee, and Jungwook Choi · 2021
Later among the works it cites.
On the usefulness of self-attention for automatic speech recognition with Transformers
Shucong Zhang, Erfan Loweimi, Peter Bell, and Steve Renals · 2021
Later among the works it cites.
Benchmarking lf-mmi, ctc and rnn-t criteria for streaming asr
Xiaohui Zhang, Frank Zhang, Chunxi Liu, Kjell Schubert, Julian Chan, Pradyot Prakash, Jun Liu, Ching-Feng Yeh, Fuchun Peng, Yatharth Saraf, et al · 2021
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli · 2022
Closest in time.
DeepNet: Scaling Transformers to 1,000 layers
Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, and Furu Wei · 2022
Closest in time.