Fetching the paper…
Reading the bibliography…
Conformer-based models have become the dominant end-to-end architecture for speech processing tasks.
“The design for the Wall Street Journal based CSR corpus,”
D. B. Paul and J. M. Baker, · 1992
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Librispeech: an ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, L ukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Xception: Deep learning with depthwise separable convolutions,”
F. Chollet, · 2017
Earlier work this paper cites.
“Xception: Deep learning with depthwise separable convolutions,”
François Chollet, · 2017
Earlier work this paper cites.
“Mozilla: A journey to less than 10% word error rate,” https://hacks.mozilla.org/2017/11/a-journey-to-10-word-error-rate/
2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
Taku Kudo and John Richardson, · 2018
Earlier work this paper cites.
“Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,”
François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Esteve, · 2018
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Cited alongside, same era.
“QuartzNet: Deep automatic speech recognition with 1D time-channel separable convolutions,”
Samuel Kriman, Stanislav Beliaev, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, and Yang Zhang, · 2020
Cited alongside, same era.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Yu Zhang, James Qin, Daniel S Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V Le, and Yonghui Wu, · 2020
Cited alongside, same era.
Sandeep Subramanian, Oleksii Hrinchuk, Virginia Adams, and Oleksii Kuchaiev, · 2021
Later among the works it cites.
“MuST-C: A multilingual corpus for end-to-end speech translation,”
Roldano Cattoni, Mattia Antonino Di Gangi, Luisa Bentivogli, Matteo Negri, and Marco Turchi, · 2021
Later among the works it cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Later among the works it cites.
Yingzhi Wang, Abdelmoumene Boumadane, and Abdelwahab Heba, · 2021
Later among the works it cites.
“Earnings-21: A practical benchmark for ASR in the wild,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Iz Beltagy, Matthew E. Peters, and Arman Cohan, · 2020
Cited alongside, same era.
“MLS: A large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert, · 2020
Cited alongside, same era.
“Libri-Light: A benchmark for ASR with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
Cited alongside, same era.
“SLURP: A Spoken Language Understanding Resource Package,”
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Recent developments on Espnet toolkit boosted by Conformer,”
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, et al., · 2021
Cited alongside, same era.
“Efficient Conformer: Progressive downsampling and grouped attention for automatic speech recognition,”
Maxime Burchi and Valentin Vielzeuf, · 2021
Cited alongside, same era.
Miguel Del Rio, Natalie Delworth, Ryan Westerman, Michelle Huang, Nishchal Bhandari, Joseph Palakapilly, Quinten McNamara, Joshua Dong, Piotr Żelasko, and Miguel Jetté, · 2021
Later among the works it cites.
“Squeezeformer: An efficient Transformer for automatic speech recognition,”
Sehoon Kim, Amir Gholami, Albert Shaw, Nicholas Lee, Karttikeya Mangalam, Jitendra Malik, Michael W Mahoney, and Kurt Keutzer, · 2022
Later among the works it cites.
“Uconv-conformer: High reduction of input sequence length for end-to-end speech recognition,”
Andrei Andrusenko, Rauf Nasretdinov, and Aleksei Romanenko, · 2022
Later among the works it cites.
“Findings of the iwslt 2022 evaluation campaign,”
Antonios Anastasopoulos, Loïc Barrault, Luisa Bentivogli, Marcely Zanon Boito, Ondřej Bojar, Roldano Cattoni, Anna Currey, Georgiana Dinu, Kevin Duh, Maha Elbayad, et al., · 2022
Later among the works it cites.
“ESPnet-SLU: Advancing spoken language understanding through ESPnet,”
Siddhant Arora, Siddharth Dalmia, Pavel Denisov, Xuankai Chang, Yushi Ueda, Yifan Peng, Yuekai Zhang, Sujay Kumar, Karthik Ganesan, Brian Yan, et al., · 2022
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision,” 2022
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2022
Later among the works it cites.
“Open automatic speech recognition leaderboard,” https://huggingface.co/spaces/huggingface.co/spaces/open-asr-leaderboard/leaderboard , 2023
Vaibhav Srivastav, Somshubra Majumdar, Nithin Koluguri, Adel Moumen, Sanchit Gandhi, Hugging Face Team, Nvidia NeMo Team, and SpeechBrain Team, · 2023
Closest in time.