Fetching the paper…
Reading the bibliography…
Attention encoder-decoder model architecture is the backbone of several recent top performing foundation speech models: Whisper, Seamless, OWSM, and Canary-1B.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Earlier work this paper cites.
Accelerating recurrent neural network training using sequence bucketing and multi-gpu data parallelization
Viacheslav Khomenko, Oleg Shyshkov, Olga Radyvonenko, and Kostiantyn Bokhan. 2016 · 2016
Earlier work this paper cites.
A comprehensive study of batch construction strategies for recurrent neural networks in mxnet
Patrick Doetsch, Pavel Golik, and Hermann Ney. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting bleu scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl. 2019 · 2019
Cited alongside, same era.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber. 2020 · 2020
Cited alongside, same era.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
Towards measuring fairness in ai: the casual conversations dataset
Caner Hazirbas, Joanna Bitton, Brian Dolhansky, Jacqueline Pan, Albert Gordo, and Cristian Canton Ferrer. 2021 · 2021
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022 · 2022
Later among the works it cites.
Seamless: Multilingual expressive and streaming speech translation
Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, and 1 others. 2023 · 2023
Later among the works it cites.
Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling
Sanchit Gandhi, Patrick von Platen, and Alexander M. Rush. 2023 · 2023
Later among the works it cites.
Reproducing whisper-style training using an open-source toolkit and publicly available data
Yifan Peng, Jinchuan Tian, Brian Yan, Dan Berrebbi, Xuankai Chang, Xinjian Li, Jiatong Shi, Siddhant Arora, William Chen, Roshan Sharma, and 1 others. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah A Smith. 2021 · 2021
Cited alongside, same era.
Covost 2 and massively multilingual speech translation
Changhan Wang, Anne Wu, Jiatao Gu, and Juan Pino. 2021 · 2021
Cited alongside, same era.
Lhotse: a speech data representation library for the modern deep learning ecosystem
Piotr Żelasko, Daniel Povey, Jan Trmal, and Sanjeev Khudanpur. 2021 · 2021
Cited alongside, same era.
Fleurs: Few-shot learning evaluation of universal representations of speech
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna. 2022 · 2022
Cited alongside, same era.
https://github.com/dao-ailab/flash-attention
flash-attention
Cited in the paper.
https://github.com/k2-fsa/k2
k2
Cited in the paper.
https://github.com/webdataset/webdataset
webdataset
Cited in the paper.
Fast conformer with linearly scalable attention for efficient speech recognition
Dima Rekesh, Nithin Rao Koluguri, Samuel Kriman, Somshubra Majumdar, Vahid Noroozi, He Huang, Oleksii Hrinchuk, Krishna Puvvada, Ankur Kumar, Jagadeesh Balam, and 1 others. 2023 · 2023
Later among the works it cites.
Owsm v3. 1: Better and faster open whisper-style speech models based on e-branchformer
Yifan Peng, Jinchuan Tian, William Chen, Siddhant Arora, Brian Yan, Yui Sudo, Muhammad Shakeel, Kwanghee Choi, Jiatong Shi, Xuankai Chang, and 1 others. 2024 · 2024
Later among the works it cites.
Less is more: Accurate speech recognition & translation without web-scale data
Krishna C. Puvvada, Piotr Żelasko, He Huang, Oleksii Hrinchuk, Nithin Rao Koluguri, Kunal Dhawan, Somshubra Majumdar, Elena Rastorgueva, Zhehuai Chen, Vitaly Lavrukhin, Jagadeesh Balam, and Boris Ginsburg. 2024 · 2024
Later among the works it cites.
How does critical batch size scale in pre-training?
Hanlin Zhang, Depen Morwani, Nikhil Vyas, Jingfeng Wu, Difan Zou, Udaya Ghai, Dean Foster, and Sham Kakade. 2024 · 2024
Later among the works it cites.
Emmett: Efficient multimodal machine translation training
Piotr Żelasko, Zhehuai Chen, Mengru Wang, Daniel Galvez, Oleksii Hrinchuk, Shuoyang Ding, Ke Hu, Jagadeesh Balam, Vitaly Lavrukhin, and Boris Ginsburg. 2025 · 2025
Closest in time.