Fetching the paper…
Reading the bibliography…
This paper introduces Moonshine, a family of speech recognition models optimized for live transcription and voice command processing.
The ami meeting corpus: A pre-announcement
Carletta, J., Ashby, S., Bourban, S., Flynn, M., Guillemot, M., Hain, T., Kadlec, J., Karaiskos, V., Kraaij, W., Kronenthal, M., et al · 2005
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization
Gupta, V., Koren, T., and Singer, Y · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2020
Earlier work this paper cites.
Mls: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2020
Cited alongside, same era.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
Chen, G., Chai, S., Wang, G., Du, J., Zhang, W.-Q., Weng, C., Su, D., Povey, D., Trmal, J., Zhang, J., et al · 2021
Cited alongside, same era.
The people’s speech: A large-scale diverse english speech recognition dataset for commercial usage
Galvez, D., Diamos, G., Ciro, J., Cerón, J. F., Achorn, K., Gopi, A., Kanter, D., Lam, M., Mazumder, M., and Reddi, V. J · 2021
Cited alongside, same era.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Gugger, S., Debut, L., Wolf, T., Schmid, P., Mueller, Z., Mangrulkar, S., Sun, M., and Bossan, B · 2022
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision, 2022
Open automatic speech recognition leaderboard
Srivastav, V., Majumdar, S., Koluguri, N., Moumen, A., Gandhi, S., et al · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., and et al., A. A · 2023
Later among the works it cites.
Defazio, A., Yang, X. A., Mehta, H., Mishchenko, K., Khaled, A., and Cutkosky, A · 2024
Closest in time.
Soap: Improving and stabilizing shampoo using adam
Vyas, N., Morwani, D., Zhao, R., Shapira, I., Brandfonbrener, D., Janson, L., and Kakade, S · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Cited alongside, same era.
Whisperx: Time-accurate speech transcription of long-form audio
Bain, M., Huh, J., Han, T., and Zisserman, A · 2023
Cited alongside, same era.
Closest in time.