Fetching the paper…
Reading the bibliography…
This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI.
“Introduction to RASC863,”
ChineseLDC.org, · 2004
Earlier work this paper cites.
“Google’s cross-dialect Arabic voice search,”
F. Biadsy, P.J. Moreno, and M. Jansche, · 2012
Earlier work this paper cites.
“Exemplar-based speech enhancement for deep neural network based automatic speech recognition,”
D. Baby, Jort F. Gemmeke, T. Virtanen, and H. Van hamme, · 2015
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Gender and dialect bias in Youtube’s automatic captions,”
R. Tatman, · 2017
Earlier work this paper cites.
“Multi-dialectical languages effect on speech recognition: Too much choice can hurt,”
M.G. Elfeky, Pedro Moreno, and V. Soto, · 2018
Earlier work this paper cites.
“Multi-dialect speech recognition with a single sequence-to-sequence model,”
B. Li, T.N. Sainath, K.C. Sim, M. Bacchiani, et al., · 2018
Earlier work this paper cites.
“Joint modeling of accents and acoustics for multi-accent speech recognition,”
X. Yang, K. Audhkhasi, A. Rosenberg, et al., · 2018
Earlier work this paper cites.
“Improved accented speech recognition using accent embeddings and multi-task learning.,”
A. Jain, M. Upreti, and P. Jyothi, · 2018
Earlier work this paper cites.
“Generalization through memorization: Nearest neighbor language models,”
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis, · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
T. Brown, B. Mann, N. Ryder, et al., · 2020
Earlier work this paper cites.
“Racial disparities in automated speech recognition,”
A. Koenecke, A. Nam, E. Lake, et al., · 2020
Earlier work this paper cites.
“Multimodal few-shot learning with frozen language models,”
M. Tsimpoukelli, J. Menick, S. Cabi, S.M. Eslami, O. Vinyals, and F. Hill, · 2021
Earlier work this paper cites.
“Calibrate before use: Improving few-shot performance of language models,”
Z. Zhao, E. Wallace, S. Feng, D. Klein, and S. Singh, · 2021
Cited alongside, same era.
“Innovative BERT-based reranking language models for speech recognition,”
S. Chiu and B. Chen, · 2021
Cited alongside, same era.
“Adapting GPT, GPT-2 and BERT language models for speech recognition,”
X. Zheng, C. Zhang, and P.C. Woodland, · 2021
Cited alongside, same era.
“Towards One Model to Rule All: Multilingual Strategy for Dialectal Code-Switching Arabic ASR,”
S.A. Chowdhury, A. Hussein, A. Abdelali, et al., · 2021
Cited alongside, same era.
“Quantifying bias in automatic speech recognition,”
S. Feng, O. Kudina, B.M. Halpern, and O. Scharenborg, · 2021
Cited alongside, same era.
“LLaMA: Open and efficient foundation language models,”
H. Touvron, T. Lavril, G. Izacard, X. Martinet, et al., · 2023
Closest in time.
“Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,”
W. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, et al., · 2023
Closest in time.
“Robust speech recognition via large-scale weak supervision,”
A. Radford, J.W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, · 2023
Closest in time.
“Scaling speech technology to 1,000+ languages,”
V. Pratap, A. Tjandra, B. Shi, P. Tomasello, et al., · 2023
Closest in time.
Z. Min and J. Wang, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer, · 2022
Cited alongside, same era.
“In-context learning for few-shot dialogue state tracking,”
U. Hu, C.-H. Lee, T. Xie, T. Yu, N.A. Smith, and M. Ostendorf, · 2022
Cited alongside, same era.
“Finetuned language models are zero-shot learners,”
J. Wei, M. Bosma, V.Y. Zhao, K. Guu, A.W. Yu, B. Lester, N. Du, A.M. Dai, and Q.V. Le, · 2022
Cited alongside, same era.
“What makes good in-context examples for GPT-3?,”
J. Liu, D. Shen, Y. Zhang, B. Dolan, L. Carin, and W. Chen, · 2022
Cited alongside, same era.
“RescoreBERT: Discriminative speech recognition rescoring with BERT,”
L. Xu, Y. Gu, J. Kolehmainen, et al., · 2022
Cited alongside, same era.
“An explanation of in-context learning as implicit Bayesian inference,”
S.M. Xie, A. Raghunathan, P. Liang, and T. Ma, · 2022
Cited alongside, same era.
“LoRA: Low-rank adaptation of large language models,”
E.J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, and W. Chen, · 2022
Cited alongside, same era.
Closest in time.
“Can ChatGPT detect intent? Evaluating large language models for spoken language understanding,”
M. He and P.N. Garner, · 2023
Closest in time.
“Prompting the hidden talent of web-scale speech models for zero-shot task generalization,”
P. Peng, B. Yan, S. Watanabe, and D. Harwath, · 2023
Closest in time.
“Prompting large language models with speech recognition abilities,”
Y. Fathullah, C. Wu, E. Lakomkin, et al., · 2023
Closest in time.
“On decoder-only architecture for speech-to-text and large language model integration,”
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu, and Y. Wu, · 2023
Closest in time.
“Making more of little data: Improving low-resource automatic speech recognition using data augmentation,”
M. Bartelds, N. San, B. McDonnell, D. Jurafsky, and M. Wieling, · 2023
Closest in time.
“Towards dialect-inclusive recognition in a low-resource language: Are balanced corpora the answer?,”
L. Lonergan, M. Qian, N.N. Chiaráin, C. Gobl, and A.N. Chasaide, · 2023
Closest in time.
“N-shot benchmarking of Whisper on diverse arabic speech recognition,”
B. Talafha, A. Waheed, and M. Abdul-Mageed, · 2023
Closest in time.