Fetching the paper…
Reading the bibliography…
Conventional end-to-end Automatic Speech Recognition (ASR) models primarily focus on exact transcription tasks, lacking flexibility for nuanced user interactions.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell,”
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al., · 2016
Earlier work this paper cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Earlier work this paper cites.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Y. Zhang, J. Qin, D. S. Park, W. Han, C.-C. Chiu, R. Pang, Q. V. Le, and Y. Wu, · 2020
Earlier work this paper cites.
“On generative spoken language modeling from raw audio,”
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed, et al., · 2021
Earlier work this paper cites.
“Wavprompt: Towards few-shot spoken language understanding with frozen language models,”
H. Gao, J. Ni, K. Qian, Y. Zhang, S. Chang, and M. Hasegawa-Johnson, · 2022
Earlier work this paper cites.
“Speechprompt: An exploration of prompt tuning on generative spoken language model for speech processing tasks,”
K.-W. Chang, W.-C. Tseng, S.-W. Li, and H.-y. Lee, · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback,”
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al., · 2022
Earlier work this paper cites.
“Multitask prompted training enables zero-shot task generalization,”
V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, T. L. Scao, A. Raja, et al., · 2022
Earlier work this paper cites.
“Finetuned language models are zero-shot learners,”
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, · 2022
Cited alongside, same era.
“Scaling instruction-finetuned language models,”
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al., · 2022
Cited alongside, same era.
“Palm: Scaling language modeling with pathways,”
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al., · 2022
Cited alongside, same era.
“Audiogpt: Understanding and generating speech, music, sound, and talking head,”
R. Huang, M. Li, D. Yang, J. Shi, X. Chang, Z. Ye, Y. Wu, Z. Hong, J. Huang, J. Liu, et al., · 2023
Cited alongside, same era.
“Listen, think, and understand,”
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. Glass, · 2023
Cited alongside, same era.
“Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,”
“Robust speech recognition via large-scale weak supervision,”
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, · 2023
Closest in time.
“Gpt-4 technical report,”
OpenAI, · 2023
Closest in time.
“Language is not all you need: Aligning perception with language models,”
S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, Q. Liu, et al., · 2023
Closest in time.
“Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,”
J. Li, D. Li, S. Savarese, and S. Hoi, · 2023
Closest in time.
“Palm-e: An embodied multimodal language model,”
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al., · 2023
Closest in time.
“Visual instruction tuning,”
H. Liu, C. Li, Q. Wu, and Y. J. Lee, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Zhang, S. Li, X. Zhang, J. Zhan, P. Wang, Y. Zhou, and X. Qiu, · 2023
Cited alongside, same era.
“Pengi: An audio language model for audio tasks,”
S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, · 2023
Cited alongside, same era.
“Audiopalm: A large language model that can speak and listen,”
P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. d. C. Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov, et al., · 2023
Cited alongside, same era.
“Prompting large language models with speech recognition abilities,”
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli, et al., · 2023
Cited alongside, same era.
“Speechprompt v2: Prompt tuning for speech classification tasks,”
K.-W. Chang, Y.-K. Wang, H. Shen, I.-t. Kang, W.-C. Tseng, S.-W. Li, and H.-y. Lee, · 2023
Closest in time.
“Prompting the hidden talent of web-scale speech models for zero-shot task generalization,”
P. Peng, B. Yan, S. Watanabe, and D. Harwath, · 2023
Closest in time.
“Llama: Open and efficient foundation language models,”
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al., · 2023
Closest in time.