Fetching the paper…
Reading the bibliography…
We present a novel Speech Augmented Language Model (SALM) with {\em multitask} and {\em in-context} learning capabilities.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Contextual speech recognition in end-to-end neural network systems using beam search.,”
Ian Williams, Anjuli Kannan, Petar S Aleksic, et al., · 2018
Earlier work this paper cites.
“Deep context: end-to-end contextual speech recognition,”
Golan Pundak, Tara N Sainath, Rohit Prabhavalkar, Anjuli Kannan, and Ding Zhao, · 2018
Earlier work this paper cites.
“On the alignment problem in multi-head attention-based neural machine translation,”
Tamer Alkhouli, Gabriel Bretschner, and Hermann Ney, · 2018
Earlier work this paper cites.
“Megatron-lm: Training multi-billion parameter language models using model parallelism,”
Mohammad Shoeybi et al., · 2019
Earlier work this paper cites.
“NeMo: a toolkit for building ai applications using neural modules,”
Oleksii Kuchaiev, Jason Li, Huyen Nguyen, et al., · 2019
Earlier work this paper cites.
“The curious case of neural text degeneration,”
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi, · 2019
Earlier work this paper cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati et al., · 2020
Earlier work this paper cites.
“Accelerating rnn transducer inference via adaptive expansion search,”
Juntae Kim, Yoonhan Lee, and Eesung Kim, · 2020
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models,”
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen, · 2021
Earlier work this paper cites.
“Finetuned language models are zero-shot learners,”
Jason Wei et al., · 2021
Earlier work this paper cites.
“Must-c: A multilingual corpus for end-to-end speech translation,”
Roldano Cattoni, Mattia Antonino Di Gangi, Luisa Bentivogli, Matteo Negri, and Marco Turchi, · 2021
Earlier work this paper cites.
“Integrating text inputs for training and adapting rnn transducer asr models,”
Samuel Thomas et al., · 2022
Earlier work this paper cites.
“Maestro: Matched speech text representations through modality matching,”
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, et al., · 2022
Cited alongside, same era.
“Speechlm: Enhanced speech pre-training with unpaired textual data,”
Ziqiang Zhang, Sanyuan Chen, Long Zhou, et al., · 2022
Cited alongside, same era.
“A survey for in-context learning,”
Qingxiu Dong, Lei Li, Damai Dai, et al., · 2022
Cited alongside, same era.
“The stack: 3 tb of permissively licensed source code,”
Denis Kocetkov et al., · 2022
Cited alongside, same era.
“Gpt-4 technical report,”
OpenAI, · 2023
Cited alongside, same era.
“Listen, think, and understand,”
Yuan Gong, Hongyin Luo, Alexander H Liu, et al., · 2023
Closest in time.
“Adapting large language model with speech for fully formatted end-to-end speech recognition,”
Shaoshi Ling, Yuxuan Hu, Shuangbei Qian, et al., · 2023
Closest in time.
“On decoder-only architecture for speech-to-text and large language model integration,”
Jian Wu, Yashesh Gaur, Zhuo Chen, et al., · 2023
Closest in time.
“Prompting large language models with speech recognition abilities,”
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, et al., · 2023
Closest in time.
“Speech-to-text adapter and speech-to-entity retriever augmented llms for speech understanding,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rohan Anil, Andrew M Dai, Orhan Firat, et al., · 2023
Cited alongside, same era.
“Audiogpt: Understanding and generating speech, music, sound, and talking head,”
Rongjie Huang, Mingze Li, Dongchao Yang, et al., · 2023
Cited alongside, same era.
Feilong Chen et al., · 2023
Cited alongside, same era.
“Large-scale language model rescoring on long-form data,”
Tongzhou Chen, Cyril Allauzen, Yinghui Huang, et al., · 2023
Cited alongside, same era.
Rao Ma et al., · 2023
Cited alongside, same era.
“Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,”
Dong Zhang, Shimin Li, Xin Zhang, et al., · 2023
Cited alongside, same era.
“Audiopalm: A large language model that can speak and listen,”
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, et al., · 2023
Cited alongside, same era.
Mingqiu Wang, Izhak Shafran, Hagen Soltau, et al., · 2023
Closest in time.
“Can contextual biasing remain effective with whisper and gpt-2?,”
Guangzhi Sun, Xianrui Zheng, Chao Zhang, and Philip C Woodland, · 2023
Closest in time.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang et al., · 2023
Closest in time.
“Voicebox: Text-guided multilingual universal speech generation at scale,”
Matthew Le et al., · 2023
Closest in time.
“Google usm: Scaling automatic speech recognition beyond 100 languages,”
Yu Zhang, Wei Han, James Qin, et al., · 2023
Closest in time.
“Seamlessm4t-massively multilingual & multimodal machine translation,”
Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, et al., · 2023
Closest in time.
“Fast conformer with linearly scalable attention for efficient speech recognition,”
Dima Rekesh, Samuel Kriman, Somshubra Majumdar, Vahid Noroozi, He Juang, Oleksii Hrinchuk, Ankur Kumar, and Boris Ginsburg, · 2023
Closest in time.
“Findings of the IWSLT 2023 Evaluation Campaign,”
Milind Agarwal et al., · 2023
Closest in time.