Fetching the paper…
Reading the bibliography…
We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models.
“A neural probabilistic language model,”
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent, · 2000
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“From wer and ril to mer and wil: improved evaluation measures for connected speech recognition,”
Andrew Cameron Morris, Viktoria Maier, and Phil Green, · 2004
Earlier work this paper cites.
“A call for clarity in reporting bleu scores,”
Matt Post, · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Earlier work this paper cites.
“Parameter-efficient transfer learning for nlp.,”
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly, · 2019
Earlier work this paper cites.
“Natural questions: a benchmark for question answering research,”
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al., · 2019
Earlier work this paper cites.
“Non-attentive tacotron: Robust and controllable neural tts synthesis including unsupervised duration modeling,”
Jonathan Shen, Ye Jia, Mike Chrzanowski, Yu Zhang, Isaac Elias, Heiga Zen, and Yonghui Wu, · 2020
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer,”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu, · 2020
Earlier work this paper cites.
“mt5: A massively multilingual pre-trained text-to-text transformer,”
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel, · 2020
Earlier work this paper cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Earlier work this paper cites.
“CoVoST 2 and Massively Multilingual Speech Translation,”
Changhan Wang, Anne Wu, Jiatao Gu, and Juan Pino, · 2021
Earlier work this paper cites.
“Png BERT: augmented BERT on phonemes and graphemes for neural TTS,”
Ye Jia, Heiga Zen, Jonathan Shen, Yu Zhang, and Yonghui Wu, · 2021
Earlier work this paper cites.
“Speechstew: Simply mix all available speech recognition data to train one large neural network,”
William Chan, Daniel Park, Chris Lee, Yu Zhang, Quoc Le, and Mohammad Norouzi, · 2021
Earlier work this paper cites.
“Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux, · 2021
Cited alongside, same era.
“Lora: Low-rank adaptation of large language models,”
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen, · 2021
Cited alongside, same era.
“Self-supervised learning with random-projection quantizer for speech recognition,”
Chung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu, and Yonghui Wu, · 2022
Cited alongside, same era.
“Crosslingual generalization through multitask finetuning,”
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al., · 2022
Cited alongside, same era.
“Fleurs: Few-shot learning evaluation of universal representations of speech,”
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna, · 2022
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2023
Closest in time.
“Speech-to-text adapter and speech-to-entity retriever augmented llms for speech understanding,”
Mingqiu Wang, Izhak Shafran, Hagen Soltau, Wei Han, Yuan Cao, Dian Yu, and Laurent El Shafey, · 2023
Closest in time.
“Speech aware dialog system technology challenge (dstc11),”
Hagen Soltau, Izhak Shafran, Mingqiu Wang, Abhinav Rastogi, Jeffrey Zhao, Ye Jia, Wei Han, Yuan Cao, and Aramys Miranda, · 2023
Closest in time.
“Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,”
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu, · 2023
Closest in time.
“Audiopalm: A large language model that can speak and listen,”
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al., · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“mslam: Massively multilingual joint pre-training for speech and text,”
Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau, · 2022
Cited alongside, same era.
“Maestro: Matched speech text representations through modality matching,”
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro Moreno, Ankur Bapna, and Heiga Zen, · 2022
Cited alongside, same era.
“Flamingo: a visual language model for few-shot learning,”
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al., · 2022
Cited alongside, same era.
“Overcoming catastrophic forgetting in zero-shot cross-lingual generation,”
Tu Vu, Aditya Barua, Brian Lester, Daniel Cer, Mohit Iyyer, and Noah Constant, · 2022
Cited alongside, same era.
“Scaling instruction-finetuned language models,”
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al., · 2022
Cited alongside, same era.
“Gpt-4 technical report,” 2023
OpenAI, · 2023
Cited alongside, same era.
“Palm 2 technical report,”
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al., · 2023
Cited alongside, same era.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al., · 2023
Closest in time.
“Pengi: An audio language model for audio tasks,”
Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang, · 2023
Closest in time.
“Imagebind: One embedding space to bind them all,”
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra, · 2023
Closest in time.
“Listen, think, and understand,”
Yuan Gong, Hongyin Luo, Alexander H Liu, Leonid Karlinsky, and James Glass, · 2023
Closest in time.
“Audiotoken: Adaptation of text-conditioned diffusion models for audio-to-image generation,”
Guy Yariv, Itai Gat, Lior Wolf, Yossi Adi, and Idan Schwartz, · 2023
Closest in time.
“Stanford alpaca: An instruction-following llama model,” 2023
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto, · 2023
Closest in time.
“Nam+: Towards scalable end-to-end contextual biasing for adaptive asr,”
Tsendsuren Munkhdalai, Zelin Wu, Golan Pundak, Khe Chai Sim, Jiayang Li, Pat Rondon, and Tara N Sainath, · 2023
Closest in time.
“Mu2 slam: Multitask, multilingual speech and language models,”
Yong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey, and Ankur Bapna, · 2023
Closest in time.