Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction”
Rohit Paturi, Sundararajan Srinivasan and Xiang Li · 1982
Earlier work this paper cites.
“2000 HUB5 English Evaluation Speech LDC2002S09”
et Cieri Christopher · 2002
Earlier work this paper cites.
“Fisher English Training Speech Part 1 Speech LDC2004S13”
et Cieri Christopher · 2004
Earlier work this paper cites.
“Europarl: A parallel corpus for statistical machine translation”
Philipp Koehn · 2005
Earlier work this paper cites.
“Fisher English Training Part 2, Speech LDC2005S13”
et Cieri Christopher · 2005
Earlier work this paper cites.
“Vox Populi: Collecting High-Quality Labels from a Crowd.”
Ofer Dekel and Ohad Shamir · 2009
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books”
Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
“Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings”
Reza Lotfian and Carlos Busso · 2017
Earlier work this paper cites.
Lukasz Kaiser et al · 2017
Earlier work this paper cites.
“Building Naturalistic Emotionally Balanced Speech Corpus by Retrieving Emotional Speech From Existing Podcast Recordings”
R. Lotfian and C. Busso · 2017
Earlier work this paper cites.
“Improving language understanding by generative pre-training”
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“Common voice: A massively-multilingual speech corpus”
Rosana Ardila et al · 2019
Earlier work this paper cites.
“SLURP: A spoken language understanding resource package”
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski and Verena Rieser · 2020
Earlier work this paper cites.
“Covost 2 and massively multilingual speech-to-text translation”
Changhan Wang, Anne Wu and Juan Pino · 2020
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer”
Colin Raffel et al · 2020
Earlier work this paper cites.
“Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates”
J. Iranzo-Sánchez et al · 2020
Earlier work this paper cites.
“SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing”
Junyi Ao et al · 2021
Earlier work this paper cites.
“LoRA: Low-Rank Adaptation of Large Language Models”, 2021
Edward. Hu et al · 2021
Earlier work this paper cites.
“WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing”
Sanyuan Chen et al · 2021
Cited alongside, same era.
“Multilingual Speech Translation from Efficient Finetuning of Pretrained Models”
Xian Li et al · 2021
Cited alongside, same era.
“Speechnet: A universal modularized model for speech processing tasks”
Yi-Chen Chen et al · 2021
Cited alongside, same era.
“Scaling instruction-finetuned language models”
Hyung Chung et al · 2022
Cited alongside, same era.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Cited alongside, same era.
“Can generative large language models perform asr error correction?”
Rao Ma et al · 2023
Later among the works it cites.
“AudioPaLM: A Large Language Model That Can Speak and Listen”
Paul Rubenstein et al · 2023
Later among the works it cites.
“VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation”
Tianrui Wang et al · 2023
Later among the works it cites.
“Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models”
Yunfei Chu et al · 2023
Later among the works it cites.
“SLM: Bridge the thin gap between speech and text foundation models”
Mingqiu Wang et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chung-Cheng Chiu et al · 2022
Cited alongside, same era.
“mslam: Massively multilingual joint pre-training for speech and text”
Ankur Bapna et al · 2022
Cited alongside, same era.
“Integration of pre-trained networks with continuous token interface for end-to-end spoken language understanding”
Seunghyun Seo, Donghyun Kwak and Bowon Lee · 2022
Cited alongside, same era.
Yingzhi Wang, Abdelmoumene Boumadane and Abdelwahab Heba · 2022
Cited alongside, same era.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei et al · 2022
Cited alongside, same era.
“Flamingo: a visual language model for few-shot learning”
Jean-Baptiste Alayrac et al · 2022
Cited alongside, same era.
“Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation”
Junnan Li, Dongxu Li, Caiming Xiong and Steven Hoi · 2022
Cited alongside, same era.
Later among the works it cites.
“Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities”
Dong Zhang et al · 2023
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision”
Alec Radford et al · 2023
Later among the works it cites.
“Stanford alpaca: An instruction-following llama model”, 2023
Rohan Taori et al · 2023
Later among the works it cites.
“SeamlessM4T: Massively Multilingual & Multimodal Machine Translation”, 2023
Seamless Communication et al · 2023
Later among the works it cites.
“Testing Speech Emotion Recognition Machine Learning Models”
Anna Derington et al · 2023
Later among the works it cites.
“Efficient Guided Generation for LLMs”
Brandon Willard and Rémi Louf · 2023
Later among the works it cites.
“Grounding language models to images for multimodal inputs and outputs”
Jing Koh, Ruslan Salakhutdinov and Daniel Fried · 2023
Later among the works it cites.
“Kosmos-2: Grounding Multimodal Large Language Models to the World”
Zhiliang Peng et al · 2023
Later among the works it cites.
“Pengi: An Audio Language Model for Audio Tasks”
Soham Deshmukh, Benjamin Elizalde, Rita Singh and Huaming Wang · 2023
Later among the works it cites.
“Listen, Think, and Understand”
Yuan Gong et al · 2023
Later among the works it cites.
“Llasm: Large language and speech model”
Yu Shu et al · 2023
Later among the works it cites.
“Large Language Model based Multi-Agents: A Survey of Progress and Challenges”
Taicheng Guo et al · 2024
Closest in time.
W Huang et al · 2024
Closest in time.