Fetching the paper…
Reading the bibliography…
Full-duplex multimodal large language models (LLMs) provide a unified framework for addressing diverse speech understanding and generation tasks, enabling more natural and seamless human-machine conversations.
“POMDP-based statistical spoken dialog systems: A review”
Steve Young, Milica Gašić, Blaise Thomson and Jason. Williams · 2013
Earlier work this paper cites.
“A neural conversational model”
Oriol Vinyals and Quoc. Le · 2015
Earlier work this paper cites.
“Building end-to-end dialogue systems using generative hierarchical neural network models”
Iulian. Serban et al · 2016
Earlier work this paper cites.
“A network-based end-to-end trainable task-oriented dialogue system”
Tsung-Hsien Wen et al · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Language models are few-shot learners”
Tom Brown et al · 2020
Earlier work this paper cites.
“A simple language model for task-oriented dialogue”
Ehsan Hosseini-Asl et al · 2020
Earlier work this paper cites.
“Understanding and predicting user dissatisfaction in a neural generative chatbot”
Abigail See and Christopher Manning · 2021
Earlier work this paper cites.
“A conformer-based ASR frontend for joint acoustic echo cancellation, speech enhancement and speech separation”
Tom O’Malley et al · 2021
Earlier work this paper cites.
“GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio”
Guoguo Chen et al · 2021
Earlier work this paper cites.
“Turn-taking prediction for natural conversational speech”
Shuo-yiin Chang et al · 2022
Cited alongside, same era.
“Streaming intended query detection using E2E modeling for continued conversation”
Shuo-yiin Chang et al · 2022
Cited alongside, same era.
“Contextual acoustic barge in classification for spoken dialog systems”
Dhanush Bekal et al · 2022
Cited alongside, same era.
“LLaMA: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Cited alongside, same era.
“Boosting large language model for speech synthesis: An empirical study”
Hongkun Hao et al · 2023
Cited alongside, same era.
Abhimanyu Dubey et al · 2024
Closest in time.
“Spirit-LM: Interleaved spoken and written language model”
Tu Nguyen et al · 2024
Closest in time.
“Llama-omni: Seamless speech interaction with large language models”
Qingkai Fang et al · 2024
Closest in time.
“Mini-omni: Language models can hear, talk while thinking in streaming”
Zhifei Xie and Changqiao Wu · 2024
Closest in time.
“Beyond the turn-based game: Enabling real-time conversations with duplex models”
Xinrong Zhang et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dong Zhang et al · 2023
Cited alongside, same era.
Albert. Jiang et al · 2024
Cited alongside, same era.
“SALMONN: Towards generic hearing abilities for large language models”
Changli Tang et al · 2024
Cited alongside, same era.
“Listen, think, and understand”
Yuan Gong et al · 2024
Cited alongside, same era.
“Qwen2-Audio technical report”
Yunfei Chu et al · 2024
Cited alongside, same era.
“Turn-taking and backchannel prediction with acoustic and large language model fusion”
Jinhan Wang et al · 2024
Closest in time.
“Language model can listen while speaking”
Ziyang Ma et al · 2024
Closest in time.
“Moshi: A speech-text foundation model for real-time dialogue”
Alexandre Défossez et al · 2024
Closest in time.
“Beyond turn-based interfaces: Synchronous LLMs as full-duplex dialogue agents”
Bandhav Veluri et al · 2024
Closest in time.
“Libriheavy: A 50,000 hours ASR corpus with punctuation casing and context”
Wei Kang et al · 2024
Closest in time.