Fetching the paper…
Reading the bibliography…
Full-Duplex Speech Dialogue Systems (Full-Duplex SDS) have significantly enhanced the naturalness of human-machine interaction by enabling real-time bidirectional communication.
The fisher corpus: A resource for the next generations of speech-to-text
Christopher Cieri, David Miller, and Kevin Walker · 2004
Earlier work this paper cites.
Efficient voice activity detection algorithms using long-term speech information
Javier Ramırez, José C Segura, Carmen Benıtez, Angel De La Torre, and Antonio Rubio · 2004
Earlier work this paper cites.
Object-based auditory and visual attention
Barbara G Shinn-Cunningham · 2008
Earlier work this paper cites.
Statistical voice activity detection based on sparse representation over learned dictionary
Shi-Wen Deng and Ji-Qing Han · 2013
Earlier work this paper cites.
Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog
Erik Ekstedt and Gabriel Skantze · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Generative spoken language modeling from raw audio, 2021
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, and Emmanuel Dupoux · 2021
Earlier work this paper cites.
Turn-taking in conversational systems and human-robot interaction: a review
Gabriel Skantze · 2021
Earlier work this paper cites.
Duplex conversation: Towards human-like interaction in spoken dialogue systems
Ting-En Lin, Yuchuan Wu, Fei Huang, Luo Si, Jian Sun, and Yongbin Li · 2022
Earlier work this paper cites.
Using voice activity detection and deep neural networks with hybrid speech feature extraction for deceptive speech detection
Serban Mihalache and Dragos Burileanu · 2022
Earlier work this paper cites.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al · 2023
Earlier work this paper cites.
Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities, 2023
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu · 2023
Cited alongside, same era.
Speechverse: A large-scale generalizable audio language model
Nilaksh Das, Saket Dingliwal, Srikanth Ronanki, Rohit Paturi, Zhaocheng Huang, Prashant Mathur, Jie Yuan, Dhanush Bekal, Xing Niu, Sai Muralidhar Jayanthi, et al · 2024
Cited alongside, same era.
Mini-omni2: Towards open-source gpt-4o with vision, speech and duplex capabilities
Zhifei Xie and Changqiao Wu · 2024
Cited alongside, same era.
Audiogpt: Understanding and generating speech, music, sound, and talking head
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al · 2024
Cited alongside, same era.
Speechgpt-gen: Scaling chain-of-information speech generation
Lauragpt: Listen, attend, understand, and regenerate audio with gpt, 2024
Zhihao Du, Jiaming Wang, Qian Chen, Yunfei Chu, Zhifu Gao, Zerui Li, Kai Hu, Xiaohuan Zhou, Jin Xu, Ziyang Ma, Wen Wang, Siqi Zheng, Chang Zhou, Zhijie Yan, and Shiliang Zhang · 2024
Later among the works it cites.
Language model can listen while speaking, 2024
Ziyang Ma, Yakun Song, Chenpeng Du, Jian Cong, Zhuo Chen, Yuping Wang, Yuxuan Wang, and Xie Chen · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dong Zhang, Xin Zhang, Jun Zhan, Shimin Li, Yaqian Zhou, and Xipeng Qiu · 2024
Cited alongside, same era.
Wavchat: A survey of spoken dialogue models, 2024
Shengpeng Ji, Yifu Chen, Minghui Fang, Jialong Zuo, Jingyu Lu, Hanting Wang, Ziyue Jiang, Long Zhou, Shujie Liu, Xize Cheng, Xiaoda Yang, Zehan Wang, Qian Yang, Jian Li, Yidi Jiang, Jingzhen He, Yunfei Chu, Jin Xu, and Zhou Zhao · 2024
Cited alongside, same era.
Moshi: a speech-text foundation model for real-time dialogue, 2024
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour · 2024
Cited alongside, same era.
A full-duplex speech dialogue scheme based on large language models, 2024
Peng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan, Wei Xia, and Yuanjun Xiong · 2024
Cited alongside, same era.
Beyond turn-based interfaces: Synchronous llms as full-duplex dialogue agents, 2024
Bandhav Veluri, Benjamin N Peloquin, Bokai Yu, Hongyu Gong, and Shyamnath Gollakota · 2024
Cited alongside, same era.
Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen llm, 2024
Xiong Wang, Yangze Li, Chaoyou Fu, Yunhang Shen, Lei Xie, Ke Li, Xing Sun, and Long Ma · 2024
Cited alongside, same era.
Vita: Towards open-source interactive omni multimodal llm, 2024
Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Meng Zhao, Yifan Zhang, Shaoqi Dong, Xiong Wang, Di Yin, Long Ma, Xiawu Zheng, Ran He, Rongrong Ji, Yunsheng Wu, Caifeng Shan, and Xing Sun · 2024
Cited alongside, same era.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot, 2024
Aohan Zeng, Zhengxiao Du, Mingdao Liu, Kedong Wang, Shengmin Jiang, Lei Zhao, Yuxiao Dong, and Jie Tang · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone, 2024
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2024
Later among the works it cites.
Minmo: A multimodal large language model for seamless voice interaction, 2025
Qian Chen, Yafeng Chen, Yanni Chen, Mengzhe Chen, Yingda Chen, Chong Deng, Zhihao Du, and et al · 2025
Closest in time.
Omniflatten: An end-to-end gpt model for seamless voice conversation, 2025
Qinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen, Wen Wang, Siqi Zheng, Jiaqing Liu, Hai Yu, Chaohong Tan, Zhihao Du, and Shiliang Zhang · 2025
Closest in time.
Real-time textless dialogue generation, 2025
Long Mai and Julie Carson-Berndsen · 2025
Closest in time.
Voice activity detection (vad) - openai api
OpenAI · 2025
Closest in time.