Fetching the paper…
Reading the bibliography…
Recent advancements in large language models (LLMs) have led to significant progress in text-based dialogue systems.
Switchboard: Telephone speech corpus for research and development
John J Godfrey, Edward C Holliman, and Jane McDaniel. 1992 · 1992
Earlier work this paper cites.
Prosodic features which cue back-channel responses in english and japanese
Nigel Ward and Wataru Tsukahara. 2000 · 2000
Earlier work this paper cites.
The fisher corpus: A resource for the next generations of speech-to-text
Christopher Cieri, David Miller, and Kevin Walker. 2004 · 2004
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M Smith, et al. 2020 · 2004
Earlier work this paper cites.
Back-channel feedback generation using linguistic and nonlinguistic information and its application to spoken dialogue system
Shinya Fujie, Kenta Fukushima, and Tetsunori Kobayashi. 2005 · 2005
Earlier work this paper cites.
From reaction to prediction: Experiments with computational models of turn-taking
David Schlangen. 2006 · 2006
Earlier work this paper cites.
Towards incremental end-of-utterance detection in dialogue systems
Michaela Atterer, Timo Baumann, and David Schlangen. 2008 · 2008
Earlier work this paper cites.
Turn-taking cues in task-oriented dialogue
Agustín Gravano and Julia Hirschberg. 2011 · 2011
Earlier work this paper cites.
Finding appropriate interaction strategies for proactive dialogue systems—an open quest
Florian Nothdurft, Stefan Ultes, and Wolfgang Minker. 2014 · 2014
Earlier work this paper cites.
Prediction of who will be the next speaker and when using gaze behavior in multiparty meetings
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, and Junji Yamato. 2016 · 2016
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al. 2017 · 2017
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018 · 2018
Earlier work this paper cites.
Prediction of turn-taking using multitask learning with prediction of backchannels and fillers
Kohei Hara, Koji Inoue, Katsuya Takanashi, and Tatsuya Kawahara. 2018 · 2018
Cited alongside, same era.
Convai dataset of topic-oriented human-to-chatbot dialogues
Varvara Logacheva, Mikhail Burtsev, Valentin Malykh, Vadim Polulyakh, and Aleksandr Seliverstov. 2018 · 2018
Cited alongside, same era.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2018 · 2018
Cited alongside, same era.
Oh, jeez! or uh-huh? a listener-aware backchannel predictor on asr transcriptions
Daniel Ortega, Chia-Yu Li, and Ngoc Thang Vu. 2020 · 2020
Cited alongside, same era.
The design and implementation of xiaoice, an social chatbot
Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum. 2020 · 2020
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022 · 2022
Later among the works it cites.
Funasr: A fundamental end-to-end speech recognition toolkit. arxiv 2023
Zhifu Gao, Z Li, J Wang, H Luo, X Shi, M Chen, Y Li, L Zuo, Z Du, Z Xiao, et al · 2023
Later among the works it cites.
Topical-chat: Towards knowledge-grounded open-domain conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tur. 2023 · 2023
Later among the works it cites.
Jungil Kong, Jihoon Park, Beomjeong Kim, Jeongmin Kim, Dohee Kong, and Sangjin Kim. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Multimodal and multitask approach to listener’s backchannel prediction: Can prediction of turn-changing and turn-management willingness improve backchannel modeling?
Ryo Ishii, Xutong Ren, Michal Muszynski, and Louis-Philippe Morency. 2021 · 2021
Cited alongside, same era.
Bpm_mt: Enhanced backchannel prediction model using multi-task learning
Jin Yea Jang, San Kim, Minyoung Jung, Saim Shin, and Gahgene Gweon. 2021 · 2021
Cited alongside, same era.
Duplex conversation in outbound agent system
Chunxiang Jin, Minghui Yang, and Zujie Wen. 2021 · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al. 2021 · 2021
Cited alongside, same era.
How much does prosody help turn-taking? investigations using voice activity projection models
Erik Ekstedt and Gabriel Skantze. 2022 · 2022
Cited alongside, same era.
Duplex conversation: Towards human-like interaction in spoken dialogue systems
Ting-En Lin, Yuchuan Wu, Fei Huang, Luo Si, Jian Sun, and Yongbin Li. 2022 · 2022
Cited alongside, same era.
Generative spoken dialogue language modeling
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al. 2023 · 2023
Later among the works it cites.
Slam-omni: Timbre-controllable voice interaction system with single-stage training
Wenxi Chen, Ziyang Ma, Ruiqi Yan, Yuzhe Liang, Xiquan Li, Ruiyang Xu, Zhikang Niu, Yanqiao Zhu, Yifan Yang, Zhanxun Liu, et al. 2024 · 2024
Later among the works it cites.
Moshi: a speech-text foundation model for real-time dialogue
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. 2024 · 2024
Later among the works it cites.
Parrot: Autoregressive spoken dialogue language modeling with decoder-only transformers
Ziqiao Meng, Qichao Wang, Wenqian Cui, Yifei Zhang, Bingzhe Wu, Irwin King, Liang Chen, and Peilin Zhao · 2024
Later among the works it cites.
A full-duplex speech dialogue scheme based on large language models
Peng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan, Wei Xia, and Yuanjun Xiong. 2024 · 2024
Later among the works it cites.
Mini-omni2: Towards open-source gpt-4o with vision, speech and duplex capabilities
Zhifei Xie and Changqiao Wu. 2024 · 2024
Later among the works it cites.
Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot
Aohan Zeng, Zhengxiao Du, Mingdao Liu, Kedong Wang, Shengmin Jiang, Lei Zhao, Yuxiao Dong, and Jie Tang. 2024 · 2024
Later among the works it cites.
Omniflatten: An end-to-end gpt model for seamless voice conversation
Qinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen, Wen Wang, Siqi Zheng, Jiaqing Liu, Hai Yu, and Chaohong Tan. 2024 · 2024
Later among the works it cites.