Fetching the paper…
Reading the bibliography…
The recent wave of audio foundation models (FMs) could provide new capabilities for conversational modeling.
Switchboard: telephone speech corpus for research and development
J.J. Godfrey, E.C. Holliman, and J. McDaniel · 1992
Earlier work this paper cites.
The fisher corpus: A resource for the next generations of speech-to-text
Christopher Cieri, David Miller, and Kevin Walker · 2004
Earlier work this paper cites.
Back-channel feedback generation using linguistic and nonlinguistic information and its application to spoken dialogue system
Shinya Fujie, Kenta Fukushima, and Tetsunori Kobayashi · 2005
Earlier work this paper cites.
On temporal aspects of turn taking in conversational dialogues
Louis Ten Bosch, Nelleke Oostdijk, and Lou Boves · 2005
Earlier work this paper cites.
Hkust/mts: A very large scale mandarin telephone speech corpus
Yi Liu, Pascale Fung, Yongsheng Yang, Christopher Cieri, Shudong Huang, and David Graff · 2006
Earlier work this paper cites.
Doing research on a deployed spoken dialogue system: one year of let’s go! experience
Antoine Raux, Dan Bohus, Brian Langner, Alan W Black, and Maxine Eskenazi · 2006
Earlier work this paper cites.
An analysis of multimodal cues of interruption in dyadic spoken interactions
Chi-Chun Lee, Sungbok Lee, and Shrikanth S. Narayanan · 2008
Earlier work this paper cites.
Turn-taking cues in task-oriented dialogue
Agustín Gravano and Julia Hirschberg · 2010
Earlier work this paper cites.
Predicting interruptions in dyadic spoken interactions
Chi-Chun Lee and Shrikanth S. Narayanan · 2010
Earlier work this paper cites.
Modeling human communication dynamics [social sciences]
Louis-Philippe Morency · 2010
Earlier work this paper cites.
A probabilistic multimodal approach for predicting listener backchannels
Louis-Philippe Morency, Iwan de Kok, and Jonathan Gratch · 2010
Earlier work this paper cites.
Margin of Error , pp. 765–765
Judith M. Tanur · 2011
Earlier work this paper cites.
A corpus-based study of interruptions in spoken dialogue
Agustín Gravano and Julia Hirschberg · 2012
Earlier work this paper cites.
Investigating the influence of pause fillers for automatic backchannel prediction
Stefan Scherer, Derya Ozkan, and Louis-Philippe Morency · 2012
Earlier work this paper cites.
Acoustic modeling for google home
Bo Li, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Haşim Sak, Golan Pundak, Kean Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim, Olivier Siohan, Mitchel Weintraub, Erik McDermott, Richard Rose, and Matt Shannon · 2017
Earlier work this paper cites.
Prediction of turn-taking using multitask learning with prediction of backchannels and fillers
Kohei Hara, Koji Inoue, Katsuya Takanashi, and Tatsuya Kawahara · 2018
Earlier work this paper cites.
Neural dialogue context online end-of-turn detection
Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Ryuichiro Higashinaka, and Yushi Aono · 2018
Earlier work this paper cites.
Hyperparameter importance across datasets
Jan N. van Rijn and Frank Hutter · 2018
Earlier work this paper cites.
End-to-end neural speaker diarization with self-attention
Yusuke Fujita, Naoyuki Kanda, Shota Horiguchi, Yawen Xue, Kenji Nagamatsu, and Shinji Watanabe · 2019
Cited alongside, same era.
Pyannote. audio: neural building blocks for speaker diarization
Hervé Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, and Marie-Philippe Gill · 2020
Cited alongside, same era.
Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog
Erik Ekstedt and Gabriel Skantze · 2020
Cited alongside, same era.
Can prediction of turn-management willingness improve turn-changing modeling?
Ryo Ishii, Xutong Ren, Michal Muszynski, and Louis-Philippe Morency · 2020
Cited alongside, same era.
The future of sensitivity analysis: An essential discipline for systems modeling and policy support
Saman Razavi, Anthony Jakeman, Andrea Saltelli, Clémentine Prieur, Bertrand Iooss, Emanuele Borgonovo, Elmar Plischke, Samuele Lo Piano, Takuya Iwanaga, William Becker, Stefano Tarantola, Joseph H.A. Guillaume, John Jakeman, Hoshin Gupta, Nicola Melillo, Giovanni Rabitti, Vincent Chabridon, Qingyun Duan, Xifu Sun, Stefán Smith, Razi Sheikholeslami, Nasim Hosseini, Masoud Asadzadeh, Arnald Puy, Sergei Kucherenko, and Holger R. Maier · 2020
Chien-yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Haibin Wu, Siddhant Arora, Kai-Wei Chang, Jiatong Shi, Yifan Peng, Roshan S. Sharma, Shinji Watanabe, Bhiksha Ramakrishnan, Shady Shehata, and Hung-yi Lee · 2023
Later among the works it cites.
Generative spoken dialogue language modeling
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoît Sagot, Abdelrahman Mohamed, and Emmanuel Dupoux · 2023
Later among the works it cites.
Augmenting transformer-transducer based speaker change detection with token-level training loss
Guanlong Zhao, Quan Wang, Han Lu, Yiling Huang, and Ignacio López-Moreno · 2023
Later among the works it cites.
Smollm - blazingly fast and remarkably powerful, 2024
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Leandro von Werra, and Thomas Wolf · 2024
Later among the works it cites.
Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Survey on evaluation methods for dialogue systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak · 2021
Cited alongside, same era.
Multimodal and multitask approach to listener’s backchannel prediction: Can prediction of turn-changing and turn-management willingness improve backchannel modeling?
Ryo Ishii, Xutong Ren, Michal Muszynski, and Louis-Philippe Morency · 2021
Cited alongside, same era.
Turn-taking in conversational systems and human-robot interaction: a review
Gabriel Skantze · 2021
Cited alongside, same era.
Superb: Speech processing universal performance benchmark
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al · 2021
Cited alongside, same era.
Voice activity projection: Self-supervised learning of turn-taking events
Erik Ekstedt and Gabriel Skantze · 2022
Cited alongside, same era.
How much does prosody help turn-taking? investigations using voice activity projection models
Erik Ekstedt and Gabriel Skantze · 2022
Cited alongside, same era.
Collection and analysis of travel agency task dialogues with age-diverse speakers
Michimasa Inaba, Yuya Chiba, Ryuichiro Higashinaka, Kazunori Komatani, Yusuke Miyao, and Takayuki Nagai · 2022
Cited alongside, same era.
Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang, Lu Lu, Yuxuan Wang, Haizhou Li, and Zhizheng Wu · 2024
Later among the works it cites.
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, Chang Zhou, and Jingren Zhou · 2024
Later among the works it cites.
Moshi: a speech-text foundation model for real-time dialogue
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour · 2024
Later among the works it cites.
Multilingual turn-taking prediction using voice activity projection
Koji Inoue, Bing’er Jiang, Erik Ekstedt, Tatsuya Kawahara, and Gabriel Skantze · 2024
Later among the works it cites.
Language model can listen while speaking, 2024
Ziyang Ma, Yakun Song, Chenpeng Du, Jian Cong, Zhuo Chen, Yuping Wang, Yuxuan Wang, and Xie Chen · 2024
Later among the works it cites.
Zahra Sadeghi and Stan Matwin · 2024
Later among the works it cites.
SALMONN: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA, and Chao Zhang · 2024
Later among the works it cites.
Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier
Silero Team · 2024
Later among the works it cites.
Turn-taking and backchannel prediction with acoustic and large language model fusion
Jinhan Wang, Long Chen, Aparna Khare, Anirudh Raju, Pranav Dheram, Di He, Minhua Wu, Andreas Stolcke, and Venkatesh Ravichandran · 2024
Later among the works it cites.
Mini-omni: Language models can hear, talk while thinking in streaming, 2024
Zhifei Xie and Changqiao Wu · 2024
Later among the works it cites.
Air-bench: Benchmarking large audio-language models via generative comprehension
Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, and Jingren Zhou · 2024
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Later among the works it cites.
Automatic evaluation of turn-taking cues in conversational speech synthesis
Erik Ekstedt, Siyang Wang, Éva Székely, Joakim Gustafson, and Gabriel Skantze · 2064
Closest in time.