Fetching the paper…
Reading the bibliography…
We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM).
“A simplest systematics for the organization of turn taking for conversation,”
Harvey Sacks, Emanuel A. Schegloff, and Gail Jefferson, · 1978
Earlier work this paper cites.
“Switchboard: Telephone speech corpus for research and development,”
John J Godfrey, Edward C. Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“Observations on overlap: findings and implications for automatic processing of multi-party conversation,”
Elizabeth Shriberg, Andreas Stolcke, and Don Baron, · 2001
Earlier work this paper cites.
“Dialog prediction for a general model of turn-taking,”
Nigel G. Ward, Olac Fuentes, and Alejandro Vega, · 2010
Earlier work this paper cites.
“Virtual agents as daily assistants for elderly or cognitively impaired people: Studies on acceptance and interaction feasibility,”
Ramin Yaghoubzadeh, Marcel Kramer, Karola Pitsch, and Stefan Kopp, · 2013
Earlier work this paper cites.
“Study of a home robot: Jibo,”
Pranav Rane, Varun Mhatre, and Lakshmi Kurup, · 2014
Earlier work this paper cites.
“Turn-taking, feedback and joint attention in situated human-robot interaction,”
Gabriel Skantze, Anna Hjalmarsson, and Catharine Oertel, · 2014
Earlier work this paper cites.
“Ten challenges in highly-interactive dialog system.,”
Nigel G. Ward and David DeVault, · 2015
Earlier work this paper cites.
“Using neural networks for data-driven backchannel prediction: A survey on input features and training techniques,”
Markus Mueller, David Leuschner, Lars Briem, Maria Schmidt, Kevin Kilgour, Sebastian Stueker, and Alex Waibel, · 2015
Earlier work this paper cites.
“Prediction and generation of backchannel form for attentive listening systems.,”
Tatsuya Kawahara, Takashi Yamaguchi, Koji Inoue, Katsuya Takanashi, and Nigel G Ward, · 2016
Earlier work this paper cites.
“Towards a general, continuous model of turn-taking in spoken dialogue using LSTM recurrent neural networks,”
Gabriel Skantze, · 2017
Earlier work this paper cites.
“Turn-taking estimation model based on joint embedding of lexical and prosodic contents.,”
Chaoran Liu, Carlos Toshinori Ishi, and Hiroshi Ishiguro, · 2017
Earlier work this paper cites.
“Online end-of-turn detection from speech based on stacked time-asynchronous sequential networks.,”
Ryo Masumura, Taichi Asami, Hirokazu Masataki, Ryo Ishii, and Ryuichiro Higashinaka, · 2017
Cited alongside, same era.
“Alexa, Siri, Cortana, and more: an introduction to voice assistants,”
Matthew B. Hoy, · 2018
Cited alongside, same era.
“Evaluation of real-time deep learning turn-taking models for multiple dialogue scenarios,”
Divesh Lala, Koji Inoue, and Tatsuya Kawahara, · 2018
Cited alongside, same era.
“Improving end-of-turn detection in spoken dialogues by detecting speaker intentions as a secondary task,”
Zakaria Aldeneh, Dimitrios Dimitriadis, and Emily Mower Provost, · 2018
Cited alongside, same era.
“Neural dialogue context online end-of-turn detection,”
Ryo Masumura, Tomohiro Tanaka, Atsushi Ando, Ryo Ishii, Ryuichiro Higashinaka, and Yushi Aono, · 2018
Cited alongside, same era.
“Turn-taking in conversational systems and human-robot interaction: a review,”
Gabriel Skantze, · 2021
Later among the works it cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Later among the works it cites.
“LoRA: Low-rank adaptation of large language models,”
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen, · 2021
Later among the works it cites.
“Turn-taking prediction for natural conversational speech,”
Shuo-yiin Chang, Bo Li, Tara N Sainath, Chao Zhang, Trevor Strohman, Qiao Liang, and Yanzhang He, · 2022
Later among the works it cites.
“Voice activity projection: Self-supervised learning of turn-taking events,”
Erik Ekstedt and Gabriel Skantze, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Multimodal continuous turn-taking prediction using multiscale RNNs,”
Matthew Roddy, Gabriel Skantze, and Naomi Harte, · 2018
Cited alongside, same era.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, · 2019
Cited alongside, same era.
“Oh, jeez! or uh-huh? a listener-aware backchannel predictor on ASR transcriptions,”
Daniel Ortega, Chia-Yu Li, and Ngoc Thang Vu, · 2020
Cited alongside, same era.
“TurnGPT: a transformer-based language model for predicting turn-taking in spoken dialog,”
Erik Ekstedt and Gabriel Skantze, · 2020
Cited alongside, same era.
“GPT-3: Its nature, scope, limits, and consequences,”
Luciano Floridi and Massimo Chiriatti, · 2020
Cited alongside, same era.
“Transformers: State-of-the-art natural language processing,”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush, · 2020
Cited alongside, same era.
“Gated multimodal fusion with contrastive learning for turn-taking prediction in human-robot dialogue,”
Jiudong Yang, Peiying Wang, Yi Zhu, Mingchao Feng, Meng Chen, and Xiaodong He, · 2022
Later among the works it cites.
“Jigsaw: Large language models meet program synthesis,”
Naman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan, Suresh Parthasarathy, Sriram Rajamani, and Rahul Sharma, · 2022
Later among the works it cites.
“Finetuned language models are zero-shot learners,”
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le, · 2022
Later among the works it cites.
“ChatGPT for good? On opportunities and challenges of large language models for education,”
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al., · 2023
Later among the works it cites.
“RedPajama: An open source recipe to reproduce LLaMA training dataset,” https://github.com/togethercomputer/RedPajama-Data, Apr. 2023
Together Computer, · 2023
Later among the works it cites.
“Recent advances in natural language processing via large pre-trained language models: A survey,”
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth, · 2023
Later among the works it cites.