Fetching the paper…
Reading the bibliography…
The advent of large language models (LLMs) has made it possible to generate natural written dialogues between two agents.
Some methods for classification and analysis of multivariate observations
James MacQueen · 1967
Earlier work this paper cites.
On getting a word in edgewise
Victor H Yngve · 1970
Earlier work this paper cites.
Laughter and dialogue: The social significance of laughter in institutional discourse
Viveka Adelswärd · 1989
Earlier work this paper cites.
Corpus of spontaneous Japanese: Its design and evaluation
Kikuo Maekawa · 2003
Earlier work this paper cites.
K-means++ the advantages of careful seeding
David Arthur and Sergei Vassilvitskii · 2007
Earlier work this paper cites.
Universals and cultural variation in turn-taking in conversation
Tanya Stivers, Nicholas J. Enfield, Penelope Brown, Christina Englert, Makoto Hayashi, Trine Heinemann, Gertie Hoymann, Federico Rossano, Jan Peter De Ruiter, Kyung-Eun Yoon, and Stephen C. Levinson · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Prediction and generation of backchannel form for attentive listening systems
Tatsuya Kawahara, Takashi Yamaguchi, Koji Inoue, Katsuya Takanashi, and Nigel G Ward · 2016
Earlier work this paper cites.
WORLD: A vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Earlier work this paper cites.
Attentive listening system with backchanneling, response generation and flexible turn-taking
Divesh Lala, Pierrick Milhorat, Koji Inoue, Masanari Ishida, Katsuya Takanashi, and Tatsuya Kawahara · 2017
Earlier work this paper cites.
Prediction of turn-taking using multitask learning with prediction of backchannels and fillers
Kohei Hara, Koji Inoue, Katsuya Takanashi, and Tatsuya Kawahara · 2018
Earlier work this paper cites.
Conversational and social laughter synthesis with WaveNet
Hiroki Mori, Tomohiro Nagata, and Yoshiko Arimoto · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Earlier work this paper cites.
Laughter synthesis: Combining seq2seq modeling with transfer learning
Noé Tits, Kevin El Haddad, and Thierry Dutoit · 2020
Cited alongside, same era.
Towards immediate backchannel generation using attention-based early prediction model
Amalia Istiqlali Adiba, Takeshi Homma, and Toshinori Miyoshi · 2021
Cited alongside, same era.
Engagement rewarded actor-critic with conservative Q-learning for speech-driven laughter backchannel generation
Öykü Zeynep Bayramoğlu, Engin Erzin, Tevfik Metin Sezgin, and Yücel Yemez · 2021
Cited alongside, same era.
Controllable context-aware conversational speech synthesis
Jian Cong, Shan Yang, Na Hu, Guangzhi Li, Lei Xie, and Dan Su · 2021
Cited alongside, same era.
Robust laughter detection in noisy environments
Jon Gillick, Wesley Deng, Kimiko Ryokai, and David Bamman · 2021
Cited alongside, same era.
Conversational end-to-end TTS for voice agents
NaturalSpeech: End-to-end text to speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, Frank Soong, Tao Qin, Sheng Zhao, and Tie-Yan Liu · 2022
Later among the works it cites.
SoundStorm: Efficient parallel audio generation
Zalán Borsos, Matt Sharifi, Damien Vincent, Eugene Kharitonov, Neil Zeghidour, and Marco Tagliasacchi · 2023
Closest in time.
AudioGPT: Understanding and generating speech, music, sound, and talking head
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, Yi Ren, Zhou Zhao, and Shinji Watanabe · 2023
Closest in time.
A generative framework for conversational laughter: Its ‘language model’ and laughter sound synthesis
Hiroki Mori and Shunya Kimura · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haohan Guo, Shaofei Zhang, Frank K Soong, Lei He, and Lei Xie · 2021
Cited alongside, same era.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Cited alongside, same era.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Cited alongside, same era.
Text-free prosody-aware generative spoken language modeling
Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Morgane Riviere, Abdelrahman Mohamed, Emmanuel Dupoux, and Wei-Ning Hsu · 2022
Cited alongside, same era.
Backchannel generation model for a third party listener agent
Divesh Lala, Koji Inoue, Tatsuya Kawahara, and Kei Sawada · 2022
Cited alongside, same era.
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, and Emmanuel Dupoux · 2023
Closest in time.
Generative Agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Closest in time.
AudioPaLM: A large language model that can speak and listen
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, Hannah Muckenhirn, Dirk Padfield, James Qin, Danny Rozenberg, Tara Sainath, Johan Schalkwyk, Matt Sharifi, Michelle Tadmor Ramanovich, Marco Tagliasacchi, Alexandru Tudor, Mihajlo Velimirović, Damien Vincent, Jiahui Yu, Yongqiang Wang, Vicky Zayats, Neil Zeghidour, Yu Zhang, Zhishuai Zhang, Lukas Zilka, and Christian Frank · 2023
Closest in time.
Response timing estimation for spoken dialog systems based on syntactic completeness prediction
Jin Sakuma, Shinya Fujie, and Tetsunori Kobayashi · 2023
Closest in time.
VioLA: Unified codec language models for speech recognition, synthesis, and translation
Tianrui Wang, Long Zhou, Ziqiang Zhang, Yu Wu, Shujie Liu, Yashesh Gaur, Zhuo Chen, Jinyu Li, and Furu Wei · 2023
Closest in time.
Laughter synthesis using pseudo phonetic tokens with a large-scale in-the-wild laughter corpus
Detai Xin, Shinnosuke Takamichi, Ai Morimatsu, and Hiroshi Saruwatari · 2023
Closest in time.
M2-CTTS: End-to-end multi-scale multi-modal conversational text-to-speech synthesis
Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li, Yingming Gao, Jianhua Tao, Jianqing Sun, and Jiaen Liang · 2023
Closest in time.
SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen · 2023
Closest in time.