Fetching the paper…
Reading the bibliography…
We introduce dGSLM, the first "textless" model able to generate audio samples of naturalistic spoken dialogues.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Dialogpt: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2019 · 1911
Earlier work this paper cites.
A statistical analysis of on-off patterns in 16 conversations
Paul T Brady. 1968 · 1968
Earlier work this paper cites.
On getting a word in edgewise
Victor H Yngve. 1970 · 1970
Earlier work this paper cites.
Some signals and rules for taking speaking turns in conversations
Starkey Duncan. 1972 · 1972
Earlier work this paper cites.
A simplest systematics for the organization of turn-taking for conversation
Harvey Sacks, Emanuel A. Schegloff, and Gail Jefferson. 1974 · 1974
Earlier work this paper cites.
Discourse as an interactional achievement: Some uses of ‘uh huh’and other things that come between sentences
Emanuel A Schegloff. 1982 · 1982
Earlier work this paper cites.
Overlapping talk and the organization of turn-taking for conversation
Emanuel A Schegloff. 2000 · 2000
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020 · 2001
Earlier work this paper cites.
Human Conversation as a System Framework: Designing Embodied Conversational Agents , page 29–63. MIT Press, Cambridge, MA, USA
Justine Cassell, Tim Bickmore, Lee Campbell, Hannes Vilhjálmsson, and Hao Yan. 2001 · 2001
Earlier work this paper cites.
Prosody-based automatic detection of annoyance and frustration in human-computer dialog
Jeremy Ang, Rajdip Dhillon, Ashley Krupski, Elizabeth Shriberg, and Andreas Stolcke. 2002 · 2002
Earlier work this paper cites.
Natural turn-taking needs no manual: Computational theory and model, from perception to action
Kristinn R. Thórisson. 2002 · 2002
Earlier work this paper cites.
The fisher corpus: a resource for the next generations of speech-to-text
Christopher Cieri, David Miller, and Kevin Walker. 2004 · 2004
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2020 · 2004
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2005
Earlier work this paper cites.
On temporal aspects of turn taking in conversational dialogues
Louis Ten Bosch, Nelleke Oostdijk, and Lou Boves. 2005 · 2005
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2006
Earlier work this paper cites.
A finite-state turn-taking model for spoken dialog systems
Antoine Raux and Maxine Eskenazi. 2009 · 2009
Earlier work this paper cites.
Universals and cultural variation in turn-taking in conversation
Tanya Stivers, Nicholas J Enfield, Penelope Brown, Christina Englert, Makoto Hayashi, Trine Heinemann, Gertie Hoymann, Federico Rossano, Jan Peter De Ruiter, Kyung-Eun Yoon, et al. 2009 · 2009
Earlier work this paper cites.
Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog
Erik Ekstedt and Gabriel Skantze. 2020 · 2010
Earlier work this paper cites.
Pauses, gaps and overlaps in conversations
Mattias Heldner and Jens Edlund. 2010 · 2010
Earlier work this paper cites.
Turn-taking cues in task-oriented dialogue
Agustín Gravano and Julia Hirschberg. 2011 · 2011
Earlier work this paper cites.
Crowdmos: An approach for crowdsourcing mean opinion score studies
Flávio Ribeiro, Dinei Florêncio, Cha Zhang, and Michael Seltzer. 2011 · 2011
Cited alongside, same era.
Paralinguistics in speech and language—state-of-the-art and the challenge
Björn Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian Müller, and Shrikanth Narayanan. 2013 · 2013
Cited alongside, same era.
Data-driven models for timing feedback responses in a map task dialogue system
Raveesh Meena, Gabriel Skantze, and Joakim Gustafson. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Timing in turn-taking and its implications for processing models of language
Stephen C Levinson and Francisco Torreira. 2015 · 2015
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux. 2020 · 2020
Later among the works it cites.
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Later among the works it cites.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee. 2020 · 2020
Later among the works it cites.
Variable-rate discrete representation learning
Sander Dieleman, Charlie Nash, Jesse Engel, and Karen Simonyan. 2021 · 2021
Later among the works it cites.
Revisiting the boundary between ASR and NLU in the age of conversational dialog systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2015 · 2015
Cited alongside, same era.
Oriol Vinyals and Quoc V. Le. 2015 · 2015
Cited alongside, same era.
Variational inference for acoustic unit discovery
Lucas Ondel, Lukás Burget, and Jan Cernocký. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Generative deep neural networks for dialogue: A short review
Iulian Vlad Serban, Ryan Lowe, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Towards a general, continuous model of turn-taking in spoken dialogue using lstm recurrent neural networks
Gabriel Skantze. 2017 · 2017
Cited alongside, same era.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Cited alongside, same era.
Manaal Faruqui and Dilek Hakkani-Tür. 2021 · 2021
Later among the works it cites.
Robust Laughter Detection in Noisy Environments
Jon Gillick, Wesley Deng, Kimiko Ryokai, and David Bamman. 2021 · 2021
Later among the works it cites.
Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, and Michael Auli. 2021b · 2021
Later among the works it cites.
Text-free prosody-aware generative spoken language modeling
Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, et al. 2021 · 2021
Later among the works it cites.
Internet-augmented dialogue generation
Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2021 · 2021
Later among the works it cites.
Textless speech emotion conversion using decomposed and discrete representations
Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu-Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi. 2021 · 2021
Later among the works it cites.
On Generative Spoken Language Modeling from Raw Audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Recent advances in deep learning based dialogue systems: A systematic survey
Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, Vinay Adiga, and Erik Cambria. 2021 · 2021
Later among the works it cites.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Towards unsupervised learning of speech features in the wild
Morgane Rivière and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Turn-taking in conversational systems and human-robot interaction: a review
Gabriel Skantze. 2021 · 2021
Later among the works it cites.
Beyond goldfish memory: Long-term open-domain conversation
Jing Xu, Arthur Szlam, and Jason Weston. 2021 · 2021
Later among the works it cites.
Superb: Speech processing universal performance benchmark
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al. 2021 · 2021
Later among the works it cites.
A brief overview of unsupervised neural speech representation learning
Lasse Borgholt, Jakob Drachmann Havtorn, Joakim Edin, Lars Maaløe, and Christian Igel. 2022 · 2022
Closest in time.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2022 · 2022
Closest in time.
A systematic review and bayesian meta-analysis of the development of turn taking in adult-child vocal interactions
Vivian Nguyen, Otto Versyp, Christopher Cox, and Riccardo Fusaroli. 2022 · 2022
Closest in time.
Language models that seek for knowledge: Modular search & generation for dialogue and prompt completion
Kurt Shuster, Mojtaba Komeili, Leonard Adolphs, Stephen Roller, Arthur Szlam, and Jason Weston. 2022 · 2022
Closest in time.