Fetching the paper…
Reading the bibliography…
Spoken dialogue modeling poses challenges beyond text-based language modeling, requiring real-time interaction, turn-taking, and backchanneling.
S. Duncan, “Some signals and rules for taking speaking turns in conversations.”
1972
Earlier work this paper cites.
E. A. Schegloff,
1982
Earlier work this paper cites.
G. Jefferson, “Preliminary notes on a possible metric which provides for a ‘standard maximum’silence of approximately one second in conversation,”
1989
Earlier work this paper cites.
E. A. Schegloff, “Overlapping talk and the organization of turn-taking for conversation,”
2000
Earlier work this paper cites.
N. Ward and W. Tsukahara, “Prosodic features which cue back-channel responses in english and japanese,”
2000
Earlier work this paper cites.
C. Cieri, D. Graff, O. Kimball, D. Miller, and K. Walker, “Fisher english training speech part 1 transcripts,”
2004
Earlier work this paper cites.
M. Heldner and J. Edlund, “Pauses, gaps and overlaps in conversations,”
2010
Earlier work this paper cites.
A. Gravano and J. Hirschberg, “Turn-taking cues in task-oriented dialogue,”
2011
Earlier work this paper cites.
A. Raux and M. Eskenazi, “Optimizing the turn-taking behavior of task-oriented spoken dialog systems,”
2012
Earlier work this paper cites.
N. Mostafazadeh, M. Roth, A. Louis, N. Chambers, and J. Allen, “LSDSem 2017 shared task: The story cloze test,” in
2017
Earlier work this paper cites.
T. A. Nguyen, M. de Seyssel, P. Rozé, M. Rivière, E. Kharitonov, A. Baevski, E. Dunbar, and E. Dupoux, “The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,” in
2020
Earlier work this paper cites.
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” in
2021
Earlier work this paper cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
2021
Earlier work this paper cites.
G.-T. Lin, Y.-S. Chuang, H.-L. Chung, S. wen Yang, H.-J. Chen, S. A. Dong, S.-W. Li, A. Mohamed, H. yi Lee, and L. shan Lee, “DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering,” in
2022
Earlier work this paper cites.
H.-S. Tsai, H.-J. Chang, W.-C. Huang, Z. Huang, K. Lakhotia, S.-w. Yang, S. Dong, A. Liu, C.-I. Lai, J. Shi
2022
Earlier work this paper cites.
S. Liu, Y. Nakajima, L. Chen, S. Arndt, M. Kakizoe, M. A. Elliott, and G. B. Remijn, “How pause duration influences impressions of english speech: Comparison between native and non-native speakers,”
2022
Earlier work this paper cites.
T. A. Nguyen, E. Kharitonov, J. Copet, Y. Adi, W.-N. Hsu, A. Elkahky, P. Tomasello, R. Algayres, B. Sagot, A. Mohamed
2023
Earlier work this paper cites.
G.-T. Lin, C.-L. Feng, W.-P. Huang, Y. Tseng, T.-H. Lin, C.-A. Li, H.-y. Lee, and N. G. Ward, “On the utility of self-supervised models for prosody-related tasks,” in
2023
Cited alongside, same era.
A. Reece, G. Cooney, P. Bull, C. Chung, B. Dawson, C. Fitzpatrick, T. Glazer, D. Knox, A. Liebscher, and S. Marin, “The candor corpus: Insights from a large multimodal dataset of naturalistic conversation,”
2023
Cited alongside, same era.
OpenAI, “Gpt-4 technical report,” 2023
2023
Cited alongside, same era.
B. Veluri, B. Peloquin, B. Yu, H. Gong, and S. Gollakota, “Beyond turn-based interfaces: Synchronous llms as full-duplex dialogue agents,” in
2024
Cited alongside, same era.
A. Défossez, L. Mazaré, M. Orsini, A. Royer, P. Pérez, H. Jégou, E. Grave, and N. Zeghidour, “Moshi: a speech-text foundation model for real-time dialogue,” Kyutai, Tech. Rep., September 2024. [Online]. Available:
G.-T. Lin, C.-H. Chiang, and H.-y. Lee, “Advancing large language models to capture varied speaking styles and respond properly in spoken conversations,” in
2024
Later among the works it cites.
G.-T. Lin, P. G. Shivakumar, A. Gandhe, C.-H. H. Yang, Y. Gu, S. Ghosh, A. Stolcke, H.-Y. Lee, and I. Bulyko, “Paralinguistics-enhanced large language modeling of spoken dialogue,” in
2024
Later among the works it cites.
G.-T. Lin and H.-y. Lee, “Can LLMs understand the implication of emphasized sentences in dialogue?” in
2024
Later among the works it cites.
P. Wang, S. Lu, Y. Tang, S. Yan, W. Xia, and Y. Xiong, “A full-duplex speech dialogue scheme based on large language model,” in
2024
Later among the works it cites.
X. Zhang, Y. Chen, S. Hu, X. Han, Z. Xu, Y. Xu, W. Zhao, M. Sun, and Z. Liu, “Beyond the turn-based game: Enabling real-time conversations with duplex models,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
P. Wang, S. Lu, Y. Tang, S. Yan, W. Xia, and Y. Xiong, “A full-duplex speech dialogue scheme based on large language model,” in
2024
Cited alongside, same era.
C. Fu, H. Lin, Z. Long, Y. Shen, M. Zhao, Y. Zhang, S. Dong, X. Wang, D. Yin, L. Ma
2024
Cited alongside, same era.
2024
Cited alongside, same era.
S. Ji, Y. Chen, M. Fang, J. Zuo, J. Lu, H. Wang, Z. Jiang, L. Zhou, S. Liu, X. Cheng
2024
Cited alongside, same era.
M. Hassid, T. Remez, T. A. Nguyen, I. Gat, A. Conneau, F. Kreuk, J. Copet, A. Defossez, G. Synnaeve, E. Dupoux
2024
Cited alongside, same era.
E. Nachmani, A. Levkovitch, R. Hirsch, J. Salazar, C. Asawaroengchai, S. Mariooryad, E. Rivlin, R. Skerry-Ryan, and M. T. Ramanovich, “Spoken question answering and speech continuation using spectrogram-powered LLM,” in
2024
Cited alongside, same era.
M.-H. Shih, H.-L. Chung, Y.-C. Pai, M.-H. Hsu, G.-T. Lin, S.-W. Li, and H. yi Lee, “Gsqa: An end-to-end model for generative spoken question answering,” in
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Q. Zhang, L. Cheng, C. Deng, Q. Chen, W. Wang, S. Zheng, J. Liu, H. Yu, C. Tan, Z. Du
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Umair, V. Sarathy, and J. Ruiter, “Large language models know what to say but not when to speak,” in
2024
Later among the works it cites.
Q. Chen, Y. Chen, Y. Chen, M. Chen, Y. Chen, C. Deng, Z. Du, R. Gao, C. Gao, Z. Gao
2025
Closest in time.
X. Cheng, R. Hu, X. Yang, J. Lu, D. Fu, Z. Wang, S. Ji, R. Huang, B. Zhang, T. Jin, and Z. Zhao, “Voxdialogue: Can spoken dialogue systems understand information beyond words?” in
2025
Closest in time.
Q. Wang, Z. Meng, W. Cui, Y. Zhang, P. Wu, B. Wu, Z. Zheng, I. King, L. Chen, and P. Zhao, “Parrot: Seamless spoken dialogue interaction with double-channel large language models,” 2025. [Online]. Available:
2025
Closest in time.
L. Mai and J. Carson-Berndsen, “Real-time textless dialogue generation,”
2025
Closest in time.
A. A. G. Intelligence, “Amazon nova sonic: Technical report and model card,” 2025
2025
Closest in time.
S. Arora, Z. Lu, C.-C. Chiu, R. Pang, and S. Watanabe, “Talking turns: Benchmarking audio foundation models on turn-taking dynamics,” in
2025
Closest in time.