Fetching the paper…
Reading the bibliography…
This study focuses on emotion-sensitive spoken dialogue in human-machine speech interaction.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
C. Bothe, S. Magg, C. Weber, and S. Wermter, “Dialogue-based neural learning to estimate the sentiment of a next upcoming utterance,” in
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline,” in
2017
Earlier work this paper cites.
X. Li and M. Zhang, “Emotion analysis for the upcoming response in open-domain human-computer conversation,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell
2020
Earlier work this paper cites.
W. Hsu, B. Bolte, Y. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,”
2021
Earlier work this paper cites.
D. Varshney, A. Ekbal, and P. Bhattacharyya, “Modelling context emotions using multi-task learning for emotion controlled dialog generation,” in
2021
Earlier work this paper cites.
R. Liu, J. Wei, C. Jia, and S. Vosoughi, “Modulating language models with emotions,” in
2021
Earlier work this paper cites.
OpenAI, “Introducing chatgpt,”
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in
2022
Cited alongside, same era.
B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng, D. Wu, and Z. Peng, “WENETSPEECH: A 10000+ hours multi-domain mandarin corpus for speech recognition,” in
2022
Cited alongside, same era.
E. Morais, R. Hoory, W. Zhu, I. Gat, M. Damasceno, and H. Aronowitz, “Speech emotion recognition using self-supervised features,” in
2022
Cited alongside, same era.
2023
Cited alongside, same era.
H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,” in
2023
Closest in time.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. C. H. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” in
2023
Closest in time.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in
2023
Closest in time.
S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An audio language model for audio tasks,” in
2023
Closest in time.
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
D. Zhang, S. Li, X. Zhang, J. Zhan, P. Wang, Y. Zhou, and X. Qiu, “Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. Glass, “Listen, think, and understand,”
2023
Cited alongside, same era.
J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in
2023
Cited alongside, same era.
Closest in time.
2023
Closest in time.
2023
Closest in time.
G.-T. Lin, P. G. Shivakumar, A. Gandhe, C.-H. H. Yang, Y. Gu, S. Ghosh, A. Stolcke, H. yi Lee, and I. Bulyko, “Paralinguistics-enhanced large language modeling of spoken dialogue,” in
2024
Closest in time.
Y. Xu, H. Chen, J. Yu, Q. Huang, Z. Wu, S. Zhang, G. Li, Y. Luo, and R. Gu, “Secap: Speech emotion captioning with large language model,”
2024
Closest in time.