Fetching the paper…
Reading the bibliography…
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.
Constants across cultures in the face and emotion
Paul Ekman and Wallace V Friesen · 1971
Earlier work this paper cites.
The role of voice input for human-machine communication
P R Cohen and S L Oviatt · 1995
Earlier work this paper cites.
Emotion recognition in human-computer interaction
R. Cowie, E. Douglas-Cowie, N. Tsapatsoulis, G. Votsis, S. Kollias, W. Fellenz, and J.G. Taylor · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Wired for Speech: How Voice Activates and Advances the Human-Computer Relationship
Clifford Nass and Scott Brave · 2005
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan · 2008
Earlier work this paper cites.
The semaine corpus of emotionally coloured character interactions
Gary McKeown, Michel F Valstar, Roderick Cowie, and Maja Pantic · 2010
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Bo-Hsiang Tseng, Sheng-Syun Shen, Hung-Yi Lee, and Lin-Shan Lee · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings
Reza Lotfian and Carlos Busso · 2017
Earlier work this paper cites.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser · 2017
Earlier work this paper cites.
The emotional voices database: Towards controlling the emotion dimension in voice generation systems
Adaeze Adigwe, Noé Tits, Kevin El Haddad, Sarah Ostadabbas, and Thierry Dutoit · 2018
Earlier work this paper cites.
An Open Source Emotional Speech Corpus for Human Robot Interaction Applications
Jesin James, Li Tian, and Catherine Inez Watson · 2018
Earlier work this paper cites.
Chia-Hsuan Li, Szu-Lin Wu, Chi-Liang Liu, and Hung-yi Lee · 2018
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
MELD: A multimodal multi-party dataset for emotion recognition in conversations
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea · 2019
Earlier work this paper cites.
Emotion and sentiment analysis from twitter text
Kashfia Sailunaz and Reda Alhajj · 2019
Earlier work this paper cites.
CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit, 2019
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald · 2019
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber · 2020
Cited alongside, same era.
Open-source Multi-speaker Corpora of the English Accents in the British Isles
Isin Demirsahin, Oddur Kjartansson, Alexander Gutkin, and Clara Rivera · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al · 2020
Cited alongside, same era.
Improving spoken question answering using contextualized word representation
Dan Su and Pascale Fung · 2020
Cited alongside, same era.
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy · 2020
Cited alongside, same era.
Heysquad: A spoken question answering dataset
Yijing Wu, SaiKrishna Rallabandi, Ravisutha Srinivasamurthy, Parag Pravin Dakle, Alolika Gon, and Preethi Raghavan · 2023
Later among the works it cites.
E-chat: Emotion-sensitive spoken dialogue system with large language models
Hongfei Xue, Yuhao Liang, Bingshen Mu, Shiliang Zhang, Qian Chen, and Lei Xie · 2023
Later among the works it cites.
SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Emotional voice conversion: Theory, databases and esd
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li · 2021
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
End-to-end spoken conversational question answering: Task, dataset and model
Chenyu You, Nuo Chen, Fenglin Liu, Shen Ge, Xian Wu, and Yuexian Zou · 2022
Cited alongside, same era.
Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al · 2024
Closest in time.
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al · 2024
Closest in time.
Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, et al · 2024
Closest in time.
Wavllm: Towards robust and adaptive speech large language model
Shujie Hu, Long Zhou, Shujie Liu, Sanyuan Chen, Hongkun Hao, Jing Pan, Xunying Liu, Jinyu Li, Sunit Sivasankaran, Linquan Liu, et al · 2024
Closest in time.
Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al · 2024
Closest in time.
Clam-tts: Improving neural codec language model for zero-shot text-to-speech
Jaehyeon Kim, Keon Lee, Seungjun Chung, and Jaewoong Cho · 2024
Closest in time.
Base tts: Lessons from building a billion-parameter text-to-speech model on 100k hours of data
Mateusz Łajszczak, Guillermo Cámbara, Yang Li, Fatih Beyhan, Arent van Korlaar, Fan Yang, Arnaud Joly, Álvaro Martín-Cortinas, Ammar Abbas, Adam Michalski, et al · 2024
Closest in time.
Guan-Ting Lin, Cheng-Han Chiang, and Hung-yi Lee · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Gpt-4o, 2024
OpenAI · 2024
Closest in time.
Voicecraft: Zero-shot speech editing and text-to-speech in the wild
Puyuan Peng, Po-Yao Huang, Daniel Li, Abdelrahman Mohamed, and David Harwath · 2024
Closest in time.
My science tutor (MyST)–a large corpus of children’s conversational speech
Sameer Pradhan, Ronald A. Cole, and Wayne H. Ward · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
SALMONN: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA, and Chao Zhang · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al · 2024
Closest in time.
Amphion: An open-source audio, music and speech generation toolkit
Xueyao Zhang, Liumeng Xue, Yicheng Gu, Yuancheng Wang, Jiaqi Li, Haorui He, Chaoren Wang, Ting Song, Xi Chen, Zihao Fang, Haopeng Chen, Junan Zhang, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Jun Han, Kai Chen, Haizhou Li, and Zhizheng Wu · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.