Fetching the paper…
Reading the bibliography…
Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding.
Bleu: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Speech emotion recognition using hidden Markov models
Nwe, T. L.; Foo, S. W.; and De Silva, L. C. 2003 · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
Optimizing speech emotion recognition using manta-ray based feature selection
Chattopadhyay, S.; Dey, A.; Basak, H.; et al. 2020 · 2009
Earlier work this paper cites.
Survey on speech emotion recognition: Features, classification schemes, and databases
El Ayadi, M.; Kamel, M. S.; et al. 2011 · 2011
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Sutskever, I.; Vinyals, O.; and Le, Q. V. 2014 · 2014
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2016 · 2016
Earlier work this paper cites.
Speaker dependent speech emotion recognition using MFCC and Support Vector Machine
Dahake, P. P.; Shaw, K.; et al. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation
Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Unsupervised feature learning based on deep models for environmental audio tagging
Xu, Y.; Huang, Q.; Wang, W.; Foster, P.; Sigtia, S.; Jackson, P. J.; and Plumbley, M. D. 2017 · 2017
Earlier work this paper cites.
Mutual information neural estimation
Belghazi, M. I.; Baratin, A.; Rajeshwar, S.; Ozair, S.; Bengio, Y.; Courville, A.; and Hjelm, D. 2018 · 2018
Earlier work this paper cites.
Speech emotion recognition based on an improved brain emotion learning model
Liu, Z.-T.; Xie, Q.; Wu, M.; Cao, W.-H.; Mei, Y.; and Mao, J.-W. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Parallelized convolutional recurrent neural network with spectral features for speech emotion recognition
Jiang, P.; Fu, H.; Tao, H.; Lei, P.; and Zhao, L. 2019 · 2019
Cited alongside, same era.
Speech emotion recognition using deep learning techniques: A review
Khalil, R. A.; Jones, E.; Babar, M. I.; Jan, T.; Zafar, M. H.; and Alhussain, T. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Cited alongside, same era.
Describing like humans: on diversity in image captioning
Wang, Q.; and Chan, A. B. 2019 · 2019
Cited alongside, same era.
A framework for the robust evaluation of sound event detection
Bilen, Ç.; Ferroni, G.; Tuveri, F.; Azcarreta, J.; and Krstulović, S. 2020 · 2020
Cited alongside, same era.
Audio Captioning Based on Transformer and Pre-Trained CNN
Chen, K.; Wu, Y.; Wang, Z.; Zhang, X.; Nian, F.; Li, S.; and Shao, X. 2020 · 2020
Large language models are zero-shot reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Later among the works it cites.
text2vec: A Tool for Text to Vector
Ming, X. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilić, S.; Hesslow, D.; Castagné, R.; Luccioni, A. S.; Yvon, F.; Gallé, M.; et al. 2022 · 2022
Later among the works it cites.
A comprehensive survey of automated audio captioning
Xu, X.; Wu, M.; and Yu, K. 2022 · 2022
Later among the works it cites.
Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Zhang, B.; Lv, H.; Guo, P.; Shao, Q.; Yang, C.; Xie, L.; Xu, X.; Bu, H.; Chen, X.; Zeng, C.; et al. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Club: A contrastive log-ratio upper bound of mutual information
Cheng, P.; Hao, W.; Dai, S.; Liu, J.; Gan, Z.; and Carin, L. 2020 · 2020
Cited alongside, same era.
Pre-training with whole word masking for chinese bert
Cui, Y.; Che, W.; Liu, T.; Qin, B.; and Yang, Z. 2021 · 2021
Cited alongside, same era.
Automated Audio Captioning with Weakly Supervised Pre-Training and Word Selection Methods
Han, Q.; Yuan, W.; Liu, D.; Li, X.; and Yang, Z. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021 · 2021
Cited alongside, same era.
Arabic speech emotion recognition employing wav2vec2. 0 and hubert based on baved dataset
Mohamed, O.; and Aly, S. A. 2021 · 2021
Cited alongside, same era.
Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information
Ye, Z.; Wang, H.; Yang, D.; and Zou, Y. 2021 · 2021
Cited alongside, same era.
Cui, Y.; Yang, Z.; and Yao, X. 2023 · 2023
Closest in time.
Audiogpt: Understanding and generating speech, music, sound, and talking head
Huang, R.; Li, M.; Yang, D.; Shi, J.; Chang, X.; Ye, Z.; Wu, Y.; Hong, Z.; Huang, J.; Liu, J.; et al. 2023 · 2023
Closest in time.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Closest in time.
Mei, X.; Meng, C.; Liu, H.; Kong, Q.; Ko, T.; Zhao, C.; Plumbley, M. D.; Zou, Y.; and Wang, W. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023 · 2023
Closest in time.
Large language models encode clinical knowledge
Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S. S.; Wei, J.; Chung, H. W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023 · 2023
Closest in time.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H.; Li, X.; and Bing, L. 2023 · 2023
Closest in time.