Fetching the paper…
Reading the bibliography…
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data.
Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals
Ary L Goldberger, Luis AN Amaral, Leon Glass, Jeffrey M Hausdorff, Plamen Ch Ivanov, Roger G Mark, Joseph E Mietus, George B Moody, Chung-Kang Peng, and H Eugene Stanley · 2000
Earlier work this paper cites.
The ami meeting corpus: A pre-announcement
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, et al · 2005
Earlier work this paper cites.
An initial study on stress detection for spoken english
Ching-Yu Tseng · 2008
Earlier work this paper cites.
Evaluation of algorithms using games: The case of music tagging
Edith Law, Kris West, Michael I Mandel, Mert Bay, and J Stephen Downie · 2009
Earlier work this paper cites.
SemEval-2012 task 6: A pilot on semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre · 2012
Earlier work this paper cites.
Stress detection of english words for a capt system using word-length dependent gmm-based bayesian classifiers
Liang-Yu CHEN and Jyh-Shing Roger JANG · 2012
Earlier work this paper cites.
Voxforge, 2018
Ken MacLean · 2012
Earlier work this paper cites.
Modal analysis and transcription of strokes of the mridangam using non-negative matrix factorization
Akshay Anantapadmanabhan, Ashwin Bellur, and Hema A Murthy · 2013
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
J. Salamon, C. Jacoby, and J. P. Bello · 2014
Earlier work this paper cites.
A study of instrument-wise onset detection in beijing opera percussion ensembles
Mi Tian, Ajay Srinivasamurthy, Mark Sandler, and Xavier Serra · 2014
Earlier work this paper cites.
Studying emotion induced by music through a crowdsourcing game
Anna Aljanaki, Frans Wiering, and Remco C. Veltkamp · 2015
Earlier work this paper cites.
Quesst2014: Evaluating query-by-example speech search in a zero-resource setting with real-life queries
Xavier Anguera, Luis-J. Rodriguez-Fuentes, Andi Buzo, Florian Metze, Igor Szöke, and Mikel Penagarikano · 2015
Earlier work this paper cites.
Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections
Peter Knees, Ángel Faraldo Pérez, Herrera Boyer, Richard Vogl, Sebastian Böck, Florian Hörschläger, Mickael Le Goff, et al · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
Musan: A music, speech, and noise corpus
David Snyder, Guoguo Chen, and Daniel Povey · 2015
Earlier work this paper cites.
English lexical stress detection and sentence-based intonation assessment based on contour shape description
Sheng-Chi Tsai · 2015
Earlier work this paper cites.
Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge
Zhizheng Wu, Tomi Kinnunen, Nicholas Evans, Junichi Yamagishi, Cemal Hanilçi, Md. Sahidullah, and Aleksandr Sizov · 2015
Earlier work this paper cites.
Fma: A dataset for music analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, and Xavier Bresson · 2017
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan · 2017
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
A study on data augmentation of reverberant speech for robust speech recognition
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L. Seltzer, and Sanjeev Khudanpur · 2017
Earlier work this paper cites.
The Fifth ’CHiME’ Speech Separation and Recognition Challenge: Dataset, Task and Baselines
Jon Barker, Shinji Watanabe, Emmanuel Vincent, and Jan Trmal · 2018
Earlier work this paper cites.
ISMIR04 Genre Identification task dataset (1.0), 2018
P. Cano, N. Wack, and P. Herrera · 2018
Earlier work this paper cites.
Asvspoof 2017 version 2.0: meta-data analysis and baseline enhancements
Héctor Delgado, Massimiliano Todisco, Md Sahidullah, Nicholas Evans, Tomi Kinnunen, Kong Aik Lee, and Junichi Yamagishi · 2018
Earlier work this paper cites.
Openmic-2018: An open data-set for multiple instrument recognition
Eric Humphrey, Simon Durand, and Brian McFee · 2018
Earlier work this paper cites.
Interpersonal relationship labels for the CALLHOME corpus
Denys Katerenchuk, David Guy Brizan, and Andrew Rosenberg · 2018
Earlier work this paper cites.
Vocal imitation set: a dataset of vocally imitated sound events using the audioset ontology
Bongjun Kim, Madhav Ghei, Bryan Pardo, and Zhiyao Duan · 2018
Earlier work this paper cites.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Steven R Livingstone and Frank A Russo · 2018
Earlier work this paper cites.
Domestic cat sound classification using transfer learning
Yagya Raj Pandeya and Joonwhoan Lee · 2018
Earlier work this paper cites.
Domestic cat sound classification using learned features from deep neural nets
Yagya Raj Pandeya, Dongwhoon Kim, and Joonwhoan Lee · 2018
Earlier work this paper cites.
LibriCount, a dataset for speaker count estimation, 2018
Fabian-Robert Stöter, Soumitro Chakrabarty, Emanuël Habets, and Bernd Edler · 2018
Earlier work this paper cites.
Speech commands: A dataset for limited-vocabulary speech recognition
Pete Warden · 2018
Earlier work this paper cites.
Vocalset: A singing voice dataset
Julia Wilkins, Prem Seetharaman, Alison Wahl, and Bryan Pardo · 2018
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
Dmitry Bogdanov, Minz Won, Philip Tovstogan, Alastair Porter, and Xavier Serra · 2019
Earlier work this paper cites.
Towards multimodal sarcasm detection (an _Obviously_ perfect paper)
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria · 2019
Earlier work this paper cites.
Enabling factorized piano music modeling and generation with the MAESTRO dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck · 2019
Earlier work this paper cites.
Audio-based identification of beehive states
Inês Nolasco, Alessandro Terenzi, Stefania Cecchi, Simone Orcioni, Helen L Bear, and Emmanouil Benetos · 2019
Earlier work this paper cites.
MELD: A multimodal multi-party dataset for emotion recognition in conversations
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea · 2019
Earlier work this paper cites.
MUSDB18-HQ - an uncompressed version of musdb18, December 2019
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner · 2019
Earlier work this paper cites.
Automatic acoustic detection of birds through deep learning: the first bird audio detection challenge
Dan Stowell, Michael D Wood, Hanna Pamuła, Yannis Stylianou, and Hervé Glotin · 2019
Earlier work this paper cites.
Sound event detection in domestic environments with weakly labeled data and soundscape synthesis
Nicolas Turpault, Romain Serizel, Ankit Parag Shah, and Justin Salamon · 2019
Cited alongside, same era.
Wham!: Extending speech separation to noisy environments
Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, and Jonathan Le Roux · 2019
Cited alongside, same era.
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald · 2019
Cited alongside, same era.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu · 2019
Cited alongside, same era.
Accentdb: A database of non-native english accents to assist neural speech recognition
Afroz Ahamad, Ankit Anand, and Pranesh Bhargava · 2020
Cited alongside, same era.
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou · 2023
Later among the works it cites.
Prosaudit, a prosodic benchmark for self-supervised speech models
Maureen de Seyssel, Marvin Lavechin, Hadrien Titeux, Arthur Thomas, Gwendal Virlet, Andrea Santos Revilla, Guillaume Wisniewski, Bogdan Ludusan, and Emmanuel Dupoux · 2023
Later among the works it cites.
Joint audio and speech understanding
Yuan Gong, Alexander H Liu, Hongyin Luo, Leonid Karlinsky, and James Glass · 2023
Later among the works it cites.
Prompttts: Controllable text-to-speech with text descriptions
Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao, and Xu Tan · 2023
Later among the works it cites.
Indicsuperb: A speech processing universal performance benchmark for indian languages
Tahir Javed, Kaushal Bhogale, Abhigyan Raman, Pratyush Kumar, Anoop Kunchukuttan, and Mitesh M Khapra · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber · 2020
Cited alongside, same era.
SLURP: A Spoken Language Understanding Resource Package
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser · 2020
Cited alongside, same era.
Children’s song dataset for singing voice research
Soonbeom Choi, Won Il Kim, Sae Byul Park, Sangeon Yong, and Juhan Nam · 2020
Cited alongside, same era.
Librimix: An open-source dataset for generalizable speech separation, 2020
Joris Cosentino, Manuel Pariente, Samuele Cornell, Antoine Deleforge, and Emmanuel Vincent · 2020
Cited alongside, same era.
Clotho: an audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen · 2020
Cited alongside, same era.
ASAP: a dataset of aligned scores and performances for piano transcription
Francesco Foscarin, Andrew McLeod, Philippe Rigaux, Florent Jacquemard, and Masahiko Sakai · 2020
Cited alongside, same era.
Cochleanet: A robust language-independent audio-visual model for real-time speech enhancement
Mandar Gogate, Kia Dashtipour, Ahsan Adeel, and Amir Hussain · 2020
Cited alongside, same era.
The corpus of regional african american language, 2023
Tyler Kendall and Charlie Farrington · 2023
Later among the works it cites.
Dailytalk: Spoken dialogue dataset for conversational text-to-speech
Keon Lee, Kyumin Park, and Daeyoung Kim · 2023
Later among the works it cites.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Later among the works it cites.
Titouan Parcollet, Ha Nguyen, Solene Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le, Sina Alisamir, Natalia Tomashenko, Marco Dinarelli, et al · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
Nonspeech7k dataset: Classification and analysis of human non-speech sound
Muhammad Mamunur Rashid, Guiqing Li, and Chengrui Du · 2023
Later among the works it cites.
General purpose audio effect removal
Matthew Rice, Christian J. Steinmetz, George Fazekas, and Joshua D. Reiss · 2023
Later among the works it cites.
Spatial librispeech: An augmented dataset for spatial audio learning
Miguel Sarabia, Elena Menyaylenko, Alessandro Toso, Skyler Seto, Zakaria Aldeneh, Shadi Pirhosseinloo, Luca Zappella, Barry-John Theobald, Nicholas Apostoloff, and Jonathan Sheaffer · 2023
Later among the works it cites.
Ml-superb: Multilingual speech universal performance benchmark
Jiatong Shi, Dan Berrebbi, William Chen, En-Pei Hu, Wei-Ping Huang, Ho-Lam Chung, Xuankai Chang, Shang-Wen Li, Abdelrahman Mohamed, Hung yi Lee, and Shinji Watanabe · 2023
Later among the works it cites.
Slue phase-2: A benchmark suite of diverse spoken language understanding tasks
Suwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad, Felix Wu, Roshan S Sharma, Wei-Lun Wu, Hung-Yi Lee, Karen Livescu, and Shinji Watanabe · 2023
Later among the works it cites.
Is chatgpt a good nlg evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou · 2023
Later among the works it cites.
Marble: Music audio representation benchmark for universal evaluation
Ruibin Yuan, Yinghao Ma, Yizhi Li, Ge Zhang, Xingran Chen, Hanzhi Yin, Yiqi Liu, Jiawen Huang, Zeyue Tian, Binyue Deng, et al · 2023
Later among the works it cites.
Siqi Zheng, Luyao Cheng, Yafeng Chen, Hui Wang, and Qian Chen · 2023
Later among the works it cites.
https://huggingface.co/datasets/speech31/Librispeech_word
Librispeech word · 2024
Closest in time.
Audiomnist: Exploring explainable artificial intelligence for audio analysis on a simple benchmark
Sören Becker, Johanna Vielhaben, Marcel Ackermann, Klaus-Robert Müller, Sebastian Lapuschkin, and Wojciech Samek · 2024
Closest in time.
High school english listening exam
CEEC · 2024
Closest in time.
Phonetic segmentation of the UCLA phonetics lab archive
Eleanor Chodroff, Blaž Pažon, Annie Baker, and Steven Moran · 2024
Closest in time.
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al · 2024
Closest in time.
Musical instrument chord classification, 2024
DeepContractor · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha · 2024
Closest in time.
Wavllm: Towards robust and adaptive speech large language model
Shujie Hu, Long Zhou, Shujie Liu, Sanyuan Chen, Hongkun Hao, Jing Pan, Xunying Liu, Jinyu Li, Sunit Sivasankaran, Linquan Liu, et al · 2024
Closest in time.
Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech
Chien-yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Haibin Wu, Siddhant Arora, Kai-Wei Chang, Jiatong Shi, Yifan Peng, et al · 2024
Closest in time.
Zero resource code-switched speech benchmark using speech utterance pairs for multiple spoken languages
Kuan-Po Huang, Chih-Kai Yang, Yu-Kuan Fu, Ewan Dunbar, and Hung-yi Lee · 2024
Closest in time.
MERT: Acoustic music understanding model with large-scale self-supervised training
Yizhi LI, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi, Wenhao Huang, Zili Wang, Yike Guo, and Jie Fu · 2024
Closest in time.
Paralinguistics-enhanced large language modeling of spoken dialogue
Guan-Ting Lin, Prashanth Gurunath Shivakumar, Ankur Gandhe, Chao-Han Huck Yang, Yile Gu, Shalini Ghosh, Andreas Stolcke, Hung-yi Lee, and Ivan Bulyko · 2024
Closest in time.
Music understanding llama: Advancing text-to-music generation with question answering and captioning
Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan · 2024
Closest in time.
A suite for acoustic language model evaluation
Gallil Maimon, Amit Roth, and Yossi Adi · 2024
Closest in time.
The nccu (national chengchi university) corpus of spoken taiwan mandarin, n.d
National Chengchi University · 2024
Closest in time.
Ml-superb 2.0: Benchmarking multilingual speech models across modeling constraints, languages, and datasets
Jiatong Shi, Shih-Heng Wang, William Chen, Martijn Bartelds, Vanya Bannihatti Kumar, Jinchuan Tian, Xuankai Chang, Dan Jurafsky, Karen Livescu, Hung yi Lee, and Shinji Watanabe · 2024
Closest in time.
Starss23: An audio-visual dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events
Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel A Krause, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, et al · 2024
Closest in time.
Sound effects library
WavSource · 2024
Closest in time.
Human screaming detection dataset, 2023
whats2000 · 2024
Closest in time.
Investigating zero-shot generalizability on mandarin-english code-switched asr and speech-to-text translation of recent foundation models with self-supervision and weak supervision
Chih-Kai Yang, Kuan-Po Huang, Ke-Han Lu, Chun-Yi Kuan, Chi-Yuan Hsiao, and Hung-yi Lee · 2024
Closest in time.
Ctrsvdd: A benchmark dataset and baseline analysis for controlled singing voice deepfake detection
Yongyi Zang, Jiatong Shi, You Zhang, Ryuichi Yamamoto, Jionghao Han, Yuxun Tang, Shengyuan Xu, Wenxiao Zhao, Jing Guo, Tomoki Toda, and Zhiyao Duan · 2024
Closest in time.
IEEE DCASE 2016 Challenge
IEEE DCASE 2016 Challenge · 2025
Closest in time.
Covost 2 and massively multilingual speech translation
Changhan Wang, Anne Wu, Jiatao Gu, and Juan Pino · 2027
Closest in time.