Fetching the paper…
Reading the bibliography…
This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs).
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Iemocap: interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Ebrahim (Abe) Kazemzadeh, Emily Mower Provost, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan · 2008
Earlier work this paper cites.
Design of speech corpus for mandarin text to speech
JTFLM Zhang and Huibin Jia · 2008
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Using audio books for training a text-to-speech system
Aimilios Chalamandaris, Pirros Tsiakoulis, Sotiris Karabetsos, and Spyros Raptis · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
ESC: dataset for environmental sound classification
Karol J. Piczak · 2015
Earlier work this paper cites.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng · 2017
Earlier work this paper cites.
A mostly data-driven approach to inverse text normalization
Ernest Pusateri, Bharat Ram Ambati, Elizabeth Brooks, Ondrej Plátek, Donald McAllaster, and Venki Nagesha · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
End-to-end automatic speech translation of audiobooks
Alexandre Berard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin · 2018
Earlier work this paper cites.
Aishell-2: Transforming mandarin asr research into industrial scale
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2019
Earlier work this paper cites.
MELD: A multimodal multi-party dataset for emotion recognition in conversations
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, and et al · 2019
Earlier work this paper cites.
Controllable time-delay transformer for real-time punctuation prediction and disfluency detection
Qian Chen, Mengzhe Chen, Bo Li, and Wen Wang · 2020
Earlier work this paper cites.
SAN-M: memory equipped self-attention for end-to-end speech recognition
Zhifu Gao, Shiliang Zhang, Ming Lei, and Ian McLoughlin · 2020
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley · 2020
Earlier work this paper cites.
The MSP-Conversation Corpus
Luz Martinez-Lucas, Mohammed Abdelwahab, and Carlos Busso · 2020
Cited alongside, same era.
Coswara — a database of breathing, cough, and voice sounds for covid-19 diagnosis
Neeraj Sharma, Prashant Krishnan, Rohit Kumar, Shreyas Ramoji, Srikanth Raj Chetupalli, Nirmala R., Prasanta Kumar Ghosh, and Sriram Ganapathy · 2020
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Sound event detection: A tutorial
Annamaria Mesaros, Toni Heittola, Tuomas Virtanen, and Mark D. Plumbley · 2021
Cited alongside, same era.
AISHELL-3: A multi-speaker mandarin TTS corpus
Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li · 2021
Cited alongside, same era.
Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset, 2021
Flow matching for generative modeling
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le · 2023
Later among the works it cites.
Matcha-tts: A fast TTS architecture with conditional flow matching
Shivam Mehta, Ruibo Tu, Jonas Beskow, Éva Székely, and Gustav Eje Henter · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
Reo Shimizu, Ryuichi Yamamoto, Masaya Kawamura, Yuma Shirahata, Hironori Doi, Tatsuya Komatsu, and Kentaro Tachibana · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li · 2021
Cited alongside, same era.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Cited alongside, same era.
Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition
Zhifu Gao, Shiliang Zhang, Ian McLoughlin, and Zhijie Yan · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
Newspod: Automatic and interactive news podcasts
Philippe Laban, Elicia Ye, Srujay Korlakunta, John F. Canny, and Marti A. Hearst · 2022
Cited alongside, same era.
Text-driven emotional style control and cross-speaker style transfer in neural TTS
Yookyung Shin, Younggun Lee, Suhee Jo, Yeongtae Hwang, and Taesu Kim · 2022
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2022
Cited alongside, same era.
Later among the works it cites.
Qwen Team · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Hifi-codec: Group-residual vector quantization for high fidelity audio codec
Dongchao Yang, Songxiang Liu, Rongjie Huang, Jinchuan Tian, Chao Weng, and Yuexian Zou · 2023
Later among the works it cites.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei · 2023
Later among the works it cites.
Shengpeng Ji, Jialong Zuo, Minghui Fang, Siqi Zheng, Qian Chen, Wen Wang, Ziyue Jiang, Hai Huang, Xize Cheng, Rongjie Huang, and Zhou Zhao · 2024
Closest in time.
Making flow-matching-based zero-shot text-to-speech laugh as you like
Naoyuki Kanda, Xiaofei Wang, Sefik Emre Eskimez, Manthan Thakker, Hemin Yang, Zirun Zhu, Min Tang, Canrun Li, Chung-Hsien Tsai, Zhen Xiao, Yufei Xia, Jinzhu Li, Yanqing Liu, Sheng Zhao, and Michael Zeng · 2024
Closest in time.
Merbench: A unified evaluation benchmark for multimodal emotion recognition, 2024
Zheng Lian, Licai Sun, Yong Ren, Hao Gu, Haiyang Sun, Lan Chen, Bin Liu, and Jianhua Tao · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy P. Lillicrap, Jean-Baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew M. Dai, Katie Millican, Ethan Dyer, Mia Glaese, Thibault Sottiaux, Benjamin Lee, Fabio Viola, Malcolm Reynolds, Yuanzhong Xu, James Molloy, Jilin Chen, Michael Isard, Paul Barham, Tom Hennigan, Ross McIlroy, Melvin Johnson, Johan Schalkwyk, Eli Collins, Eliza Rutherford, Erica Moreira, Kareem Ayoub, Megha Goel, Clemens Meyer, Gregory Thornton, Zhen Yang, Henryk Michalewski, Zaheer Abbas, Nathan Schucher, Ankesh Anand, Richard Ives, James Keeling, Karel Lenc, Salem Haykal, Siamak Shakeri, Pranav Shyam, Aakanksha Chowdhery, Roman Ring, Stephen Spencer, Eren Sezener, and et al · 2024
Closest in time.
Seaco-paraformer: A non-autoregressive asr system with flexible and effective hotword customization ability
Xian Shi, Yexin Yang, Zerui Li, Yanni Chen, Zhifu Gao, and Shiliang Zhang · 2024
Closest in time.
SALMONN: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA, and Chao Zhang · 2024
Closest in time.
Emo-superb: An in-depth look at speech emotion recognition, 2024
Haibin Wu, Huang-Cheng Chou, Kai-Wei Chang, Lucas Goncalves, Jiawei Du, Jyh-Shing Roger Jang, Chi-Chun Lee, and Hung-Yi Lee · 2024
Closest in time.
E-chat: Emotion-sensitive spoken dialogue system with large language models
Hongfei Xue, Yuhao Liang, Bingshen Mu, Shiliang Zhang, Mengzhe Chen, Qian Chen, and Lei Xie · 2024
Closest in time.