Fetching the paper…
Reading the bibliography…
The sound codec's dual roles in minimizing data transmission latency and serving as tokenizers underscore its critical importance.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 1904
Earlier work this paper cites.
Mel-cepstral distance measure for objective speech quality assessment
R. Kubichek. 1993 · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra. 2001 · 2001
Earlier work this paper cites.
Musical genre classification of audio signals
George Tzanetakis and Perry Cook. 2002 · 2002
Earlier work this paper cites.
Librimix: An open-source dataset for generalizable speech separation
Joris Cosentino, Manuel Pariente, Samuele Cornell, Antoine Deleforge, and Emmanuel Vincent. 2020 · 2005
Earlier work this paper cites.
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020 · 2005
Earlier work this paper cites.
Short-time phase spectrum in speech processing: A review and some experimental results
Leigh D Alsteris and Kuldip K Paliwal. 2007 · 2007
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008 · 2008
Earlier work this paper cites.
Seanet: A multi-modal speech enhancement network
Marco Tagliasacchi, Yunpeng Li, Karolis Misiunas, and Dominik Roblek. 2020 · 2009
Earlier work this paper cites.
A short-time objective intelligibility measure for time-frequency weighted noisy speech
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen. 2010 · 2010
Earlier work this paper cites.
Learn about pearson’s correlation coefficient in spss with data from the global health observatory data (2012)
2015 · 2012
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma. 2014 · 2014
Earlier work this paper cites.
Quesst2014: Evaluating query-by-example speech search in a zero-resource setting with real-life queries
Xavier Anguera, Luis-J Rodriguez-Fuentes, Andi Buzo, Florian Metze, Igor Szöke, and Mikel Penagarikano. 2015 · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
Karol J. Piczak. 2015 · 2015
Earlier work this paper cites.
Neural audio synthesis of musical notes with WaveNet autoencoders
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and Karen Simonyan. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Earlier work this paper cites.
Voxceleb: A large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Speech commands: A public dataset for single-word speech recognition
Pete Warden. 2017 · 2017
Cited alongside, same era.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018 · 2018
Cited alongside, same era.
Introducing Parselmouth: A Python interface to Praat
Yannick Jadoul, Bill Thompson, and Bart de Boer. 2018 · 2018
Cited alongside, same era.
Vocal imitation set: a dataset of vocally imitated sound events using the audioset ontology
Bongjun Kim, Madhav Ghei, Bryan Pardo, and Zhiyao Duan. 2018 · 2018
Cited alongside, same era.
X-vectors: Robust dnn embeddings for speaker recognition
Semi-supervised spoken language understanding via self-supervised speech and language model pretraining
Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, and James R. Glass. 2021 · 2021
Later among the works it cites.
VoxLingua107: a dataset for spoken language recognition
Jörgen Valk and Tanel Alumäe. 2021 · 2021
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021 · 2021
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. 2022 · 2022
Later among the works it cites.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. 2018 · 2018
Cited alongside, same era.
Classification vs. regression in supervised learning for single channel speaker count estimation
Fabian-Robert Stöter, Soumitro Chakrabarty, Bernd Edler, and Emanuël AP Habets. 2018 · 2018
Cited alongside, same era.
Vocalset: A singing voice dataset
Julia Wilkins, Prem Seetharaman, Alison Wahl, and Bryan Pardo. 2018 · 2018
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre De Brebisson, Yoshua Bengio, and Aaron C Courville. 2019 · 2019
Cited alongside, same era.
Speech model pre-training for end-to-end spoken language understanding
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio. 2019 · 2019
Cited alongside, same era.
The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation
Davide Chicco and Giuseppe Jurman. 2020 · 2020
Cited alongside, same era.
Gunshots recorded in an open field using ipod touch devices
Seth Cooper and Steven Shaw. 2020 · 2020
Cited alongside, same era.
FSD50K: an open dataset of human-labeled sound events
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra. 2022 · 2022
Later among the works it cites.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2022 · 2022
Later among the works it cites.
M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus
Lichao Zhang, Ruiqi Li, Shoutong Wang, Liqun Deng, Jinglin Liu, Yi Ren, Jinzheng He, Rongjie Huang, Jieming Zhu, Xiao Chen, and Zhou Zhao. 2022 · 2022
Later among the works it cites.
Musiclm: Generating music from text
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al. 2023 · 2023
Later among the works it cites.
Lauragpt: Listen, attend, understand, and regenerate audio with gpt
Qian Chen, Yunfei Chu, Zhifu Gao, Zerui Li, Kai Hu, Xiaohuan Zhou, Jin Xu, Ziyang Ma, Wen Wang, Siqi Zheng, et al. 2023 · 2023
Later among the works it cites.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2023 · 2023
Later among the works it cites.
Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec
Zhihao Du, Shiliang Zhang, Kai Hu, and Siqi Zheng. 2023 · 2023
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. 2023 · 2023
Later among the works it cites.
Stack-and-delay: a new codebook pattern for music generation
Gael Le Lan, Varun Nagaraja, Ernie Chang, David Kant, Zhaoheng Ni, Yangyang Shi, Forrest Iandola, and Vikas Chandra. 2023 · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Later among the works it cites.
Audiopalm: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. 2023 · 2023
Later among the works it cites.
Zero-shot singing voice synthesis from musical score
Jun-You Wang, Hung-Yi Lee, Jyh-Shing Roger Jang, and Li Su. 2023b · 2023
Later among the works it cites.
Audiodec: An open-source streaming high-fidelity neural audio codec
Yi-Chiao Wu, Israel D Gebru, Dejan Marković, and Alexander Richard. 2023 · 2023
Later among the works it cites.