Fetching the paper…
Reading the bibliography…
Discrete audio representation, aka audio tokenization, has seen renewed interest driven by its potential to facilitate the application of text language modeling approaches in audio domain.
“Callhome american english speech,”
Alexandra Canavan, David Graff, and George Zipperlen, · 1997
Earlier work this paper cites.
“The nist speaker recognition evaluations: 1996-2001,”
Alvin Martin and Mark Przybocki, · 2001
Earlier work this paper cites.
“The ami meeting corpus: A pre-announcement,”
Jean Carletta, Simone Ashby, et al., · 2005
Earlier work this paper cites.
“Fisher spanish speech (ldc2010s01),”
David Graff, Shudong Huang, et al., · 2010
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” 2018
Taku Kudo and John Richardson, · 2018
Earlier work this paper cites.
“VoxCeleb2: Deep speaker recognition,”
J.S. Chung, A. Nagrani, and A. Zisserman, · 2018
Earlier work this paper cites.
“Arcface: Additive angular margin loss for deep face recognition,”
Jiankang Deng, Jia Guo, et al., · 2019
Earlier work this paper cites.
“Common voice: A massively-multilingual speech corpus,”
Rosana Ardila, Megan Branson, et al., · 2019
Earlier work this paper cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, et al., · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown et al., · 2020
Earlier work this paper cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, et al., · 2020
Cited alongside, same era.
“Mls: A large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, et al., · 2020
Cited alongside, same era.
“Pyannote. audio: neural building blocks for speaker diarization,”
Hervé Bredin, Ruiqing Yin, et al., · 2020
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, et al., · 2021
Cited alongside, same era.
“W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,”
Yu-An Chung, Yu Zhang, et al., · 2021
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models,” 2023
Hugo Touvron et al., · 2023
Closest in time.
“Google usm: Scaling automatic speech recognition beyond 100 languages,”
Yu Zhang, Wei Han, et al., · 2023
Closest in time.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, et al., · 2023
Closest in time.
“Speak foreign languages with your own voice: Cross-lingual neural codec language modeling,”
Ziqiang Zhang et al., · 2023
Closest in time.
“Viola: Unified codec language models for speech recognition, synthesis, and translation,”
Tianrui Wang, Long Zhou, et al., · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neil Zeghidour, Alejandro Luebs, et al., · 2021
Cited alongside, same era.
Changhan Wang, Morgane Riviere, et al., · 2021
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,” 2022
Alec Radford, Jong Wook Kim, et al., · 2022
Cited alongside, same era.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, et al., · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, et al., · 2022
Cited alongside, same era.
“Titanet: Neural model for speaker representation with 1d depth-wise separable convolutions and global context,”
Nithin Rao Koluguri, Taejin Park, and Boris Ginsburg, · 2022
Cited alongside, same era.
“NeMo: a toolkit for Conversational AI and Large Language Models,”
Eric Harper, Somshubra Majumdar, et al.,
Cited in the paper.
Closest in time.
“Musiclm: Generating music from text,”
Andrea Agostinelli, Timo I Denk, et al., · 2023
Closest in time.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, et al., · 2023
Closest in time.
“Audiopalm: A large language model that can speak and listen,”
Paul K Rubenstein, Chulayuth Asawaroengchai, et al., · 2023
Closest in time.
“High-fidelity audio compression with improved rvqgan,”
Rithesh Kumar, Prem Seetharaman, et al., · 2023
Closest in time.
“Fast conformer with linearly scalable attention for efficient speech recognition,”
Dima Rekesh, Nithin Rao Koluguri, et al., · 2023
Closest in time.