Fetching the paper…
Reading the bibliography…
Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks.
“Statistical theory of quantization,”
Bernard Widrow, István Kollár, and Ming-Chang Liu, · 1996
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean, · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Rethinking the inception architecture for computer vision,”
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Subword regularization: Improving neural network translation models with multiple subword candidates,”
Taku Kudo, · 2018
Earlier work this paper cites.
“Effectiveness of self-supervised pre-training for speech recognition,”
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed, · 2019
Earlier work this paper cites.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, · 2019
Cited alongside, same era.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2020
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“E-branchformer: Branchformer with enhanced merging for speech recognition,”
Kwangyoun Kim, Felix Wu, Yifan Peng, Jing Pan, Prashant Sridhar, Kyu Jeong Han, and Shinji Watanabe, · 2022
Cited alongside, same era.
Rohan Anil, Andrew M. Dai, Orhan Firat, and et al., · 2023
Closest in time.
“Llama: Open and efficient foundation language models,”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample, · 2023
Closest in time.
“Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,”
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu, · 2023
Closest in time.
“Viola: Unified codec language models for speech recognition, synthesis, and translation,”
Tianrui Wang, Long Zhou, Ziqiang Zhang, Yu Wu, Shujie Liu, Yashesh Gaur, Zhuo Chen, Jinyu Li, and Furu Wei, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei, · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
OpenAI, · 2023
Cited alongside, same era.
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, and et al., · 2023
Closest in time.
“Lauragpt: Listen, attend, understand, and regenerate audio with GPT,”
Jiaming Wang, Zhihao Du, Qian Chen, Yunfei Chu, Zhifu Gao, Zerui Li, Kai Hu, Xiaohuan Zhou, Jin Xu, Ziyang Ma, Wen Wang, Siqi Zheng, Chang Zhou, Zhijie Yan, and Shiliang Zhang, · 2023
Closest in time.
“Exploration of efficient end-to-end ASR using discretized input from self-supervised learning,”
Xuankai Chang, Brian Yan, Yuya Fujita, Takashi Maekaku, and Shinji Watanabe, · 2023
Closest in time.
“Google USM: scaling automatic speech recognition beyond 100 languages,”
Yu Zhang, Wei Han, James Qin, and et al., · 2023
Closest in time.