Fetching the paper…
Reading the bibliography…
Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR).
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Deep speaker embeddings for short-duration speaker verification.,”
Gautam Bhattacharya, Md Jahangir Alam, and Patrick Kenny, · 2017
Earlier work this paper cites.
“Voxceleb: A large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Earlier work this paper cites.
“Evaluating deep learning architectures for speech emotion recognition,”
Haytham M Fayek, Margaret Lech, and Lawrence Cavedon, · 2017
Earlier work this paper cites.
“An attention pooling based representation learning method for speech emotion recognition,”
Pengcheng Li, Yan Song, Ian Vince McLoughlin, Wu Guo, and Lirong Dai, · 2018
Earlier work this paper cites.
“Towards end-to-end spoken language understanding,”
Dmitriy Serdyuk, Yongqiang Wang, Christian Fuegen, Anuj Kumar, Baiyang Liu, and Yoshua Bengio, · 2018
Earlier work this paper cites.
Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al., · 2018
Earlier work this paper cites.
“An attention pooling based representation learning method for speech emotion recognition,”
Pengcheng Li, Yan Song, Ian Vince McLoughlin, Wu Guo, and Li-Rong Dai, · 2018
Earlier work this paper cites.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Earlier work this paper cites.
“wav2vec: Unsupervised Pre-Training for Speech Recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Earlier work this paper cites.
“Speech emotion recognition using capsule networks,”
Xixin Wu, Songxiang Liu, Yuewen Cao, Xu Li, Jianwei Yu, Dongyang Dai, Xi Ma, Shoukang Hu, Zhiyong Wu, Xunying Liu, et al., · 2019
Cited alongside, same era.
“Recent advances in end-to-end spoken language understanding,”
Natalia Tomashenko, Antoine Caubrière, Yannick Estève, Antoine Laurent, and Emmanuel Morin, · 2019
Cited alongside, same era.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee, · 2020
Cited alongside, same era.
“Multi-task self-supervised learning for robust speech recognition,”
Mirco Ravanelli, Jianyuan Zhong, Santiago Pascual, Pawel Swietojanski, Joao Monteiro, Jan Trmal, and Yoshua Bengio, · 2020
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.
“Superb: Speech processing universal performance benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Lai, Kushal Lakhotia, Yist Lin, Andy Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung-yi Lee, · 2021
Closest in time.
“Emotion Recognition from Speech Using wav2vec 2.0 Embeddings,”
Leonardo Pepino, Pablo Riera, and Luciana Ferrer, · 2021
Closest in time.
“Temporal context in speech emotion recognition,”
Yangyang Xia, Li-Wei Chen, Alexander Rudnicky, and Richard M Stern, · 2021
Closest in time.
“Exploring wav2vec 2.0 on Speaker Verification and Language Identification,”
Zhiyun Fan, Meng Li, Shiyu Zhou, and Bo Xu, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
A. Baevski, S. Schneider, and M. Auli, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification,”
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck, · 2020
Cited alongside, same era.
“Slurp: A spoken language understanding resource package,”
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
Cited alongside, same era.
“Non-Autoregressive Predictive Coding for Learning Speech Representations from Local Dependencies,”
Alexander H. Liu, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“Tera: Self-supervised learning of transformer encoder representation for speech,”
Andy T Liu, Shang-Wen Li, and Hung-yi Lee, · 2021
Cited alongside, same era.
“Speechbrain: A general-purpose speech toolkit,”
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, et al., · 2021
Closest in time.
“Timers and Such: A Practical Benchmark for Spoken Language Understanding with Numbers,”
Loren Lugosch, Piyush Papreja, Mirco Ravanelli, Abdelwahab Heba, and Titouan Parcollet, · 2021
Closest in time.
“Head fusion: Improving the accuracy and robustness of speech emotion recognition on the iemocap and ravdess dataset,”
Mingke Xu, Fan Zhang, and Wei Zhang, · 2021
Closest in time.
“Siamese capsule network for end-to-end speaker recognition in the wild,”
Amirhossein Hajavi and Ali Etemad, · 2021
Closest in time.
Seunghyun Seo, Donghyun Kwak, and Bowon Lee, · 2021
Closest in time.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Closest in time.
“Fine-tuning wav2vec2 for speaker recognition,”
Nik Vaessen and David A Van Leeuwen, · 2022
Closest in time.