Fetching the paper…
Reading the bibliography…
Transfer learning has proven to be crucial in advancing the state of speech and natural language processing research in recent years.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 1905
Earlier work this paper cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 1906
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Wham!: Extending speech separation to noisy environments
Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, and Jonathan Le Roux. 2019 · 1907
Earlier work this paper cites.
Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen. 2017 · 1913
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Librimix: An open-source dataset for generalizable speech separation
Joris Cosentino, Manuel Pariente, Samuele Cornell, Antoine Deleforge, and Emmanuel Vincent. 2020 · 2005
Earlier work this paper cites.
Santa Barbara corpus of spoken American English
John W Du Bois, Wallace L Chafe, Charles Meyer, Sandra A Thompson, and Nii Martey. 2000 – 2005 · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Tera: Self-supervised learning of transformer encoder representation for speech
Andy T Liu, Shang-Wen Li, and Hung-yi Lee. 2020b · 2007
Earlier work this paper cites.
Yist Y Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-yi Lee, and Lin-shan Lee. 2020 · 2010
Earlier work this paper cites.
Non-autoregressive predictive coding for learning speech representations from local dependencies
Alexander H Liu, Yu-An Chung, and James Glass. 2020a · 2011
Earlier work this paper cites.
Exploring wav2vec 2.0 on speaker verification and language identification
Zhiyun Fan, Meng Li, Shiyu Zhou, and Bo Xu. 2020 · 2012
Earlier work this paper cites.
DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization
Shaoshi Ling and Yuzong Liu. 2020 · 2012
Earlier work this paper cites.
The voice bank corpus: Design, collection and data analysis of a large regional accent speech database
Christophe Veaux, Junichi Yamagishi, and Simon King. 2013 · 2013
Earlier work this paper cites.
Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks
Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux. 2015 · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. 2017 · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Permutation invariant training of deep models for speaker-independent multi-talker speech separation
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen. 2017 · 2017
Cited alongside, same era.
Exploring the limits of weakly supervised pretraining
Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten. 2018 · 2018
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020 · 2020
Later among the works it cites.
Vector-quantized autoregressive predictive coding
Yu-An Chung, Hao Tang, and James Glass. 2020 · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
J. Kong, J. Kim, and J. Bae. 2020 · 2020
Later among the works it cites.
Deep contextualized acoustic representations for semi-supervised speech recognition
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Natural TTS Synthesis by Conditioning WaveNet on MEL Spectrogram Predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu. 2018 · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
Supervised speech separation based on deep learning: An overview
DeLiang Wang and Jitong Chen. 2018 · 2018
Cited alongside, same era.
Problem-agnostic speech embeddings for multi-speaker text-to-speech with samplernn
David Álvarez et al. 2019 · 2019
Cited alongside, same era.
An Unsupervised Autoregressive Model for Speech Representation Learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019 · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Multi-task self-supervised learning for robust speech recognition
Mirco Ravanelli, Jianyuan Zhong, Santiago Pascual, Pawel Swietojanski, Joao Monteiro, Jan Trmal, and Yoshua Bengio. 2020 · 2020
Later among the works it cites.
Unsupervised pretraining transfers well across languages
Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, and Emmanuel Dupoux. 2020 · 2020
Later among the works it cites.
CoVoST 2: A massively multilingual speech-to-text translation corpus
Changhan Wang, Anne Wu, and Juan Pino. 2020 · 2020
Later among the works it cites.
Voice Conversion Challenge 2020 - Intra-lingual semi-parallel and cross-lingual voice conversion -
Y. Zhao, W.-C. Huang, X. Tian, J. Yamagishi, R. K. Das, T. Kinnunen, Z. Ling, and T. Toda. 2020 · 2020
Later among the works it cites.
DistilHuBERT: Speech representation learning by layer-wise distillation of hidden-unit BERT
Heng-Jui Chang, Shu-wen Yang, and Hung-yi Lee. 2021 · 2021
Later among the works it cites.
Solene Evain, Ha Nguyen, Hang Le, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Esteve, Benjamin Lecouteux, Francois Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, and Laurent Besacier. 2021 · 2021
Later among the works it cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Later among the works it cites.
Semi-supervised spoken language understanding via self-supervised speech and language model pretraining
Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, and James Glass. 2021 · 2021
Later among the works it cites.
On the use of self-supervised pre-trained acoustic and linguistic features for continuous speech emotion recognition
Manon Macary, Marie Tahon, Yannick Estève, and Anthony Rousseau. 2021 · 2021
Later among the works it cites.
The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling
Tu Anh Nguyen et al. 2020 · 2021
Later among the works it cites.
SUPERB: Speech processing universal performance benchmark
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung yi Lee. 2021 · 2021
Later among the works it cites.