Fetching the paper…
Reading the bibliography…
Recently proposed self-supervised learning approaches have been successful for pre-training speech representation models.
“Relations between two sets of variates,”
Harold Hotelling, · 1936
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Rapid evaluation of speech representations for spoken term discovery,”
Michael A Carlin, Samuel Thomas, Aren Jansen, and Hynek Hermansky, · 2011
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality,”
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean, · 2013
Earlier work this paper cites.
“GloVe: Global vectors for word representation,”
Jeffrey Pennington, Richard Socher, and Christopher D Manning, · 2014
Earlier work this paper cites.
“Community evaluation and exchange of word vectors at wordvectors. org,”
Manaal Faruqui and Chris Dyer, · 2014
Earlier work this paper cites.
“Unsupervised visual representation learning by context prediction,”
Carl Doersch, Abhinav Gupta, and Alexei A Efros, · 2015
Earlier work this paper cites.
“Unsupervised neural network based feature extraction using weak top-down constraints,”
Herman Kamper, Micha Elsner, Aren Jansen, and Sharon Goldwater, · 2015
Earlier work this paper cites.
“LibriSpeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Representations of language in a model of visually grounded speech signal,”
Grzegorz Chrupała, Lieke Gelderloos, and Afra Alishahi, · 2017
Earlier work this paper cites.
“SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability,”
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein, · 2017
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi.,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Earlier work this paper cites.
“Improving language understanding by generative pre-training,”
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Insights on representational similarity in neural networks with canonical correlation,”
Ari S Morcos, Maithra Raghu, and Samy Bengio, · 2018
Earlier work this paper cites.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Yu-An Chung and James Glass, · 2018
Earlier work this paper cites.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
Santiago Pascual, Mirco Ravanelli, Joan Serra, Antonio Bonafonte, and Yoshua Bengio, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“Analysis methods in neural language processing: A survey,”
Yonatan Belinkov and James Glass, · 2019
Cited alongside, same era.
“Learned in speech recognition: Contextual acoustic word embeddings,”
Shruti Palaskar, Vikas Raunak, and Florian Metze, · 2019
Cited alongside, same era.
“BERT rediscovers the classical NLP pipeline,”
Ian Tenney, Dipanjan Das, and Ellie Pavlick, · 2019
Cited alongside, same era.
“The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives,”
“How accents confound: Probing for accent information in end-to-end speech recognition systems,”
Archiki Prasad and Preethi Jyothi, · 2020
Later among the works it cites.
“The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,”
Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Evgeny Kharitonov, Alexei Baevski, Ewan Dunbar, and Emmanuel Dupoux, · 2020
Later among the works it cites.
“Multilingual jointly trained acoustic and written word embeddings,”
Yushi Hu, Shane Settle, and Karen Livescu, · 2020
Later among the works it cites.
“Evaluating the reliability of acoustic speech embeddings,”
Robin Algayres, Mohamed Zaiem, Benoît Sagot, and Emmanuel Dupoux, · 2020
Later among the works it cites.
“Libri-light: A benchmark for ASR with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elena Voita, Rico Sennrich, and Ivan Titov, · 2019
Cited alongside, same era.
“Similarity of neural network representations revisited,”
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton, · 2019
Cited alongside, same era.
“Acoustically grounded word embeddings for improved acoustics-to-word speech recognition,”
Shane Settle, Kartik Audhkhasi, Karen Livescu, and Michael Picheny, · 2019
Cited alongside, same era.
“Speech model pre-training for end-to-end spoken language understanding,”
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio, · 2019
Cited alongside, same era.
“How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings,”
Kawin Ethayarajh, · 2019
Cited alongside, same era.
“Generative pre-training for speech with autoregressive predictive coding,”
Yu-An Chung and James Glass, · 2020
Cited alongside, same era.
“Unsupervised pre-training of bidirectional speech encoders via masked reconstruction,”
Weiran Wang, Qingming Tang, and Karen Livescu, · 2020
Cited alongside, same era.
“Revisiting few-sample BERT fine-tuning,”
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi, · 2020
Later among the works it cites.
“HuBERT: How much can a bad teacher benefit ASR pre-training?,”
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.
“SUPERB: Speech processing universal PERformance benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al., · 2021
Closest in time.
“Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,”
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, et al., · 2021
Closest in time.
“Unsupervised speech recognition,”
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau, and Michael Auli, · 2021
Closest in time.
“Large-scale self-and semi-supervised learning for speech translation,”
Changhan Wang, Anne Wu, Juan Pino, Alexei Baevski, Michael Auli, and Alexis Conneau, · 2021
Closest in time.
“Probing acoustic representations for phonetic properties,”
Danni Ma, Neville Ryant, and Mark Liberman, · 2021
Closest in time.
Jui Shah, Yaman Kumar Singla, Changyou Chen, and Rajiv Ratn Shah, · 2021
Closest in time.
“Similarity analysis of self-supervised speech representations,”
Yu-An Chung, Yonatan Belinkov, and James Glass, · 2021
Closest in time.
“Acoustic word embeddings for zero-resource languages using self-supervised contrastive learning and multilingual adaptation,”
Christiaan Jacobs, Yevgen Matusevych, and Herman Kamper, · 2021
Closest in time.
“SpeechBrain: A general-purpose speech toolkit,”
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, et al., · 2021
Closest in time.