Fetching the paper…
Reading the bibliography…
This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages.
Towards language independent acoustic modeling
William Byrne, Peter Beyerlein, Juan M. Huerta, Sanjeev Khudanpur, B. Marthi, John Morgan, Nino Peterek, Joseph Picone, Dimitra Vergyri, and T. Wang · 2000
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2006
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, and Faustino Gomez · 2006
Earlier work this paper cites.
Automatic speech recognition for under-resourced languages: Application to vietnamese language
Viet-Bac Le and Laurent Besacier · 2009
Earlier work this paper cites.
Multilingual acoustic modeling for speech recognition based on subspace gaussian mixture models
Lukas Burget, Petr Schwarz, Mohit Agarwal, Pinar Akyazi, Feng Kai, Arnab Ghoshal, Ondrej Glembek, Nagendra Goel, Martin Karafiát, Daniel Povey, Ariya Rastrow, Richard Rose, and Samuel Thomas · 2010
Earlier work this paper cites.
Current trends in multilingual speech processing
Herve Bourlard, John Dines, Mathew Magimai-Doss, Philip Garner, David Imseng, Petr Motlicek, Hui Liang, Lakshmi Saheer, and Fabio Valente · 2011
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid · 2011
Earlier work this paper cites.
Multilingual training of deep neural networks
Arnab Ghoshal, Pawel Swietojanski, and Stephen Renals · 2013
Earlier work this paper cites.
Scalable modified Kneser-Ney language model estimation
Kenneth Heafield, Ivan Pouzyrevsky, Jonathan H. Clark, and Philipp Koehn · 2013
Earlier work this paper cites.
Multilingual acoustic models using distributed deep neural networks
Georg Heigold, V. Vanhoucke, Andrew Senior, Phuongtrang Nguyen, M. Ranzato, M. Devin, and Jeff Dean · 2013
Earlier work this paper cites.
Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers
Jui-Ting Huang, Jinyu Li, Dong Yu, Li Deng, and Yifan Gong · 2013
Earlier work this paper cites.
Speech recognition and keyword spotting for low-resource languages: Babel project research at cued
Mark J. F. Gales, Kate M. Knill, Anton Ragni, and Shakti P. Rath · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
The 2016 bbn georgian telephone speech keyword spotting system
T. Alumäe, D. Karakos, W. Hartmann, R. Hsiao, L. Zhang, L. Nguyen, S. Tsakalidis, and R. Schwartz · 2017
Earlier work this paper cites.
Multilingually trained bottleneck features in spoken language recognition
Radek Fer, Pavel Matějka, František Grézl, Oldřich Plchot, Karel Veselỳ, and Jan Honza Černockỳ · 2017
Earlier work this paper cites.
Low-resource speech recognition and keyword-spotting
Mark JF Gales, Kate M Knill, and Anton Ragni · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Multilingual sequence-to-sequence speech recognition: Architecture, transfer learning, and language modeling
J. Cho, M. K. Baskar, R. Li, M. Wiesner, S. H. Mallidi, N. Yalta, M. Karafiát, S. Watanabe, and T. Hori · 2018
Cited alongside, same era.
Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech
Yu-An Chung and James Glass · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2019
Later among the works it cites.
Transfer learning of language-independent end-to-end asr with language model fusion
Hirofumi Inaguma, Jaejin Cho, Murali Karthick Baskar, Tatsuya Kawahara, and Shinji Watanabe · 2019
Later among the works it cites.
Improving transformer-based speech recognition using unsupervised pre-training
Dongwei Jiang, Xiaoning Lei, Wubo Li, Ne Luo, Yuxuan Hu, Wei Zou, and Xiangang Li · 2019
Later among the works it cites.
Large-scale multilingual speech recognition with a streaming end-to-end model
Anjuli Kannan, Arindrima Datta, Tara N. Sainath, Eugene Weinstein, Bhuvana Ramabhadran, Yonghui Wu, Ankur Bapna, Zhifeng Chen, and Seungji Lee · 2019
Later among the works it cites.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The challenge of realistic music generation: modelling raw audio at scale
Sander Dieleman, Aäron van den Oord, and Karen Simonyan · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Confidence estimation and deletion prediction using bidirectional recurrent neural networks
Anton Ragni, Qiujia Li, Mark J. F. Gales, and Yu Wang · 2018
Cited alongside, same era.
An end-to-end language-tracking speech recognizer for mixed-language speech
H. Seki, S. Watanabe, T. Hori, J. L. Roux, and J. R. Hershey · 2018
Cited alongside, same era.
Multilingual speech recognition with a single end-to-end model
Shubham Toshniwal, Tara N. Sainath, Ron J. Weiss, Bo Li, Pedro Moreno, Eugene Weinstein, and Kanishka Rao · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M. Tyers, and Gregor Weber · 2019
Cited alongside, same era.
Later among the works it cites.
Transformers with convolutional context for ASR
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Later among the works it cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Later among the works it cites.
Contrastive multiview coding
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Later among the works it cites.
Vqvae unsupervised unit discovery and multi-scale code2spec inverter for zerospeech challenge 2019
Andros Tjandra, Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, and Satoshi Nakamura · 2019
Later among the works it cites.
Ccnet: Extracting high quality monolingual datasets from web crawl data, 2019
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave · 2019
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Learning hierarchical discrete linguistic units from visually-grounded speech
David Harwath, Wei-Ning Hsu, and James Glass · 2020
Closest in time.
Learning robust and multilingual speech representations
Kazuya Kawakami, Luyu Wang, Chris Dyer, Phil Blunsom, and Aaron van den Oord · 2020
Closest in time.
Mls: A large-scale multilingual dataset for speech research
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert · 2020
Closest in time.
Unsupervised pretraining transfers well across languages
Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, and Emmanuel Dupoux · 2020
Closest in time.