Fetching the paper…
Reading the bibliography…
In this paper, we describe our submissions to the ZeroSpeech 2021 Challenge and SUPERB benchmark.
Problem-Agnostic Speech Embeddings for Multi-Speaker Text-to-Speech with SampleRNN
Álvarez, D.; Pascual, S.; and Bonafonte, A. 2019 · 1906
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
Baevski, A.; Schneider, S.; and Auli, M. 2020 · 1910
Earlier work this paper cites.
Multi-task self-supervised learning for Robust Speech Recognition
Ravanelli, M.; Zhong, J.; Pascual, S.; Swietojanski, P.; Monteiro, J.; Trmal, J.; and Bengio, Y. 2020 · 2001
Earlier work this paper cites.
Vector-Quantized Autoregressive Predictive Coding
Chung, Y.-A.; Tang, H.; and Glass, J. 2020 · 2005
Earlier work this paper cites.
Lai, C.-I.; Chuang, Y.-S.; Lee, H.-Y.; Li, S.-W.; and Glass, J. 2020 · 2010
Earlier work this paper cites.
Lin, Y. Y.; Chien, C.-M.; Lin, J.-H.; Lee, H.-Y.; and Lee, L.-S. 2021 · 2010
Earlier work this paper cites.
Exploring wav2vec 2.0 on speaker verification and language identification
Fan, Z.; Li, M.; Zhou, S.; and Xu, B. 2021 · 2012
Earlier work this paper cites.
DeCoAR 2.0: Deep Contextualized Acoustic Representations with Vector Quantization
Ling, S.; and Liu, Y. 2020 · 2012
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.-Y.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Learning words from images and speech
Synnaeve, G.; Versteegh, M.; and Dupoux, E. 2014 · 2014
Earlier work this paper cites.
Deep multimodal semantic embeddings for speech and images
Harwath, D. F.; and Glass, J. R. 2015 · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Panayotov, V.; Chen, G.; Povey, D.; and Khudanpur, S. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Unsupervised Learning of Spoken Language with Visual Context
Harwath, D. F.; Torralba, A.; and Glass, J. R. 2016 · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Hendrycks, D.; and Gimpel, K. 2016 · 2016
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; et al. 2016 · 2016
Earlier work this paper cites.
Look, Listen, and Decode: Multimodal Speech Recognition with Images
Sun, F.; Harwath, D.; and Glass, J. 2016 · 2016
Earlier work this paper cites.
Representations of language in a model of visually grounded speech signal
Chrupała, G.; Gelderloos, L.; and Alishahi, A. 2017 · 2017
Earlier work this paper cites.
Learning Word-Like Units from Joint Audio-Visual Analysis
Harwath, D. F.; and Glass, J. R. 2017 · 2017
Earlier work this paper cites.
Visually Grounded Learning of Keyword Prediction from Untranscribed Speech
Kamper, H.; Settle, S.; Shakhnarovich, G.; and Livescu, K. 2017 · 2017
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Karpathy, A.; and Fei-Fei, L. 2017 · 2017
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Vision as an Interlingua: Learning Multilingual Semantic Embeddings of Untranscribed Speech
Harwath, D. F.; Chuang, G.; and Glass, J. R. 2018 · 2018
Cited alongside, same era.
Visually grounded cross-lingual keyword spotting in speech
Kamper, H.; and Roth, M. 2018 · 2018
Cited alongside, same era.
End-to-End Multimodal Speech Recognition
Palaskar, S.; Sanabria, R.; and Metze, F. 2018 · 2018
Cited alongside, same era.
Representation Learning with Contrastive Predictive Coding
v. d. Oord, A.; Li, Y.; and Vinyals, O. 2018 · 2018
Libri-Light: A Benchmark for ASR with Limited or No Supervision
Kahn, J.; Riviere, M.; Zheng, W.; Kharitonov, E.; Xu, Q.; Mazare, P.; Karadayi, J.; Liptchinsky, V.; Collobert, R.; Fuegen, C.; and et al. 2020 · 2020
Later among the works it cites.
Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
Liu, A. T.; Yang, S.-W.; Chi, P.-H.; Hsu, P.-C.; and Lee, H.-Y. 2020 · 2020
Later among the works it cites.
Speech-Image Semantic Alignment Does Not Depend on Any Prior Classification Tasks
Mortazavi, M. 2020 · 2020
Later among the works it cites.
Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms
Ohishi, Y.; Kimura, A.; Kawanishi, T.; Kashino, K.; Harwath, D. F.; and Glass, J. 2020 · 2020
Later among the works it cites.
A DNN-HMM-DNN Hybrid Model for Discovering Word-Like Units from Spoken Captions and Image Regions
Wang, L.; and Hasegawa-Johnson, M. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Symbolic Inductive Bias for Visually Grounded Learning of Spoken Language
Chrupała, G. 2019 · 2019
Cited alongside, same era.
An Unsupervised Autoregressive Model for Speech Representation Learning
Chung, Y.-A.; Hsu, W.-N.; Tang, H.; and Glass, J. R. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Towards Visually Grounded Sub-word Speech Unit Discovery
Harwath, D. F.; and Glass, J. R. 2019 · 2019
Cited alongside, same era.
Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input
Harwath, D. F.; Recasens, A.; Surís, D.; Chuang, G.; Torralba, A.; and Glass, J. 2019 · 2019
Cited alongside, same era.
Transfer Learning from Audio-Visual Grounding to Speech Recognition
Hsu, W.-N.; Harwath, D. F.; and Glass, J. 2019 · 2019
Cited alongside, same era.
Alishahi, A.; et al. 2021 · 2021
Later among the works it cites.
DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT
Chang, H.-J.; Yang, S.-W.; and Lee, H.-Y. 2021 · 2021
Later among the works it cites.
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
Chen, S.; Wang, C.; Chen, Z.; Wu, Y.; Liu, S.; Chen, Z.; Li, J.; Kanda, N.; Yoshioka, T.; Xiao, X.; Wu, J.; Zhou, L.; Ren, S.; Qian, Y.; Qian, Y.; Wu, J.; Zeng, M.; and Wei, F. 2021 · 2021
Later among the works it cites.
XLM-E: Cross-lingual Language Model Pre-training via ELECTRA
Chi, Z.; Huang, S.; Dong, L.; Ma, S.; Singhal, S.; Bajaj, P.; Song, X.; and Wei, F. 2021 · 2021
Later among the works it cites.
Information Retrieval for ZeroSpeech 2021: The Submission by University of Wroclaw
Chorowski, J.; et al. 2021 · 2021
Later among the works it cites.
Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques
Chrupała, G. 2021 · 2021
Later among the works it cites.
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
Khorrami, K.; and Räsänen, O. 2021 · 2021
Later among the works it cites.
TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
Liu, A. T.; Li, S.-W.; and Lee, H.-Y. 2021 · 2021
Later among the works it cites.
Maekaku, T.; Chang, X.; Fujita, Y.; Chen, L.-W.; Watanabe, S.; and Rudnicky, A. I. 2021 · 2021
Later among the works it cites.
The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling
Nguyen, T. A.; de Seyssel, M.; Rozé, P.; Rivière, M.; Kharitonov, E.; Baevski, A.; Dunbar, E.; and Dupoux, E. 2020 · 2021
Later among the works it cites.
Attention-Based Keyword Localisation in Speech using Visual Grounding
Olaleye, K.; and Kamper, H. 2021 · 2021
Later among the works it cites.
Fast-Slow Transformer for Visually Grounding Speech
Peng, P.; and Harwath, D. 2021 · 2021
Later among the works it cites.
Talk, Don’t Write: A Study of Direct Speech-Based Image Retrieval
Sanabria, R.; Waters, A.; and Baldridge, J. 2021 · 2021
Later among the works it cites.
Analyzing speaker information in self-supervised models to improve zero-resource speech processing
v. Niekerk, B.; Nortje, L.; Baas, M.; and Kamper, H. 2021 · 2021
Later among the works it cites.
Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval
Wang, L.; Wang, X.; Hasegawa-Johnson, M.; Scharenborg, O.; and Dehak, N. 2021 · 2021
Later among the works it cites.
SUPERB: Speech processing Universal PERformance Benchmark
Yang, S.-W.; Chi, P.-H.; Chuang, Y.-S.; Lai, C.-I. J.; Lakhotia, K.; Lin, Y. Y.; Liu, A. T.; Shi, J.; Chang, X.; Lin, G.-T.; Huang, T.-H.; Tseng, W.-C.; tik Lee, K.; Liu, D.-R.; Huang, Z.; Dong, S.; Li, S.-W.; Watanabe, S.; Mohamed, A.; and Lee, H.-Y. 2021 · 2021
Later among the works it cites.