Fetching the paper…
Reading the bibliography…
The large amount of audiovisual content being shared online today has drawn substantial attention to the prospect of audiovisual self-supervised learning.
“Self-Supervised Learning by Cross-Modal Audio-Video Clustering”
Humam Alwassel et al · 1911
Earlier work this paper cites.
Soo-Whan Chung et al · 2004
Earlier work this paper cites.
“Does Visual Self-Supervision Improve Learning of Speech Representations?”
Abhinav Shukla et al · 2005
Earlier work this paper cites.
“Dlib-ml: A Machine Learning Toolkit”
Davis. King · 2009
Earlier work this paper cites.
“Prediction-based classification for audiovisual discrimination between laughter and speech”
S. Petridis et al · 2011
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books”
Vassil Panayotov et al · 2015
Earlier work this paper cites.
“Prediction-Based Audiovisual Fusion for Classification of Non-Linguistic Vocalisations”
Stavros Petridis and Maja Pantic · 2015
Earlier work this paper cites.
“Lip Reading in the Wild”
Joon Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
“Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles”
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
“Colorful Image Colorization”
Richard Zhang et al · 2016
Earlier work this paper cites.
“Look, Listen and Learn”
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
“Combining Residual Networks with LSTMs for Lipreading”
Themos Stafylakis and Georgios Tzimiropoulos · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Deep Audio-Visual Speech Recognition”
T. Afouras et al · 2018
Earlier work this paper cites.
“LRS3-TED: a large-scale dataset for visual speech recognition”
T. Afouras et al · 2018
Earlier work this paper cites.
“Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization”
Bruno Korbar et al · 2018
Cited alongside, same era.
“Online Hybrid CTC/Attention Architecture for End-to-End Speech Recognition”
Haoran Miao et al · 2018
Cited alongside, same era.
“Representation Learning with Contrastive Predictive Coding”
A“”aron van Oord et al · 2018
Cited alongside, same era.
“Audio-Visual Scene Analysis with Self-Supervised Multisensory Features”
Andrew Owens and Alexei. Efros · 2018
Cited alongside, same era.
“Learning Sight from Sound: Ambient Sound Provides Supervision for Visual Learning”
Andrew Owens et al · 2018
Cited alongside, same era.
“Audio-Visual Speech Recognition with a Hybrid CTC/Attention Architecture”
“wav2vec: Unsupervised Pre-Training for Speech Recognition”
Steffen Schneider et al · 2019
Later among the works it cites.
“Learning Spatio-Temporal Features with Two-Stream Deep 3D CNNs for Lipreading”
Xinshuo Weng and Kris Kitani · 2019
Later among the works it cites.
“Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading”
Xingxuan Zhang et al · 2019
Later among the works it cites.
“ASR is All You Need: Cross-Modal Distillation for Lip Reading”
Triantafyllos Afouras et al · 2020
Later among the works it cites.
“Perfect Match: Self-Supervised Embeddings for Cross-Modal Retrieval”
Soo-Whan Chung et al · 2020
Later among the works it cites.
“Conformer: Convolution-augmented Transformer for Speech Recognition”
A. Gulati et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stavros Petridis et al · 2018
Cited alongside, same era.
“Pushing the boundaries of audiovisual word recognition using Residual Networks and LSTMs”
Themos Stafylakis et al · 2018
Cited alongside, same era.
“Transformer-XL: Attentive Language Models beyond a Fixed-Length Context”
Zihang Dai et al · 2019
Cited alongside, same era.
“Revisiting Self-Supervised Visual Representation Learning”
Alexander Kolesnikov et al · 2019
Cited alongside, same era.
“Decoupled Weight Decay Regularization”
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
“Recurrent Neural Network Transducer for Audio-Visual Speech Recognition”
Takaki Makino et al · 2019
Cited alongside, same era.
“Learning Problem-Agnostic Speech Representations from Multiple Self-Supervised Tasks”
Santiago Pascual et al · 2019
Cited alongside, same era.
“Lipreading Using Temporal Convolutional Networks”
Brais Mart“’nez et al · 2020
Later among the works it cites.
“Evolving Losses for Unsupervised Video Representation Learning”
A.. Piergiovanni et al · 2020
Later among the works it cites.
“Multi-Task Self-Supervised Learning for Robust Speech Recognition”
Mirco Ravanelli et al · 2020
Later among the works it cites.
“Audio-Visual Recognition of Overlapped Speech for the LRS2 Dataset”
J. Yu et al · 2020
Later among the works it cites.
“Towards Practical Lipreading with Distilled and Efficient Models”
Pingchuan Ma et al · 2021
Closest in time.
“End-To-End Audio-Visual Speech Recognition with Conformers”
Pingchuan Ma et al · 2021
Closest in time.
“Lip-reading with Densely Connected Temporal Convolutional Networks”
Pingchuan Ma et al · 2021
Closest in time.
“Audio-Visual Predictive Coding for Self-Supervised Visual Representation Learning”
Mani Tellamekala et al · 2021
Closest in time.