Fetching the paper…
Reading the bibliography…
Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information.
“Hearing lips and seeing voices,”
Harry McGurk and John MacDonald, · 1976
Earlier work this paper cites.
“Recent advances in the automatic recognition of audiovisual speech,”
G. Potamianos et al., · 2003
Earlier work this paper cites.
“Is neocortex essentially multisensory?,”
Asif A. Ghazanfar and Charles E. Schroeder, · 2006
Earlier work this paper cites.
“An audio-visual corpus for speech perception and automatic speech recognition,”
Martin Cooke et al., · 2006
Earlier work this paper cites.
“Ucf101: A dataset of 101 human actions classes from videos in the wild,”
Khurram Soomro et al., · 2012
Earlier work this paper cites.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar et al., · 2016
Earlier work this paper cites.
“Look, listen and learn,”
Relja Arandjelovic and Andrew Zisserman, · 2017
Earlier work this paper cites.
“The kinetics human action video dataset,”
Will Kay et al., · 2017
Earlier work this paper cites.
“Lip reading sentences in the wild,”
Joon Son Chung et al., · 2017
Earlier work this paper cites.
“Voxceleb: a large-scale speaker identification dataset,”
A. Nagrani et al., · 2017
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F. Gemmeke et al., · 2017
Earlier work this paper cites.
“Cooperative learning of audio and video models from self-supervised synchronization,”
Bruno Korbar et al., · 2018
Earlier work this paper cites.
“Lrs3-ted: a large-scale dataset for visual speech recognition,”
Triantafyllos Afouras et al., · 2018
Earlier work this paper cites.
“Voxceleb2: Deep speaker recognition,”
Joon Son Chung et al., · 2018
Earlier work this paper cites.
“Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks,”
Jason Phang et al., · 2018
Earlier work this paper cites.
“Audiovisual emotion recognition in wild,”
Egils Avots et al., · 2019
Cited alongside, same era.
“Can you tell me how to get past sesame street? sentence-level pretraining beyond language modeling,”
Alex Wang et al., · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski et al., · 2020
Cited alongside, same era.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy et al., · 2020
Cited alongside, same era.
“Self-supervised learning by cross-modal audio-video clustering,”
Humam Alwassel et al., · 2020
Cited alongside, same era.
“Disentangled speech embeddings using cross-modal self-supervision,”
Arsha Nagrani et al., · 2020
Cited alongside, same era.
“Parameter efficient multimodal transformers for video representation learning,”
Sangho Lee et al., · 2021
Later among the works it cites.
“Layer-wise analysis of a self-supervised speech representation model,”
Ankita Pasad et al., · 2021
Later among the works it cites.
“Masked autoencoders that listen,”
Po-Yao Huang et al., · 2022
Later among the works it cites.
“Masked autoencoders are scalable vision learners,”
Kaiming He et al., · 2022
Later among the works it cites.
“Superb-sg: Enhanced speech processing universal performance benchmark for semantic and generative capabilities,”
Hsiang-Sheng Tsai et al., · 2022
Later among the works it cites.
“How severe is benchmark-sensitivity in video self-supervised learning?,”
Fida Mohammad Thoker et al., · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Vggsound: A large-scale audio-visual dataset,”
Honglie Chen et al., · 2020
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu et al., · 2021
Cited alongside, same era.
“Byol for audio: Self-supervised learning for general-purpose audio representation,”
Daisuke Niizumi et al., · 2021
Cited alongside, same era.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu wen Yang et al., · 2021
Cited alongside, same era.
“Hear: Holistic evaluation of audio representations,”
Joseph Turian et al., · 2021
Cited alongside, same era.
“Space-time crop & attend: Improving cross-modal video representation learning,”
Mandela Patrick et al., · 2021
Cited alongside, same era.
“Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100,”
Dima Damen et al., · 2022
Later among the works it cites.
“Superb@ slt 2022: Challenge on generalization and efficiency of self-supervised speech representation learning,”
Tzu-hsun Feng et al., · 2022
Later among the works it cites.
“Ego4d: Around the world in 3,000 hours of egocentric video,”
Kristen Grauman et al., · 2022
Later among the works it cites.
“VLUE: A multi-task multi-dimension benchmark for evaluating vision-language pre-training,”
Wangchunshu Zhou et al., · 2022
Later among the works it cites.
“Learning audio-visual speech representation by masked multimodal cluster prediction,”
Bowen Shi et al., · 2022
Later among the works it cites.
“Learning state-aware visual representations from audible interactions,”
Himangi Mittal et al., · 2022
Later among the works it cites.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen et al., · 2022
Later among the works it cites.
“Benchmarking self-supervised video representation learning,”
Akash Kumar et al., · 2023
Closest in time.
“Mavil: Masked audio-video learners,”
Po-Yao Huang et al., · 2023
Closest in time.