Fetching the paper…
Reading the bibliography…
Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse.
Fundamentals of speech recognition
L. Rabiner and B.-H. Juang, · 1993
Earlier work this paper cites.
“Audio vision: Using audio-visual synchrony to locate sounds,”
J. Hershey and J. Movellan, · 1999
Earlier work this paper cites.
“Facesync: A linear operator for measuring synchronization of video facial images and audio tracks,”
M. Slaney and M. Covell, · 2000
Earlier work this paper cites.
“Out of time: automated lip sync in the wild,”
J. S. Chung and A. Zisserman, · 2016
Earlier work this paper cites.
“Lip reading in the wild,”
J. S. Chung and A. Zisserman, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
J. Gemmeke, D. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“The "something something" video database for learning and evaluating visual common sense,”
R. Goyal, S. Kahou, V. Michalski, J. Materzynska, et al., · 2017
Earlier work this paper cites.
“Objects that sound,”
R. Arandjelovic and A. Zisserman, · 2018
Earlier work this paper cites.
“Audio-visual scene analysis with self-supervised multisensory features,”
A. Owens and A. Efros, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
A. Oord, Y. Li, and O. Vinyals, · 2018
Cited alongside, same era.
“LRS3-TED: a large-scale dataset for visual speech recognition,”
T. Afouras, J. S. Chung, and A. Zisserman, · 2018
Cited alongside, same era.
“Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,”
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, · 2018
Cited alongside, same era.
“Perfect match: Improved cross-modal embeddings for audio-visual synchronisation,”
S.-W. Chung, J. S. Chung, and H.-G. Kang, · 2019
Cited alongside, same era.
“End-to-end object detection with transformers,”
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, · 2020
Later among the works it cites.
“Audio-visual synchronisation in the wild,”
H. Chen, W. Xie, T. Afouras, A. Nagrani, A. Vedaldi, and A. Zisserman, · 2021
Later among the works it cites.
“End-to-end lip synchronisation based on pattern classification,”
Y. J. Kim, H. S. Heo, S.-W. Chung, and B.-J. Lee, · 2021
Later among the works it cites.
“AST: Audio Spectrogram Transformer,”
Y. Gong, Y. Chung, and J. Glass, · 2021
Later among the works it cites.
“Keeping your eye on the ball: Trajectory attention in video transformers,”
P. Mandela, D. Campbell, Y. Asano, I. Misra, F. Metze, C. Feichtenhofer, A. Vedaldi, and J. F. Henriques, · 2021
Later among the works it cites.
“Is space-time attention all you need for video understanding?,”
G. Bertasius, H. Wang, and L. Torresani, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Dynamic temporal alignment of speech to lips,”
T. Halperin, A. Ephrat, and S. Peleg, · 2019
Cited alongside, same era.
“On attention modules for audio-visual synchronization,”
N. Khosravan, S. Ardeshir, and R. Puri, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, · 2019
Cited alongside, same era.
“Self-supervised learning of audio-visual objects from video,”
T. Afouras, A. Owens, J. S. Chung, and A. Zisserman, · 2020
Cited alongside, same era.
“VGG-Sound: A large-scale audio-visual dataset,”
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, · 2020
Cited alongside, same era.
Later among the works it cites.
“Learning transferable visual models from natural language supervision,”
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, et al., · 2021
Later among the works it cites.
“Sparse in space and time: Audio-visual synchronisation with trainable selectors,”
V. Iashin, W. Xie, E. Rahtu, and A. Zisserman, · 2022
Later among the works it cites.
“Vocalist: An audio-visual synchronisation model for lips and voices,”
V. S. Kadandale, J. F. Montesinos, and G. Haro, · 2022
Later among the works it cites.
“ModEFormer: Modality-preserving embedding for audio-video synchronization using transformers,”
A. Gupta, R. Tripathi, and W. Jang, · 2023
Later among the works it cites.