Fetching the paper…
Reading the bibliography…
The objective of this paper is audio-visual synchronisation of general videos 'in the wild'.
Fundamentals of speech recognition
Lawrence Rabiner and Biing-Hwang Juang · 1993
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
John Hershey and Javier Movellan · 1999
Earlier work this paper cites.
Facesync: A linear operator for measuring synchronization of video facial images and audio tracks
Malcolm Slaney and Michele Covell · 2000
Earlier work this paper cites.
Synchronization of multiple camera videos using audio-visual features
Prarthana Shrestha, Mauro Barbieri, Hans Weda, and Dragan Sekulovski · 2009
Earlier work this paper cites.
Assessing the importance of audio/video synchronization for simultaneous translation of video sequences
Nicolas Staelens, Jonas De Meulenaere, Lizzy Bleumers, Glenn Van Wallendael, Jan De Cock, Koen Geeraert, Nick Vercammen, Wendy Van den Broeck, Brecht Vermeulen, Rik Van de Walle, et al · 2012
Earlier work this paper cites.
Audio-visual events for multi-camera synchronization
Anna Llagostera Casanovas and Andrea Cavallaro · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Cited alongside, same era.
Objects that sound
Relja Arandjelovic and Andrew Zisserman · 2018
Cited alongside, same era.
Co-training of audio and video representations from self-supervised temporal synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A. Efros · 2018
Cited alongside, same era.
Learning and using the arrow of time
Donglai Wei, Joseph J Lim, Andrew Zisserman, and William T Freeman · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Self-supervised learning of audio-visual objects from video
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
VGG-Sound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Later among the works it cites.
Audio-visual synchronisation in the wild
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman · 2021
Later among the works it cites.
Detection of audio-video synchronization errors via event detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Cited alongside, same era.
Perfect match: Improved cross-modal embeddings for audio-visual synchronisation
Soo-Whan Chung, Joon Son Chung, and Hong-Goo Kang · 2019
Cited alongside, same era.
Automated composition of picture-synched music soundtracks for movies
Vansh Dassani, Jon Bird, and Dave Cliff · 2019
Cited alongside, same era.
Dynamic temporal alignment of speech to lips
Tavi Halperin, Ariel Ephrat, and Shmuel Peleg · 2019
Cited alongside, same era.
On attention modules for audio-visual synchronization
Naji Khosravan, Shervin Ardeshir, and Rohit Puri · 2019
Cited alongside, same era.
SpecAugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Cited alongside, same era.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman
Cited in the paper.
Joshua P. Ebeneze, Yongjun Wu, Hai Wei, Sriram Sethuraman, and Zongyi Liu · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Later among the works it cites.
End-to-end lip synchronisation based on pattern classification
You Jin Kim, Hee Soo Heo, Soo-Whan Chung, and Bong-Jin Lee · 2021
Later among the works it cites.
Long short-term transformer for online action detection
Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Xia, Zhuowen Tu, and Stefano Soatto · 2021
Later among the works it cites.
VocaLiST: An audio-visual synchronisation model for lips and voices
Venkatesh S Kadandale, Juan F Montesinos, and Gloria Haro · 2022
Closest in time.