Fetching the paper…
Reading the bibliography…
This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization.
“Reducing the dimensionality of data with neural networks,”
G. E. Hinton and R. R. Salakhutdinov, · 2006
Earlier work this paper cites.
“Return of the devil in the details: Delving deep into convolutional nets,”
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman, · 2014
Earlier work this paper cites.
“Context encoders: Feature learning by inpainting,”
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, · 2016
Earlier work this paper cites.
“Colorful image colorization,”
R. Zhang, P. Isola, and A. A. Efros, · 2016
Earlier work this paper cites.
“Out of time: automated lip sync in the wild,”
J. S. Chung and A. Zisserman, · 2016
Earlier work this paper cites.
“Lip reading in the wild,”
J. S. Chung and A. Zisserman, · 2016
Cited alongside, same era.
“Lip reading sentences in the wild,”
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, · 2017
Cited alongside, same era.
“Dynamic temporal alignment of speech to lips,”
T. Halperin, A. Ephrat, and S. Peleg, · 2018
Cited alongside, same era.
“Objects that sound,”
R. Arandjelović and A. Zisserman, · 2018
Cited alongside, same era.
“Learning to localize sound source in visual scenes,”
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, · 2018
Cited alongside, same era.
“Co-training of audio and video representations from self-supervised temporal synchronization,”
B. Korbar, D. Tran, and L. Torresani, · 2018
Closest in time.
“Audio-visual scene analysis with self-supervised multisensory features,”
A. Owens and A. A. Efros, · 2018
Closest in time.
“Seeing voices and hearing faces: Cross-modal biometric matching,”
A. Nagrani, S. Albanie, and A. Zisserman, · 2018
Closest in time.
“On learning associations of faces and voices,”
C. Kim, H. V. Shin, T.-H. Oh, A. Kaspar, M. Elgharib, and W. Matusik, · 2018
Closest in time.
“Learning to lip read words by watching videos,”
J. S. Chung and A. Zisserman, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…