Fetching the paper…
Reading the bibliography…
The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data.
Y. Tian, D. Krishnan, and P. Isola, “Contrastive multiview coding,”
1906
Earlier work this paper cites.
D. A. Reynolds, “An overview of automatic speaker recognition technology,” in
2002
Earlier work this paper cites.
P. Belin, S. Fecteau, and C. Bedard, “Thinking the voice: neural correlates of voice perception,”
2004
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,”
2006
Earlier work this paper cites.
P. S. Aleksic and A. K. Katsaggelos, “Audio-visual biometrics,”
2006
Earlier work this paper cites.
L. Shams and R. Kim, “Crossmodal influences on visual perception,”
2010
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in
2013
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, C. Gulçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” in
2014
Earlier work this paper cites.
N. Goncalves, J. Nikkilä, and R. Vigario, “Self-supervised mri tissue segmentation by discriminative clustering,”
2014
Earlier work this paper cites.
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman, “Return of the devil in the details: Delving deep into convolutional nets,” in
2014
Earlier work this paper cites.
M. Liang and X. Hu, “Recurrent convolutional neural network for object recognition,” in
2015
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Earlier work this paper cites.
C. Doersch, A. Gupta, and A. A. Efros, “Unsupervised visual representation learning by context prediction,” in
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in
2016
Cited alongside, same era.
R. Zhang, P. Isola, and A. A. Efros, “Colorful image colorization,” in
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in
2016
Cited alongside, same era.
K. Sohn, “Improved deep metric learning with multi-class n-pair loss objective,” in
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Signs in time: Encoding human motion as a temporal image,” in
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in
2018
Later among the works it cites.
A. Nagrani, S. Albanie, and A. Zisserman, “Learnable pins: Cross-modal embeddings for person identity,” in
2018
Later among the works it cites.
H. Wang, Y. Wang, Z. Zhou, X. Ji, D. Gong, J. Zhou, Z. Li, and W. Liu, “Cosface: Large margin cosine loss for deep face recognition,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Nagrani, S. Albanie, and A. Zisserman, “Seeing voices and hearing faces: Cross-modal biometric matching,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
K. Gong, X. Liang, D. Zhang, X. Shen, and L. Lin, “Look into person: Self-supervised structure-sensitive learning and a new benchmark for human parsing,” in
2017
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “VoxCeleb: a large-scale speaker identification dataset,” in
2017
Cited alongside, same era.
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” in
2017
Cited alongside, same era.
D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superpoint: Self-supervised interest point detection and description,” in
2018
Cited alongside, same era.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in
2018
Cited alongside, same era.
R. Arandjelović and A. Zisserman, “Objects that sound,” in
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep speaker recognition,” in
2018
Later among the works it cites.
——, “Learning to lip read words by watching videos,”
2018
Later among the works it cites.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in
2019
Later among the works it cites.
T.-H. Oh, T. Dekel, C. Kim, I. Mosseri, W. T. Freeman, M. Rubinstein, and W. Matusik, “Speech2face: Learning the face behind a voice,” in
2019
Later among the works it cites.
S.-W. Chung, J. S. Chung, and H.-G. Kang, “Perfect match: Self-supervised embeddings for cross-modal retrieval,”
2020
Closest in time.
A. Nagrani, J. S. Chung, S. Albanie, and A. Zisserman, “Disentangled speech embeddings using cross-modal self-supervision,” in
2020
Closest in time.
2020
Closest in time.
C. Doersch and A. Zisserman, “Multi-task self-supervised visual learning,” in
2060
Closest in time.