Fetching the paper…
Reading the bibliography…
We propose a method for learning landmark detectors for visual objects (such as the eyes and the nose in a face) without any manual supervision.
Splines minimizing rotation-invariant semi-norms in sobolev spaces
J. Duchon · 1977
Earlier work this paper cites.
Spline models for observational data , volume 59
G. Wahba · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Learning methods for generic object recognition with invariance to pose and lighting
Y. LeCun, F. J. Huang, and L. Bottou · 2004
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y.-W. Teh · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Earlier work this paper cites.
The recurrent temporal restricted boltzmann machine
I. Sutskever, G. E. Hinton, and G. W. Taylor · 2009
Earlier work this paper cites.
Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization
M. Koestinger, P. Wohlhart, P. M. Roth, and H. Bischof · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Earlier work this paper cites.
Articulated pose estimation with flexible mixtures-of-parts
Y. Yang and D. Ramanan · 2011
Earlier work this paper cites.
Robust face landmark estimation under occlusion
X. P. Burgos-Artizzu, P. Perona, and P. Dollár · 2013
Earlier work this paper cites.
Domain adaptation for upper body pose tracking in signed TV broadcasts
J. Charles, T. Pfister, D. Magee, D. Hogg, and A. Zisserman · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Large-scale learning of sign language by watching TV (using co-occurrences)
T. Pfister, J. Charles, and A. Zisserman · 2013
Earlier work this paper cites.
Deep convolutional network cascade for facial point detection
Y. Sun, X. Wang, and X. Tang · 2013
Earlier work this paper cites.
Articulated pose estimation by a graphical model with image dependent pairwise relations
X. Chen and A. L. Yuille · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Deep convolutional neural networks for efficient pose estimation in gesture videos
T. Pfister, K. Simonyan, J. Charles, and A. Zisserman · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Learning what and where to draw
S. E. Reed, Z. Akata, S. Mohan, S. Tenka, B. Schiele, and H. Lee · 2016
Later among the works it cites.
Generating videos with scene dynamics
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Later among the works it cites.
Understanding visual concepts with continuation learning
W. F. Whitney, M. Chang, T. Kulkarni, and J. B. Tenenbaum · 2016
Later among the works it cites.
Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks
T. Xue, J. Wu, K. L. Bouman, and W. T. Freeman · 2016
Later among the works it cites.
Learning Deep Representation for Face Alignment with Auxiliary Attributes
Z. Zhang, P. Luo, C. C. Loy, and X. Tang · 2016
Later among the works it cites.
Photographic image synthesis with cascaded refinement networks
Q. Chen and V. Koltun · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning face attributes in the wild
Z. L., P. L., X. W., and X. T · 2015
Cited alongside, same era.
Spatio-temporal video autoencoder with differentiable memory
V. Patraucean, A. Handa, and R. Cipolla · 2015
Cited alongside, same era.
Flowing convnets for human pose estimation in videos
T. Pfister, J. Charles, and A. Zisserman · 2015
Cited alongside, same era.
Deep visual analogy-making
S. E. Reed, Y. Zhang, Y. Zhang, and H. Lee · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Cited alongside, same era.
Super-resolution with deep convolutional sufficient statistics
J. Bruna, P. Sprechmann, and Y. LeCun · 2016
Cited alongside, same era.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
Unsupervised learning of disentangled representations from video
E. L. Denton and V. Birodkar · 2017
Later among the works it cites.
Image-to-image translation with conditional adversarial networks
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros · 2017
Later among the works it cites.
Photo-realistic single image super-resolution using a generative adversarial network
C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, and W. Shi · 2017
Later among the works it cites.
Voxceleb: a large-scale speaker identification dataset
A. Nagrani, J. S. Chung, and A. Zisserman · 2017
Later among the works it cites.
Plug & play generative networks: Conditional iterative generation of images in latent space
A. Nguyen, J. Yosinski, Y. Bengio, A. Dosovitskiy, and J. Clune · 2017
Later among the works it cites.
Learning to generate long-term future via hierarchical prediction
R. Villegas, J. Yang, Y. Zou, S. Sohn, X. Lin, and H. Lee · 2017
Later among the works it cites.
VoxCeleb2: Deep speaker recognition
J. S. Chung, A. Nagrani, and A. Zisserman · 2018
Closest in time.
Deforming autoencoders: Unsupervised disentangling of shape and appearance
Z. Shu, M. Sahasrabudhe, A. Guler, D. Samaras, N. Paragios, and I. Kokkinos · 2018
Closest in time.
Discovery of latent 3d keypoints via end-to-end geometric reasoning
S. Suwajanakorn, N. Snavely, J. Tompson, and M. Norouzi · 2018
Closest in time.
Self-supervised learning of a facial attribute embedding from video
O. Wiles, A. S. Koepke, and A. Zisserman · 2018
Closest in time.
Unsupervised discovery of object landmarks as structural representations
Y. Zhang, Y. Guo, Y. Jin, Y. Luo, Z. He, and H. Lee · 2018
Closest in time.