Fetching the paper…
Reading the bibliography…
We present a new model DrNET that learns disentangled image representations from video.
Multiple view geometry in computer vision, 2000
R. Hartley and A. Zisserman · 2000
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariance
L. Wiskott and T. Sejnowski · 2002
Earlier work this paper cites.
Learning methods for generic object recognition with invariance to pose and lighting
Y. LeCun, F. Huang, and L. Bottou · 2004
Earlier work this paper cites.
Recognizing human actions: A local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
Beyond pixels: exploring new representations and applications for motion analysis
C. Liu · 2009
Earlier work this paper cites.
Transforming auto-encoders
G. E. Hinton, A. Krizhevsky, and S. Wang · 2011
Earlier work this paper cites.
Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis
Q. V. Le, W. Y. Zou, S. Y. Yeung, and A. Y. Ng · 2011
Earlier work this paper cites.
Generative adversarial nets
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep convolutional inverse graphics network
T. D. Kulkarni, W. F. Whitney, P. Kohli, and J. Tenenbaum · 2015
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
M. Mathieu, C. Couprie, and Y. LeCun · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in Atari games
J. Oh, X. Guo, H. Lee, R. Lewis, and S. Singh · 2015
Cited alongside, same era.
Semi-supervised learning with ladder network
A. Rasmus, M. Berglund, M. Honkala, H. Valpola, and T. Raiko · 2015
Cited alongside, same era.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu · 2016
Later among the works it cites.
Disentangling factors of variation in deep representations using adversarial training
M. Mathieu, P. S. Junbo Zhao, A. Ramesh, and Y. LeCun · 2016
Later among the works it cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
A. Radford, L. Metz, and S. Chintala · 2016
Later among the works it cites.
Improved techniques for training gans
T. Salimans, I. J. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Later among the works it cites.
Pixel recurrent neural networks
A. van den Oord, N. Kalchbrenner, and K. Kavukcuoglu · 2016
Later among the works it cites.
Generating videos with scene dynamics
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
Cited alongside, same era.
Dense optical flow prediction from a static image
J. Walker, A. Gupta, and M. Hebert · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
X. Wang and A. Gupta · 2015
Cited alongside, same era.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Cited alongside, same era.
F. Cricri, M. Honkala, X. Ni, E. Aksu, and M. Gabbouj · 2016
Cited alongside, same era.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Cited alongside, same era.
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Later among the works it cites.
Understanding visual concepts with continuation learning
W. F. Whitney, M. Chang, T. Kulkarni, and J. B. Tenenbaum · 2016
Later among the works it cites.
Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks
T. Xue, J. Wu, K. L. Bouman, and W. T. Freeman · 2016
Later among the works it cites.
Stacked what-where auto-encoders
J. Zhao, M. Mathieu, R. Goroshin, and Y. LeCun · 2016
Later among the works it cites.
Recurrent environment simulators
S. Chiappa, S. Racaniere, D. Wierstra, and S. Mohamed · 2017
Closest in time.
T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma · 2017
Closest in time.
Semantic scene completion from a single depth image
S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser · 2017
Closest in time.
Decomposing motion and content for natural video sequence prediction
R. Villegas, J. Yang, S. Hong, X. Lin, and H. Lee · 2017
Closest in time.