Fetching the paper…
Reading the bibliography…
We propose a deep video prediction model conditioned on a single image and an action class.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” in
1997
Earlier work this paper cites.
C. Schuldt, I. Laptev, and B. Caputo, “Recognizing human actions: a local svm approach,” in
2004
Earlier work this paper cites.
K. Soomro, A. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,” in
2012
Earlier work this paper cites.
H. Dibeklioğlu, A. Salah, and T. Gevers, “Are you really smiling at me? spontaneous versus posed enjoyment smiles,” in
2012
Earlier work this paper cites.
W. Zhang, M. Zhu, and K. Derpanis, “From actemes to action: A strongly-supervised representation for detailed action understanding,” in
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
D. Kingma and M. Welling, “Auto-encoding variational bayes,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in
2014
Earlier work this paper cites.
N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using lstms,” in
2015
Earlier work this paper cites.
S. Reed, Y. Zhang, Y. Zhang, and H. Lee, “Deep visual analogy-making,” in
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and L. Fei-Fei, “Imagenet large scale visual recognition challenge,” in
2015
Earlier work this paper cites.
K. Sohn, X. Yan, and H. Lee, “Learning structured output representation using deep conditional generative models,” in
2015
Earlier work this paper cites.
C. Finn, I. Goodfellow, and S. Levine, “Unsupervised learning for physical interaction through video prediction,” in
2016
Earlier work this paper cites.
B. de Brabandere, X. Jia, T. Tuytelaars, and L. van Gool, “Dynamic filter networks,” in
2016
Cited alongside, same era.
M. Mathieu, C. Couprie, and Y. LeCun, “Deep multi-scale video prediction beyond mean square error,” in
2016
Cited alongside, same era.
A. Newell, K. Yang, and J. Deng, “Stacked hourglass networks for human pose estimation,” in
2016
Cited alongside, same era.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in
2016
Cited alongside, same era.
N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu, “Video pixel networks,” in
2017
Cited alongside, same era.
F. Ebert, C. Finn, A. Lee, and S. Levine, “Self-supervised visual planning with temporal skip connections,” in
2018
Later among the works it cites.
E. Denton and R. Fergus, “Stochastic video generation with a learned prior,” in
2018
Later among the works it cites.
S. Tulyakov, M. Liu, X. Yang, and J. Kautz, “MoCoGAN: Decomposing motion and content for video generation,” in
2018
Later among the works it cites.
N. Wichers, R. Villegas, D. Erhan, and H. Lee, “Hierarchical long-term video prediction without supervision,” in
2018
Later among the works it cites.
G. Balakrishnan, A. Zhao, A. Dalca, F. Durand, and J. Guttag, “Synthesizing images of humans in unseen poses,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
R. Villegas, J. Yang, S. Hong, X. Lin, and H. Lee, “Decomposing motion and content for natural video sequence prediction,” in
2017
Cited alongside, same era.
E. Denton and V. Birodkar, “Unsupervised learning of disentangled representations from video,” in
2017
Cited alongside, same era.
L. Ma, X. Jia, Q. Sun, B. Schiele, T. Tuytelaars, and L. van Gool, “Pose guided person image generation,” in
2017
Cited alongside, same era.
S. Reed, A. van den Oord, N. Kalchbrenner, S. Colmenarejo, Z. Wang, Y. Chen, D. Belov, and N. de Freitas, “Parallel multiscale autoregressive density estimation,” in
2017
Cited alongside, same era.
R. Villegas, J. Yang, Y. Zou, S. Sohn, X. Lin, and H. Lee, “Learning to generate long-term future via hierarchical prediction,” in
2017
Cited alongside, same era.
J. Thewlis, H. Bilen, and A. Vedaldi, “Unsupervised learning of object landmarks by factorized spatial embeddings,” in
2017
Cited alongside, same era.
2018
Later among the works it cites.
H. Cai, C. Bai, Y. Tai, and C. Tang, “Deep video generation, prediction and completion of human action sequences,” in
2018
Later among the works it cites.
W. Wei, X. Alameda-Pineda, D. Xu, P. Fua, E. Ricci, and N. Sebe, “Every smile is unique: Landmark-guided diverse smile generation,” in
2018
Later among the works it cites.
Y. Zhang, Y. Guo, Y. Jin, Y. Luo, Z. He, and H. Lee, “Unsupervised discovery of object landmarks as structural representations,” in
2018
Later among the works it cites.
T. Jakab, A. Gupta, H. Bilen, and A. Vedaldi, “Unsupervised learning of object landmarks through conditional image generation,” in
2018
Later among the works it cites.
Y. A. Mejjati, C. Richardt, J. Tompkin, D. Cosker, and K. I. Kim, “Unsupervised attention-guided image-to-image translation,” in
2018
Later among the works it cites.
Y. Li, C. Fang, J. Yang, Z. Wang, X. Lu, and M. Yang, “Flow-grounded spatial-temporal video prediction from still images,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe, “Animating arbitrary objects via deep motion transfer,” in
2019
Closest in time.