Fetching the paper…
Reading the bibliography…
This paper proposes the novel task of video generation conditioned on a SINGLE semantic label map, which provides a good balance between flexibility and quality in the generation process.
Recognizing human actions: a local svm approach
I. Laptev, B. Caputo, et al · 2004
Earlier work this paper cites.
An unbiased second-order prior for high-accuracy motion estimation
W. Trobin, T. Pock, D. Cremers, and H. Bischof · 2008
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
As-rigid-as-possible stereo under second order smoothness priors
C. Zhang, Z. Li, R. Cai, H. Chao, and Y. Rui · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Dynamic filter networks
X. Jia, B. De Brabandere, T. Tuytelaars, and L. V. Gool · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
J. Johnson, A. Alahi, and L. Fei-Fei · 2016
Cited alongside, same era.
Pixel recurrent neural networks
A. v. d. Oord, N. Kalchbrenner, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Neural autoregressive distribution estimation
B. Uria, M.-A. Côté, K. Gregor, I. Murray, and H. Larochelle · 2016
Cited alongside, same era.
Generating videos with scene dynamics
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Cited alongside, same era.
An uncertain future: Forecasting from static images using variational autoencoders
J. Walker, C. Doersch, A. Gupta, and M. Hebert · 2016
Cited alongside, same era.
Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks
Temporal generative adversarial nets with singular value clipping
M. Saito, E. Matsumoto, and S. Saito · 2017
Later among the works it cites.
Learning from simulated and unsupervised images through adversarial training
A. Shrivastava, T. Pfister, O. Tuzel, J. Susskind, W. Wang, and R. Webb · 2017
Later among the works it cites.
Mocogan: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2017
Later among the works it cites.
Decomposing motion and content for natural video sequence prediction
R. Villegas, J. Yang, S. Hong, X. Lin, and H. Lee · 2017
Later among the works it cites.
The pose knows: Video forecasting by generating pose futures
J. Walker, K. Marino, A. Gupta, and M. Hebert · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Xue, J. Wu, K. Bouman, and B. Freeman · 2016
Cited alongside, same era.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine · 2017
Cited alongside, same era.
Unsupervised pixel-level domain adaptation with generative adversarial networks
K. Bousmalis, N. Silberman, D. Dohan, D. Erhan, and D. Krishnan · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros · 2017
Cited alongside, same era.
Unsupervised learning of long-term motion dynamics for videos
Z. Luo, B. Peng, D.-A. Huang, A. Alahi, and L. Fei-Fei · 2017
Cited alongside, same era.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
H. Zhang, T. Xu, H. Li, S. Zhang, X. Huang, X. Wang, and D. Metaxas · 2017
Later among the works it cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros · 2017
Later among the works it cites.
Synthesizing images of humans in unseen poses
G. Balakrishnan, A. Zhao, A. V. Dalca, F. Durand, and J. Guttag · 2018
Later among the works it cites.
Encoder-decoder with atrous separable convolution for semantic image segmentation
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam · 2018
Later among the works it cites.
Stochastic video generation with a learned prior
E. Denton and R. Fergus · 2018
Later among the works it cites.
Image generation from scene graphs
J. Johnson, A. Gupta, and L. Fei-Fei · 2018
Later among the works it cites.
Flow-grounded spatial-temporal video prediction from still images
Y. Li, C. Fang, J. Yang, Z. Wang, X. Lu, and M.-H. Yang · 2018
Later among the works it cites.
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, G. Liu, A. Tao, J. Kautz, and B. Catanzaro · 2018
Later among the works it cites.