Fetching the paper…
Reading the bibliography…
Video generation is one of the most challenging tasks in Machine Learning and Computer Vision fields of study.
Whistler: A trainable text-to-speech system
X. Huang, A. Acero, J. Adcock, H.-W. Hon, J. Goldsmith, J. Liu, and M. Plumpe · 1996
Earlier work this paper cites.
Recognizing human actions: a local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Earlier work this paper cites.
Can humans fly? Action understanding with multiple classes of actors
C. Xu, S.-H. Hsieh, C. Xiong, and J. J. Corso · 2015
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
W. Lotter, G. Kreiman, and D. Cox · 2016
Earlier work this paper cites.
Improved techniques for training gans
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Earlier work this paper cites.
Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network
W. Shi, J. Caballero, F. Huszár, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
S. Ö. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, et al · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
Modulating early visual processing by language
H. De Vries, F. Strub, J. Mary, H. Larochelle, O. Pietquin, and A. C. Courville · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Cited alongside, same era.
Progressive growing of gans for improved quality, stability, and variation
T. Karras, T. Aila, S. Laine, and J. Lehtinen · 2017
Cited alongside, same era.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Cited alongside, same era.
Attentive semantic video generation using captions
T. Marwah, G. Mittal, and V. N. Balasubramanian · 2017
Cited alongside, same era.
To create what you tell: Generating videos from captions
Y. Pan, Z. Qiu, T. Yao, H. Li, and T. Mei · 2017
Cited alongside, same era.
C. Spampinato, S. Palazzo, P. D’Oro, F. Murabito, D. Giordano, and M. Shah · 2018
Later among the works it cites.
Mocogan: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2018
Later among the works it cites.
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, G. Liu, A. Tao, J. Kautz, and B. Catanzaro · 2018
Later among the works it cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He · 2018
Later among the works it cites.
Generative image inpainting with contextual attention
J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. S. Huang · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to generate long-term future via hierarchical prediction
R. Villegas, J. Yang, Y. Zou, S. Sohn, X. Lin, and H. Lee · 2017
Cited alongside, same era.
C. Chan, S. Ginosar, T. Zhou, and A. A. Efros · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Actor and action video segmentation from a sentence
K. Gavrilyuk, A. Ghodrati, Z. Li, and C. G. Snoek · 2018
Cited alongside, same era.
Stochastic adversarial video prediction
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Cited alongside, same era.
Video generation from text
Y. Li, M. R. Min, D. Shen, D. Carlson, and L. Carin · 2018
Cited alongside, same era.
B. McIntosh, K. Duarte, Y. S. Rawat, and M. Shah · 2018
Cited alongside, same era.
Later among the works it cites.
Pay attention!-robustifying a deep visuomotor policy through task-focused visual attention
P. Abolghasemi, A. Mazaheri, M. Shah, and L. Boloni · 2019
Later among the works it cites.
Efficient video generation on complex datasets
A. Clark, J. Donahue, and K. Simonyan · 2019
Later among the works it cites.
Edge-aware deep image deblurring
Z. Fu, Y. Zheng, H. Ye, Y. Kong, J. Yang, and L. He · 2019
Later among the works it cites.
Semantic object accuracy for generative text-to-image synthesis
T. Hinz, S. Heinrich, and S. Wermter · 2019
Later among the works it cites.
Deep video inpainting
D. Kim, S. Woo, J.-Y. Lee, and I. S. Kweon · 2019
Later among the works it cites.
Cross-modal dual learning for sentence-to-video generation
Y. Liu, X. Wang, Y. Yuan, and W. Zhu · 2019
Later among the works it cites.
Video generation from single semantic label map
J. Pan, C. Wang, X. Jia, J. Shao, L. Sheng, J. Yan, and X. Wang · 2019
Later among the works it cites.
Animating arbitrary objects via deep motion transfer
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe · 2019
Later among the works it cites.