Fetching the paper…
Reading the bibliography…
Large-scale pretrained transformers have created milestones in text (GPT-3) and text-to-image (DALL-E and CogView) generation.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. Hinton · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Generative adversarial networks
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Improved techniques for training gans
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
Temporal generative adversarial nets with singular value clipping
M. Saito, E. Matsumoto, and S. Saito · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Predrnn: Recurrent neural networks for predictive learning using spatiotemporal lstms
Y. Wang, M. Long, J. Wang, Z. Gao, and P. S. Yu · 2017
Cited alongside, same era.
A short note about kinetics-600
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman · 2018
Cited alongside, same era.
Mocogan: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2018
Cited alongside, same era.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Cited alongside, same era.
Adversarial video generation on complex datasets
A. Clark, J. Donahue, and K. Simonyan · 2019
Cited alongside, same era.
R. Rakhimov, D. Volkhonskiy, A. Artemov, D. Zorin, and E. Burnaev · 2020
Later among the works it cites.
Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan
M. Saito, S. Saito, M. Koyama, and S. Kobayashi · 2020
Later among the works it cites.
Cogview: Mastering text-to-image generation via transformers
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang, et al · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tsm: Temporal shift module for efficient video understanding
J. Lin, C. Gan, and S. Han · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Cited alongside, same era.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang · 2019
Cited alongside, same era.
Scaling autoregressive video models
D. Weissenborn, O. Täckström, and J. Uszkoreit · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2020
Cited alongside, same era.
Transformation-based adversarial video prediction on large-scale data
P. Luc, A. Clark, S. Dieleman, D. d. L. Casas, Y. Doron, A. Cassirer, and K. Simonyan · 2020
Cited alongside, same era.
A good image generator is what you need for high-resolution video synthesis
Y. Tian, J. Ren, M. Chai, K. Olszewski, X. Peng, D. N. Metaxas, and S. Tulyakov · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas · 2021
Later among the works it cites.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
M. Ding, W. Zheng, W. Hong, and J. Tang · 2022
Closest in time.
Long video generation with time-agnostic vqgan and time-sensitive transformer
S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh · 2022
Closest in time.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Closest in time.
Generating videos with dynamics-aware implicit generative adversarial networks
S. Yu, J. Tack, S. Mo, H. Kim, J. Kim, J.-W. Ha, and J. Shin · 2022
Closest in time.