Fetching the paper…
Reading the bibliography…
We present VideoGPT: a conceptually simple architecture for scaling likelihood based generative modeling to natural videos.
Ho, J., Chen, X., Srinivas, A., Duan, Y., and Abbeel, P · 1902
Earlier work this paper cites.
Axial attention in multidimensional transformers
Ho, J., Kalchbrenner, N., Weissenborn, D., and Salimans, T · 1912
Earlier work this paper cites.
Mpeg: A video compression standard for multimedia applications
Le Gall, D · 1991
Earlier work this paper cites.
The jpeg still picture compression standard
Wallace, G. K · 1992
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A. R., and Shah, M · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Dinh, L., Krueger, D., and Bengio, Y · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Srivastava, N., Mansimov, E., and Salakhudinov, R · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Density estimation using Real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Finn, C., Goodfellow, I., and Levine, S · 2016
Earlier work this paper cites.
Kalchbrenner, N., Oord, A. v. d., Simonyan, K., Danihelka, I., Vinyals, O., Graves, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaśkowski, W · 2016
Earlier work this paper cites.
Improving variational inference with inverse autoregressive flow
Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M · 2016
Earlier work this paper cites.
Tgif: A new dataset and benchmark on animated gif description
Li, Y., Song, Y., Cao, L., Tetreault, J., Goldberg, L., Jaimes, A., and Luo, J · 2016
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Vondrick, C., Pirsiavash, H., and Torralba, A · 2016
Earlier work this paper cites.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2017
Earlier work this paper cites.
Pixelsnail: An improved autoregressive generative model
Chen, X., Mishra, N., Rohaninejad, M., and Abbeel, P · 2017
Earlier work this paper cites.
Unsupervised learning of disentangled representations from video
Denton, E. L. et al · 2017
Cited alongside, same era.
Self-supervised visual planning with temporal skip connections
Ebert, F., Finn, C., Lee, A. X., and Levine, S · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Cited alongside, same era.
Video pixel networks
Kalchbrenner, N., Oord, A., Simonyan, K., Danihelka, I., Vinyals, O., Graves, A., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Progressive growing of gans for improved quality, stability, and variation
Karras, T., Aila, T., Laine, S., and Lehtinen, J · 2017
Cited alongside, same era.
High fidelity speech synthesis with adversarial networks
Bińkowski, M., Donahue, J., Dieleman, S., Clark, A., Elsen, E., Casagrande, N., Cobo, L. C., and Simonyan, K · 2019
Later among the works it cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Later among the works it cites.
Adversarial video generation on complex datasets, 2019
Clark, A., Donahue, J., and Simonyan, K · 2019
Later among the works it cites.
Ccnet: Criss-cross attention for semantic segmentation
Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., and Liu, W · 2019
Later among the works it cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Predicting deeper into the future of semantic segmentation
Luc, P., Neverova, N., Couprie, C., Verbeek, J., and LeCun, Y · 2017
Cited alongside, same era.
Attentive semantic video generation using captions
Marwah, T., Mittal, G., and Balasubramanian, V. N · 2017
Cited alongside, same era.
Sync-draw: Automatic video generation using deep recurrent attentive architectures
Mittal, G., Marwah, T., and Balasubramanian, V. N · 2017
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
Oord, A. v. d., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G. v. d., Lockhart, E., Cobo, L. C., Stimberg, F., et al · 2017
Cited alongside, same era.
Temporal generative adversarial nets with singular value clipping
Saito, M., Matsumoto, E., and Saito, S · 2017
Cited alongside, same era.
Salimans, T., Karpathy, A., Chen, X., and Kingma, D. P · 2017
Cited alongside, same era.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Cited alongside, same era.
Waveglow: A flow-based generative network for speech synthesis
Prenger, R., Valle, R., and Catanzaro, B · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Generating diverse high-fidelity images with vq-vae-2
Razavi, A., van den Oord, A., and Vinyals, O · 2019
Later among the works it cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Later among the works it cites.
Scaling autoregressive video models
Weissenborn, D., Täckström, O., and Uszkoreit, J · 2019
Later among the works it cites.
Markov decision process for video generation
Yushchenko, V., Araslanov, N., and Roth, S · 2019
Later among the works it cites.
Self-attention generative adversarial networks
Zhang, H., Goodfellow, I., Metaxas, D., and Odena, A · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Generative pretraining from pixels
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Dhariwal, P., Luan, D., and Sutskever, I · 2020
Later among the works it cites.
Very deep vaes generalize autoregressive models and can outperform them on images
Child, R · 2020
Later among the works it cites.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Later among the works it cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Later among the works it cites.
Lower dimensional kernels for video discriminators
Kahembwe, E. and Ramamoorthy, S · 2020
Later among the works it cites.
Transformation-based adversarial video prediction on large-scale data
Luc, P., Clark, A., Dieleman, S., Casas, D. d. L., Doron, Y., Cassirer, A., and Simonyan, K · 2020
Later among the works it cites.
Rakhimov, R., Volkhonskiy, D., Artemov, A., Zorin, D., and Burnaev, E · 2020
Later among the works it cites.
Metnet: A neural weather model for precipitation forecasting
Sønderby, C. K., Espeholt, L., Heek, J., Dehghani, M., Oliver, A., Salimans, T., Agrawal, S., Hickey, J., and Kalchbrenner, N · 2020
Later among the works it cites.
Nvae: A deep hierarchical variational autoencoder
Vahdat, A. and Kautz, J · 2020
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Closest in time.
Walker, J., Razavi, A., and Oord, A. v. d · 2021
Closest in time.