Fetching the paper…
Reading the bibliography…
To generate accurate videos, algorithms have to understand the spatial and temporal dependencies in the world.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P · 2004
Earlier work this paper cites.
Scope of validity of psnr in image/video quality assessment
Huynh-Thu, Q. and Ghanbari, M · 2008
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons
Bengio, Y · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Modeling arbitrary probability distributions using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2014
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al · 2016
Earlier work this paper cites.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., and Zhang, Y · 2017
Earlier work this paper cites.
Video pixel networks
Kalchbrenner, N., Oord, A., Simonyan, K., Danihelka, I., Vinyals, O., Graves, A., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Stochastic video generation with a learned prior
Denton, E. and Fergus, R · 2018
Earlier work this paper cites.
Neural scene representation and rendering
Eslami, S. A., Jimenez Rezende, D., Besse, F., Viola, F., Morcos, A. S., Garnelo, M., Ruderman, A., Rusu, A. A., Danihelka, I., Gregor, K., et al · 2018
Earlier work this paper cites.
Stochastic adversarial video prediction
Lee, A. X., Zhang, R., Ebert, F., Abbeel, P., Finn, C., and Levine, S · 2018
Earlier work this paper cites.
Tganv2: Efficient training of large models for video generation with multiple subsampling layers
Saito, M. and Saito, S · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
Tulyakov, S., Liu, M.-Y., Yang, X., and Kautz, J · 2018
Earlier work this paper cites.
Gibson env: Real-world perception for embodied agents
Xia, F., Zamir, A. R., He, Z., Sax, A., Malik, J., and Savarese, S · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Adversarial video generation on complex datasets, 2019
Clark, A., Donahue, J., and Simonyan, K · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
Guss, W. H., Houghton, B., Topin, N., Wang, P., Codel, C., Veloso, M., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Videoflow: A flow-based generative model for video
Kumar, M., Babaeizadeh, M., Erhan, D., Finn, C., Levine, S., Dinh, L., and Kingma, D · 2019
Cited alongside, same era.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., and Lillicrap, T. P · 2019
Cited alongside, same era.
Habitat: A Platform for Embodied AI Research
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., and Batra, D · 2019
Cited alongside, same era.
Fvd: A new metric for video generation
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2019
Cited alongside, same era.
High fidelity video prediction with large stochastic recurrent neural networks
Villegas, R., Pathak, A., Kannan, H., Erhan, D., Le, Q. V., and Lee, H · 2019
Cited alongside, same era.
Walker, J., Razavi, A., and Oord, A. v. d · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Later among the works it cites.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Later among the works it cites.
Zhai, S., Talbott, W., Srivastava, N., Huang, C., Goh, H., Zhang, R., and Susskind, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling autoregressive video models
Weissenborn, D., Täckström, O., and Uszkoreit, J · 2019
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Transformation-based adversarial video prediction on large-scale data
Luc, P., Clark, A., Dieleman, S., Casas, D. d. L., Doron, Y., Cassirer, A., and Simonyan, K · 2020
Cited alongside, same era.
Rakhimov, R., Volkhonskiy, D., Artemov, A., Zorin, D., and Burnaev, E · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Fitvid: Overfitting in pixel-level video prediction
Babaeizadeh, M., Saffar, M. T., Nair, S., Levine, S., Finn, C., and Erhan, D · 2021
Cited alongside, same era.
Brooks, T., Hellsten, J., Aittala, M., Wang, T.-C., Aila, T., Lehtinen, J., Liu, M.-Y., Efros, A. A., and Karras, T · 2022
Closest in time.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Closest in time.
Long video generation with time-agnostic vqgan and time-sensitive transformer
Ge, S., Hayes, T., Yang, H., Yin, X., Pang, G., Jacobs, D., Huang, J.-B., and Parikh, D · 2022
Closest in time.
Maskvit: Masked visual pre-training for video prediction
Gupta, A., Tian, S., Zhang, Y., Wu, J., Martín-Martín, R., and Fei-Fei, L · 2022
Closest in time.
Flexible diffusion modeling of long videos
Harvey, W., Naderiparizi, S., Masrani, V., Weilbach, C., and Wood, F · 2022
Closest in time.
General-purpose, long-context autoregressive modeling with perceiver ar
Hawthorne, C., Jaegle, A., Cangea, C., Borgeaud, S., Nash, C., Malinowski, M., Dieleman, S., Vinyals, O., Botvinick, M., Simon, I., et al · 2022
Closest in time.
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Closest in time.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Closest in time.
Diffusion models for video prediction and infilling
Höppe, T., Mehrjou, A., Bauer, S., Nielsen, D., and Dittadi, A · 2022
Closest in time.
Hutchins, D., Schlag, I., Wu, Y., Dyer, E., and Neyshabur, B · 2022
Closest in time.
Draft-and-revise: Effective image generation with contextual rq-transformer
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Closest in time.
Long movie clip classification with state-space video models
Mohaiminul Islam, M. and Bertasius, G · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Closest in time.
Harp: Autoregressive latent video prediction with high-fidelity image generator
Seo, Y., Lee, K., Liu, F., James, S., and Abbeel, P · 2022
Closest in time.
Phenaki: Variable length video generation from open domain textual description
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2022
Closest in time.
Masked conditional video diffusion for prediction, generation, and interpolation
Voleti, V., Jolicoeur-Martineau, A., and Pal, C · 2022
Closest in time.
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Wu, C.-Y., Li, Y., Mangalam, K., Fan, H., Xiong, B., Malik, J., and Feichtenhofer, C · 2022
Closest in time.
Generating videos with dynamics-aware implicit generative adversarial networks
Yu, S., Tack, J., Mo, S., Kim, H., Kim, J., Ha, J.-W., and Shin, J · 2022
Closest in time.