Fetching the paper…
Reading the bibliography…
In this work, we present Patch-based Object-centric Video Transformer (POVT), a novel region-based video generation architecture that leverages object-centric information to efficiently model temporal dynamics in videos.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Space: Unsupervised object-oriented scene representation via spatial attention and decomposition
Lin, Z., Wu, Y.-F., Peri, S. V., Sun, W., Singh, G., Deng, F., Jiang, J., and Ahn, S · 2001
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P · 2004
Earlier work this paper cites.
Scope of validity of psnr in image/video quality assessment
Huynh-Thu, Q. and Ghanbari, M · 2008
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Spatial transformer networks
Jaderberg, M., Simonyan, K., Zisserman, A., et al · 2015
Earlier work this paper cites.
Simple online and realtime tracking
Bewley, A., Ge, Z., Ott, L., Ramos, F., and Upcroft, B · 2016
Earlier work this paper cites.
Density estimation using Real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Vondrick, C., Pirsiavash, H., and Torralba, A · 2016
Earlier work this paper cites.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al · 2017
Earlier work this paper cites.
Video pixel networks
Kalchbrenner, N., Oord, A., Simonyan, K., Danihelka, I., Vinyals, O., Graves, A., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Predicting deeper into the future of semantic segmentation
Luc, P., Neverova, N., Couprie, C., Verbeek, J., and LeCun, Y · 2017
Earlier work this paper cites.
Temporal generative adversarial nets with singular value clipping
Saito, M., Matsumoto, E., and Saito, S · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Learning to generate long-term future via hierarchical prediction
Villegas, R., Yang, J., Zou, Y., Sohn, S., Lin, X., and Lee, H · 2017
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2018
Earlier work this paper cites.
Stochastic video generation with a learned prior
Denton, E. and Fergus, R · 2018
Earlier work this paper cites.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P · 2018
Earlier work this paper cites.
Sequential attend, infer, repeat: Generative modelling of moving objects
Kosiorek, A. R., Kim, H., Posner, I., and Teh, Y. W · 2018
Earlier work this paper cites.
Stochastic adversarial video prediction
Lee, A. X., Zhang, R., Ebert, F., Abbeel, P., Finn, C., and Levine, S · 2018
Earlier work this paper cites.
Predicting future instance segmentation by forecasting convolutional features
Luc, P., Couprie, C., Lecun, Y., and Verbeek, J · 2018
Cited alongside, same era.
Tganv2: Efficient training of large models for video generation with multiple subsampling layers
Saito, M. and Saito, S · 2018
Cited alongside, same era.
Mocogan: Decomposing motion and content for video generation
Tulyakov, S., Liu, M.-Y., Yang, X., and Kautz, J · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation
Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A · 2019
Cited alongside, same era.
Compositional video synthesis with action graphs
Bar, A., Herzig, R., Wang, X., Rohrbach, A., Chechik, G., Darrell, T., and Globerson, A · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improved conditional vrnns for video prediction
Castrejon, L., Ballas, N., and Courville, A · 2019
Cited alongside, same era.
Adversarial video generation on complex datasets, 2019
Clark, A., Donahue, J., and Simonyan, K · 2019
Cited alongside, same era.
Genesis: Generative scene inference and sampling with object-centric latent representations
Engelcke, M., Kosiorek, A. R., Jones, O. P., and Posner, I · 2019
Cited alongside, same era.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
Girdhar, R. and Ramanan, D · 2019
Cited alongside, same era.
Multi-object representation learning with iterative variational inference
Greff, K., Kaufman, R. L., Kabra, R., Watters, N., Burgess, C., Zoran, D., Matthey, L., Botvinick, M., and Lerchner, A · 2019
Cited alongside, same era.
Scalor: Generative world models with scalable object representations
Jiang, J., Janghorbani, S., De Melo, G., and Ahn, S · 2019
Cited alongside, same era.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Cited alongside, same era.
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Later among the works it cites.
Transformation-based adversarial video prediction on large-scale data
Luc, P., Clark, A., Dieleman, S., Casas, D. d. L., Doron, Y., Cassirer, A., and Simonyan, K · 2020
Later among the works it cites.
Something-else: Compositional action recognition with spatial-temporal interaction networks
Materzynska, J., Xiao, T., Herzig, R., Xu, H., Wang, X., and Darrell, T · 2020
Later among the works it cites.
Rakhimov, R., Volkhonskiy, D., Artemov, A., Zorin, D., and Burnaev, E · 2020
Later among the works it cites.
Fitvid: Overfitting in pixel-level video prediction
Babaeizadeh, M., Saffar, M. T., Nair, S., Levine, S., Finn, C., and Erhan, D · 2021
Later among the works it cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Later among the works it cites.
Kubric: a scalable dataset generator
Greff, K., Belletti, F., Beyer, L., Doersch, C., Du, Y., Duckworth, D., Fleet, D. J., Gnanapragasam, D., Golemo, F., Herrmann, C., Kipf, T., Kundu, A., Lagun, D., Laradji, I., Liu, H.-T. D., Meyer, H., Miao, Y., Nowrouzezahrai, D., Oztireli, C., Pot, E., Radwan, N., Rebain, D., Sabour, S., Sajjadi, M. S. M., Sela, M., Sitzmann, V., Stone, A., Sun, D., Vora, S., Wang, Z., Wu, T., Yi, K. M., Zhong, F., and Tagliasacchi, A · 2021
Later among the works it cites.
Kabra, R., Zoran, D., Erdogan, G., Matthey, L., Creswell, A., Botvinick, M., Lerchner, A., and Burgess, C. P · 2021
Later among the works it cites.
Conditional object-centric learning from video
Kipf, T., Elsayed, G. F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., and Greff, K · 2021
Later among the works it cites.
Ccvs: Context-aware controllable video synthesis
Le Moing, G., Ponce, J., and Schmid, C · 2021
Later among the works it cites.
Revisiting hierarchical approach for persistent long-term video prediction
Lee, W., Jung, W., Zhang, H., Chen, T., Koh, J. Y., Huang, T., Yoon, H., Lee, H., and Hong, S · 2021
Later among the works it cites.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M · 2021
Later among the works it cites.
A good image generator is what you need for high-resolution video synthesis
Tian, Y., Ren, J., Chai, M., Olszewski, K., Peng, X., Metaxas, D. N., and Tulyakov, S · 2021
Later among the works it cites.
Walker, J., Razavi, A., and Oord, A. v. d · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Later among the works it cites.
Autoregressive latent video prediction with high-fidelity image generator
Anonymous · 2022
Closest in time.
Generating videos with dynamics-aware implicit generative adversarial networks
Yu, S., Tack, J., Mo, S., Kim, H., Kim, J., Ha, J.-W., and Shin, J · 2022
Closest in time.