Fetching the paper…
Reading the bibliography…
Tokenizer, serving as a translator to map the intricate visual data into a compact latent space, lies at the core of visual generative models.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation
T. Karras, T. Aila, S. Laine, and J. Lehtinen · 2017
Earlier work this paper cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
A short note about kinetics-600
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz · 2018
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
A. Brock, J. Donahue, and K. Simonyan · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Earlier work this paper cites.
Moments in time dataset: one million videos for event understanding
M. Monfort, A. Andonian, B. Zhou, K. Ramakrishnan, S. A. Bargal, T. Yan, L. Brown, Q. Fan, D. Gutfreund, C. Vondrick, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
A. Razavi, A. Van den Oord, and O. Vinyals · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Earlier work this paper cites.
Generative pretraining from pixels
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever · 2020
Earlier work this paper cites.
Generative adversarial networks
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
R. Rakhimov, D. Volkhonskiy, A. Artemov, D. Zorin, and E. Burnaev · 2020
Earlier work this paper cites.
Scaling autoregressive video models
D. Weissenborn, O. Täckström, and J. Uszkoreit · 2020
Cited alongside, same era.
Is space-time attention all you need for video understanding?
G. Bertasius, H. Wang, and L. Torresani · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Objectformer for image manipulation detection and localization
J. Wang, Z. Wu, J. Chen, X. Han, A. Shrivastava, S.-N. Lim, and Y.-G. Jiang · 2022
Later among the works it cites.
Bevt: Bert pretraining of video transformers
R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, Y.-G. Jiang, L. Zhou, and L. Yuan · 2022
Later among the works it cites.
Vector-quantized image modeling with improved vqgan
J. Yu, X. Li, J. Y. Koh, H. Zhang, R. Pang, J. Qin, A. Ku, Y. Xu, J. Baldridge, and Y. Wu · 2022
Later among the works it cites.
Generating videos with dynamics-aware implicit generative adversarial networks
S. Yu, J. Tack, S. Mo, H. Kim, J. Kim, J.-W. Ha, and J. Shin · 2022
Later among the works it cites.
Scaling vision transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Later among the works it cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Cited alongside, same era.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Cited alongside, same era.
A good image generator is what you need for high-resolution video synthesis
Y. Tian, J. Ren, M. Chai, K. Olszewski, X. Peng, D. N. Metaxas, and S. Tulyakov · 2021
Cited alongside, same era.
J. Walker, A. Razavi, and A. v. d. Oord · 2021
Cited alongside, same era.
Videogpt: Video generation using vq-vae and transformers
W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas · 2021
Cited alongside, same era.
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, et al · 2023
Later among the works it cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang · 2023
Later among the works it cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2023
Later among the works it cites.
Mostgan-v: Video generation with temporal motion styles
X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Resformer: Scaling vits with multi-resolution training
R. Tian, Z. Wu, Q. Dai, H. Hu, Y. Qiao, and Y.-G. Jiang · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Magvit: Masked generative video transformer
L. Yu, Y. Cheng, K. Sohn, J. Lezama, H. Zhang, H. Chang, A. G. Hauptmann, M.-H. Yang, Y. Hao, I. Essa, et al · 2023
Later among the works it cites.
Video probabilistic diffusion models in projected latent space
S. Yu, K. Sohn, S. Kim, and J. Shin · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Latte: Latent diffusion transformer for video generation
X. Ma, Y. Wang, G. Jia, X. Chen, Z. Liu, Y.-F. Li, C. Chen, and Y. Qiao · 2024
Closest in time.
Autoregressive model beats diffusion: Llama for scalable image generation
P. Sun, Y. Jiang, S. Chen, S. Zhang, B. Peng, P. Luo, and Z. Yuan · 2024
Closest in time.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
K. Tian, Y. Jiang, Z. Yuan, B. Peng, and L. Wang · 2024
Closest in time.
Omnivid: A generative framework for universal video understanding
J. Wang, D. Chen, C. Luo, B. He, L. Yuan, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Simda: Simple diffusion adapter for efficient video generation
Z. Xing, Q. Dai, H. Hu, Z. Wu, and Y.-G. Jiang · 2024
Closest in time.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2024
Closest in time.
Language model beats diffusion–tokenizer is key to visual generation
L. Yu, J. Lezama, N. B. Gundavarapu, L. Versari, K. Sohn, D. Minnen, Y. Cheng, A. Gupta, X. Gu, A. G. Hauptmann, et al · 2024
Closest in time.