Fetching the paper…
Reading the bibliography…
Next-Token Prediction (NTP) is a de facto approach for autoregressive (AR) video generation, but it suffers from suboptimal unidirectional dependencies and slow inference speed.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A., and Shah, M · 2012
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2015
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the Kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O. K., and Socher, R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
A short note about kinetics-600
Carreira, J., Noland, E., Banki-Horvath, A., Hillier, C., and Zisserman, A · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2018
Earlier work this paper cites.
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N. M., and Uszkoreit, J · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
Semi-autoregressive neural machine translation
Wang, C., Zhang, J., and Chen, H · 2018
Earlier work this paper cites.
Adversarial video generation on complex datasets
Clark, A., Donahue, J., and Simonyan, K · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., teusz Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Generative pretraining from pixels
Chen, M., Radford, A., Wu, J., Jun, H., Dhariwal, P., Luan, D., and Sutskever, I · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Rakhimov, R., Volkhonskiy, D., Artemov, A., Zorin, D., and Burnaev, E · 2020
Earlier work this paper cites.
Scaling autoregressive video models
Weissenborn, D., Täckström, O., and Uszkoreit, J · 2020
Earlier work this paper cites.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M · 2021
Earlier work this paper cites.
Walker, J., Razavi, A., and Oord, A. v. d · 2021
Earlier work this paper cites.
Videogpt: Video generation using vq-vae and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Cited alongside, same era.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and R’e, C · 2022
Cited alongside, same era.
Long video generation with time-agnostic vqgan and time-sensitive transformer
Ge, S., Hayes, T., Yang, H., Yin, X., Pang, G., Jacobs, D., Huang, J.-B., and Parikh, D · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A. A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., and Salimans, T · 2022
Finite scalar quantization: Vq-vae made simple
Mentzer, F., Minnen, D. C., Agustsson, E., and Tschannen, M · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Testa: Temporal-spatial token aggregation for long-form video-language understanding
Ren, S., Chen, S., Li, S., Sun, X., and Hou, L · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Cited alongside, same era.
xformers: A modular and hackable transformer modelling library
Lefaudeux, B., Massa, F., Liskovich, D., Xiong, W., Caggiano, V., Naren, S., Xu, M., Hu, J., Tintore, M., Zhang, S., Labatut, P., Haziza, D., Wehrstedt, L., Reizenstein, J., and Sizov, G · 2022
Cited alongside, same era.
Improved masked image generation with token-critic
Lezama, J., Chang, H., Jiang, L., and Essa, I · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Cited alongside, same era.
Phenaki: Variable length video generation from open domain textual descriptions
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Cited alongside, same era.
Diffusion models: A comprehensive survey of methods and applications
Yang, L., Zhang, Z., Hong, S., Xu, R., Zhao, Y., Shao, Y., Zhang, W., Yang, M.-H., and Cui, B · 2022
Cited alongside, same era.
Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S · 2023
Later among the works it cites.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation
Xia, H., Ge, T., Wang, P., Chen, S.-Q., Wei, F., and Sui, Z · 2023
Later among the works it cites.
Pre-trained language models do not help auto-regressive text-to-image generation
Zhang, Y., McKinzie, B., Gan, Z., Shankar, V., and Toshev, A · 2023
Later among the works it cites.
Genie: Generative interactive environments
Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al · 2024
Later among the works it cites.
PCA-bench: Evaluating multimodal large language models in perception-cognition-action chain
Chen, L., Zhang, Y., Ren, S., Zhao, H., Cai, Z., Wang, Y., Wang, P., Meng, X., Liu, T., and Chang, B · 2024
Later among the works it cites.
Next token prediction towards multimodal intelligence: A comprehensive survey
Chen, L., Wang, Z., Ren, S., Li, L., Zhao, H., Li, Y., Cai, Z., Guo, H., Zhang, L., Xiong, Y., et al · 2024
Later among the works it cites.
Causal diffusion transformers for generative modeling
Deng, C., Zh, D., Li, K., Guan, S., and Fan, H · 2024
Later among the works it cites.
Better & faster large language models via multi-token prediction
Gloeckle, F., Idrissi, B. Y., Rozière, B., Lopez-Paz, D., and Synnaeve, G · 2024
Later among the works it cites.
Autoregressive image generation without vector quantization
Li, T., Tian, Y., Li, H., Deng, M., and He, K · 2024
Later among the works it cites.
Liu, D., Zhao, S., Zhuo, L., Lin, W., Qiao, Y., Li, H., and Gao, P · 2024
Later among the works it cites.
Chatgpt: Chat generative pre-trained transformer
OpenAI · 2024
Later among the works it cites.
Openai o1
OpenAI · 2024
Later among the works it cites.
Hello gpt-4o
OpenAI · 2024
Later among the works it cites.
Randar: Decoder-only autoregressive visual generation in random orders
Pang, Z., Zhang, T., Luan, F., Man, Y., Tan, H., Zhang, K., Freeman, W. T., and Wang, Y.-X · 2024
Later among the works it cites.
Hierarchical patch diffusion models for high-resolution video generation
Skorokhodov, I., Menapace, W., Siarohin, A., and Tulyakov, S · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Team, C · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., Yin, D., Gu, X., Zhang, Y., Wang, W., Cheng, Y., Liu, T., Xu, B., Dong, Y., and Tang, J · 2024
Later among the works it cites.
Language model beats diffusion–tokenizer is key to visual generation
Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Gupta, A., Gu, X., Hauptmann, A. G., et al · 2024
Later among the works it cites.