Fetching the paper…
Reading the bibliography…
Long-context video modeling is essential for enabling generative models to function as world simulators, as they must maintain temporal coherence over extended time spans.
2012
Earlier work this paper cites.
F. Ebert, C. Finn, A. X. Lee, and S. Levine, “Self-supervised visual planning with temporal skip connections.” CoRL , vol. 12, no. 16, p. 23, 2017
2017
Earlier work this paper cites.
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “Fvd: A new metric for video generation,” 2019
2019
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
V. Saxena, J. Ba, and D. Hafner, “Clockwork variational autoencoders,” Advances in Neural Information Processing Systems , vol. 34, pp. 29 246–29 257, 2021
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
C. Hawthorne, A. Jaegle, C. Cangea, S. Borgeaud, C. Nash, M. Malinowski, S. Dieleman, O. Vinyals, M. Botvinick, I. Simon et al. , “General-purpose, long-context autoregressive modeling with perceiver ar,” in International Conference on Machine Learning . PMLR, 2022, pp. 8535–8558
2022
Earlier work this paper cites.
W. Harvey, S. Naderiparizi, V. Masrani, C. Weilbach, and F. Wood, “Flexible diffusion modeling of long videos,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 953–27 965, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh, “Long video generation with time-agnostic vqgan and time-sensitive transformer,” in European Conference on Computer Vision . Springer, 2022, pp. 102–118
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
V. Voleti, A. Jolicoeur-Martineau, and C. Pal, “Mcvd-masked conditional video diffusion for prediction, generation, and interpolation,” Advances in neural information processing systems , vol. 35, pp. 23 371–23 385, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Wu, Z. Fan, X. Liu, H.-T. Zheng, Y. Gong, J. Jiao, J. Li, J. Guo, N. Duan, W. Chen et al. , “Ar-diffusion: Auto-regressive diffusion model for text generation,” Advances in Neural Information Processing Systems , vol. 36, pp. 39 957–39 974, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
bloc97, “NTK-Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation.” 2023. [Online]. Available: https://www.reddit.com/r/LocalLLaMA/comments/14lz7j5/ntkaware_scaled_rope_allows_llama_models_to_have/
2023
Cited alongside, same era.
2023
L. Barrault, P.-A. Duquenne, M. Elbayad, A. Kozhevnikov, B. Alastruey, P. Andrews, M. Coria, G. Couairon, M. R. Costa-jussà, D. Dale et al. , “Large concept models: Language modeling in a sentence representation space,” arXiv e-prints , pp. arXiv–2412, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps et al. , “Genie: Generative interactive environments,” in Forty-first International Conference on Machine Learning , 2024
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2023
Cited alongside, same era.
W. Yan, D. Hafner, S. James, and P. Abbeel, “Temporally consistent transformers for video generation,” in International Conference on Machine Learning . PMLR, 2023, pp. 39 062–39 098
2023
Cited alongside, same era.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 4195–4205
2023
Cited alongside, same era.
H. Ni, C. Shi, K. Li, S. X. Huang, and M. R. Min, “Conditional image-to-video generation with latent flow diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2023, pp. 18 444–18 455
2023
Cited alongside, same era.
K. Mei and V. Patel, “Vidm: Video implicit diffusion models,” in Proceedings of the AAAI conference on artificial intelligence , vol. 37, no. 8, 2023, pp. 9117–9125
2023
Cited alongside, same era.
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh, “Video generation models as world simulators,” 2024. [Online]. Available: https://openai.com/research/video-generation-models-as-world-simulators
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Later among the works it cites.
J. Parker-Holder, P. Ball, J. Bruce, V. Dasagi, K. Holsheimer, C. Kaplanis, A. Moufarek, G. Scully, J. Shar, J. Shi, S. Spencer, J. Yung, M. Dennis, S. Kenjeyev, S. Long, V. Mnih, H. Chan, M. Gazeau, B. Li, F. Pardo, L. Wang, L. Zhang, F. Besse, T. Harley, A. Mitenkova, J. Wang, J. Clune, D. Hassabis, R. Hadsell, A. Bolton, S. Singh, and T. Rocktäschel, “Genie 2: A large-scale foundation world model,” 2024. [Online]. Available: https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/
2024
Later among the works it cites.
N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie, “Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers,” in European Conference on Computer Vision . Springer, 2024, pp. 23–40
2024
Later among the works it cites.
2024
Later among the works it cites.
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel et al. , “Scaling rectified flow transformers for high-resolution image synthesis,” in Forty-first international conference on machine learning , 2024
2024
Later among the works it cites.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing , vol. 568, p. 127063, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Zhang, J. Hu, W. Cheng, D. Paudel, and J. Yang, “Extdm: Distribution extrapolation diffusion model for video prediction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 19 310–19 320
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. Chen, D. Martí Monsó, Y. Du, M. Simchowitz, R. Tedrake, and V. Sitzmann, “Diffusion forcing: Next-token prediction meets full-sequence diffusion,” Advances in Neural Information Processing Systems , vol. 37, pp. 24 081–24 125, 2025
2025
Closest in time.
T. Li, Y. Tian, H. Li, M. Deng, and K. He, “Autoregressive image generation without vector quantization,” Advances in Neural Information Processing Systems , vol. 37, pp. 56 424–56 445, 2025
2025
Closest in time.