Fetching the paper…
Reading the bibliography…
OpenAI's Sora highlights the potential of video generation for developing world models that adhere to fundamental physical laws.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P · 2004
Earlier work this paper cites.
Estimating the material properties of fabric from video
Bouman, K. L., Xiao, B., Battaglia, P., and Freeman, W. T · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Material recognition in the wild with the materials in context database
Bell, S., Upchurch, P., Snavely, N., and Bala, K · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik, P. K. and Ba, J · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Physics 101: Learning physical object properties from unlabeled videos
Wu, J., Lim, J. J., Zhang, H., Tenenbaum, J. B., and Freeman, W. T · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Video enhancement with task-oriented flow
Xue, T., Chen, B., Wu, J., Wei, D., and Freeman, W. T · 2017
Earlier work this paper cites.
Shapestacks: Learning vision-based physical intuition for generalised object stacking
Groth, O., Fuchs, F. B., Posner, I., and Vedaldi, A · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Phyre: A new benchmark for physical reasoning
Bakhtin, A., van der Maaten, L., Johnson, J., Gustafson, L., and Girshick, R · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Earlier work this paper cites.
Clevrer: Collision events for video representation and reasoning
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B · 2019
Earlier work this paper cites.
Rapid trial-and-error learning with simulation supports flexible tool use and physical reasoning
Allen, K. R., Smith, K. A., and Tenenbaum, J. B · 2020
Earlier work this paper cites.
Craft: A benchmark for causal reasoning about forces and interactions
Ates, T., Atesoglu, M. S., Yigit, C., Kesen, I., Kobas, M., Erdem, E., Erdem, A., Goksun, T., and Yuret, D · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B · 2020
Earlier work this paper cites.
Discovery of physics from data: Universal laws and discrepancies
de Silva, B. M., Higdon, D. M., Brunton, S. L., and Kutz, J. N · 2020
Earlier work this paper cites.
Forward prediction for physical reasoning
Girdhar, R., Gustafson, L., Adcock, A., and van der Maaten, L · 2020
Earlier work this paper cites.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Swap attention in spatiotemporal diffusions for text-to-video generation
Wang, W., Yang, H., Tuo, Z., He, H., Zhu, J., Fu, J., and Liu, J · 2023
Later among the works it cites.
Perception and simulation during concept learning
Weitnauer, E., Goldstone, R. L., and Ritter, H · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Later among the works it cites.
Phy-q as a measure for physical reasoning intelligence
Xue, C., Pinto, V., Gamage, C., Nikonova, E., Zhang, P., and Renz, J · 2023
Later among the works it cites.
Learning interactive real-world simulators
Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Panda: A gigapixel-level human-centric video dataset
Wang, X., Zhang, X., Zhu, Y., Guo, Y., Yuan, X., Xiang, L., Wang, Z., Ding, G., Brady, D., Dai, Q., and Fang, L · 2020
Cited alongside, same era.
How neural networks extrapolate: From feedforward to graph neural networks
Xu, K., Zhang, M., Li, J., Du, S. S., Kawarabayashi, K.-i., and Jegelka, S · 2020
Cited alongside, same era.
Learning in high dimension always amounts to extrapolation
Balestriero, R., Pesenti, J., and LeCun, Y · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Visual representation learning does not generalize strongly within the same domain
Schott, L., Von Kügelgen, J., Träuble, F., Gehler, P., Russell, C., Bethge, M., Schölkopf, B., Locatello, F., and Brendel, W · 2021
Cited alongside, same era.
Laion-aesthetics v1
Beaumont, R. and Schuhmann, C · 2022
Cited alongside, same era.
Coyo-700m: Image-text pair dataset
Byeon, M., Park, B., Kim, H., Lee, S., Baek, W., and Kim, S · 2022
Cited alongside, same era.
Value-consistent representation learning for data-efficient reinforcement learning
Yue, Y., Kang, B., Xu, Z., Huang, G., and Yan, S · 2023
Later among the works it cites.
Videophy: Evaluating physical commonsense for video generation
Bansal, H., Lin, Z., Xie, T., Zong, Z., Yarom, M., Bitton, Y., Jiang, C., Sun, Y., Chang, K.-W., and Grover, A · 2024
Closest in time.
Video generation models as world simulators
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A · 2024
Closest in time.
Genie: Generative interactive environments
Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al · 2024
Closest in time.
Teaching video diffusion model with latent physical phenomenon knowledge
Cao, Q., Wang, D., Li, X., Chen, Y., Ma, C., and Yang, X · 2024
Closest in time.
Compositional generative modeling: A single model is not all you need
Du, Y. and Kaelbling, L · 2024
Closest in time.
Vista: A generalizable driving world model with high fidelity and versatile controllability
Gao, S., Yang, J., Chen, L., Chitta, K., Qiu, Y., Geiger, A., Zhang, J., and Li, H · 2024
Closest in time.
Diffusion models in low-level vision: A survey
He, C., Shen, Y., Fang, C., Xiao, F., Tang, L., Zhang, Y., Zuo, W., Guo, Z., and Li, X · 2024
Closest in time.
Case-based or rule-based: How do transformers do the math?
Hu, Y., Tang, X., Yang, H., and Zhang, M · 2024
Closest in time.
Vbench: Comprehensive benchmark suite for video generative models
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al · 2024
Closest in time.
Evaluation of text-to-video generation models: A dynamics perspective
Liao, M., Ye, Q., Zuo, W., Wan, F., Wang, T., Zhao, Y., Wang, J., Zhang, X., et al · 2024
Closest in time.
Common diffusion noise schedules and sample steps are flawed
Lin, S., Liu, B., Li, J., and Yang, X · 2024
Closest in time.
Evalcrafter: Benchmarking and evaluating large video generation models
Liu, Y., Cun, X., Liu, X., Wang, X., Zhang, Y., Chen, H., Liu, Y., Zeng, T., Chan, R., and Shan, Y · 2024
Closest in time.
Natural language instructions induce compositional generalization in networks of neurons
Riveland, R. and Pouget, A · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
T2v-compbench: A comprehensive benchmark for compositional text-to-video generation
Sun, K., Huang, K., Liu, X., Wu, Y., Xu, Z., Li, Z., and Liu, X · 2024
Closest in time.
Generalized predictive model for autonomous driving
Yang, J., Gao, S., Qiu, Y., Chen, L., Li, T., Dai, B., Chitta, K., Wu, P., Zeng, J., Luo, P., et al · 2024
Closest in time.
Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution
Yue, Y., Wang, Y., Kang, B., Han, Y., Wang, S., Song, S., Feng, J., and Huang, G · 2024
Closest in time.
Make pixels dance: High-dynamic video generation
Zeng, Y., Wei, G., Zheng, J., Zou, J., Wei, Y., Zhang, Y., and Li, H · 2024
Closest in time.
Genad: Generative end-to-end autonomous driving
Zheng, W., Song, R., Guo, X., and Chen, L · 2024
Closest in time.