Fetching the paper…
Reading the bibliography…
This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions.
Determining optical flow
Berthold KP Horn and Brian G Schunck · 1981
Earlier work this paper cites.
An iterative image registration technique with an application to stereo vision
Bruce D Lucas and Takeo Kanade · 1981
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
A framework for the robust estimation of optical flow
Michael J Black and Padmanabhan Anandan · 1993
Earlier work this paper cites.
Three-dimensional scene flow
Sundar Vedula, Simon Baker, Peter Rander, Robert Collins, and Takeo Kanade · 1999
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
Lucas/kanade meets horn/schunck: Combining local and global optic flow methods
Andrés Bruhn, Joachim Weickert, and Christoph Schnörr · 2005
Earlier work this paper cites.
Scope of validity of psnr in image/video quality assessment
Quan Huynh-Thu and Mohammed Ghanbari · 2008
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Se3-nets: Learning rigid body motion using deep neural networks
Arunkumar Byravan and Dieter Fox · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Optical flow estimation using a spatial pyramid network
Anurag Ranjan and Michael J Black · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
A Waswani, N Shazeer, N Parmar, J Uszkoreit, L Jones, A Gomez, L Kaiser, and I Polosukhin · 2017
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Frederik Ebert, Chelsea Finn, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Controllable video generation with sparse trajectories
Zekun Hao, Xun Huang, and Serge Belongie · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2018
Earlier work this paper cites.
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
4d spatio-temporal convnets: Minkowski convolutional neural networks
Christopher Choy, JunYoung Gwak, and Silvio Savarese · 2019
Earlier work this paper cites.
Flownet3d: Learning scene flow in 3d point clouds
Xingyu Liu, Charles R Qi, and Leonidas J Guibas · 2019
Earlier work this paper cites.
Unsupervised learning of object structure and dynamics from videos
Matthias Minderer, Chen Sun, Ruben Villegas, Forrester Cole, Kevin P Murphy, and Honglak Lee · 2019
Earlier work this paper cites.
High fidelity video prediction with large stochastic recurrent neural networks
Ruben Villegas, Arkanath Pathak, Harini Kannan, Dumitru Erhan, Quoc V Le, and Honglak Lee · 2019
Earlier work this paper cites.
Volumetric correspondence networks for optical flow
Gengshan Yang and Deva Ramanan · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Model based reinforcement learning for atari
Łukasz Kaiser, Mohammad Babaeizadeh, Piotr Miłos, Błażej Osiński, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Cited alongside, same era.
Diffusion for world modeling: Visual details matter in atari
Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce, and François Fleuret · 2024
Later among the works it cites.
Zero-shot robotic manipulation with pre-trained image-editing diffusion models
Kevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Rich Walke, Chelsea Finn, Aviral Kumar, and Sergey Levine · 2024
Later among the works it cites.
Video generation models as world simulators, 2024
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Later among the works it cites.
Genie: Generative interactive environments
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al · 2024
Later among the works it cites.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng · 2020
Cited alongside, same era.
Learning 3d dynamic scene representations for robot manipulation
Zhenjia Xu, Zhanpeng He, Jiajun Wu, and Shuran Song · 2020
Cited alongside, same era.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, Roberto Martín-Martín, Abhishek Joshi, Soroush Nasiriany, and Yifeng Zhu · 2020
Cited alongside, same era.
Fitvid: Overfitting in pixel-level video prediction
Mohammad Babaeizadeh, Mohammad Taghi Saffar, Suraj Nair, Sergey Levine, Chelsea Finn, and Dumitru Erhan · 2021
Cited alongside, same era.
Mastering atari with discrete world models
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Cited alongside, same era.
Scalable scene flow from point clouds in the real world
Philipp Jund, Chris Sweeney, Nichola Abdo, Zhifeng Chen, and Jonathon Shlens · 2021
Cited alongside, same era.
Robodesk: A multi-task reinforcement learning benchmark, 2021
Harini Kannan, Danijar Hafner, Chelsea Finn, and Dumitru Erhan · 2021
Cited alongside, same era.
Later among the works it cites.
Automated creation of digital cousins for robust policy learning
Tianyuan Dai, Josiah Wong, Yunfan Jiang, Chen Wang, Cem Gokmen, Ruohan Zhang, Jiajun Wu, and Li Fei-Fei · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Later among the works it cites.
Motion prompting: Controlling video generation with motion trajectories
Daniel Geng, Charles Herrmann, Junhwa Hur, Forrester Cole, Serena Zhang, Tobias Pfaff, Tatiana Lopez-Guevara, Carl Doersch, Yusuf Aytar, Michael Rubinstein, et al · 2024
Later among the works it cites.
I can’t believe it’s not scene flow!
Ishan Khatri, Kyle Vedder, Neehar Peri, Deva Ramanan, and James Hays · 2024
Later among the works it cites.
Dense optical tracking: connecting the dots
Guillaume Le Moing, Jean Ponce, and Cordelia Schmid · 2024
Later among the works it cites.
Learning to model the world with language
Jessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner, Pieter Abbeel, Dan Klein, and Anca Dragan · 2024
Later among the works it cites.
Icp-flow: Lidar scene flow estimation with icp
Yancong Lin and Holger Caesar · 2024
Later among the works it cites.
Efficient world models with context-aware tokenization
Vincent Micheli, Eloi Alonso, and François Fleuret · 2024
Later among the works it cites.
Genie 2: A large-scale foundation world model, 2024
Jack Parker-Holder, Philip Ball, Jake Bruce, Vibhavari Dasagi, Kristian Holsheimer, Christos Kaplanis, Alexandre Moufarek, Guy Scully, Jeremy Shar, Jimmy Shi, Stephen Spencer, Jessica Yung, Michael Dennis, Sultan Kenjeyev, Shangbang Long, Vlad Mnih, Harris Chan, Maxime Gazeau, Bonnie Li, Fabio Pardo, Luyu Wang, Lei Zhang, Frederic Besse, Tim Harley, Anna Mitenkova, Jane Wang, Jeff Clune, Demis Hassabis, Raia Hadsell, Adrian Bolton, Satinder Singh, and Tim Rocktäschel · 2024
Later among the works it cites.
Avid: Adapting video diffusion models to world models
Marc Rigter, Tarun Gupta, Agrin Hilmkil, and Chao Ma · 2024
Later among the works it cites.
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling
Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian, Dasong Li, Yi Zhang, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, et al · 2024
Later among the works it cites.
Learning latent dynamic robust representations for world models
Ruixiang Sun, Hongyu Zang, Xin Li, and Riashat Islam · 2024
Later among the works it cites.
Diffusion models are real-time game engines
Dani Valevski, Yaniv Leviathan, Moab Arar, and Shlomi Fruchter · 2024
Later among the works it cites.
Zeroflow: Scalable scene flow via distillation
Kyle Vedder, Neehar Peri, Nathaniel Eliot Chodosh, Ishan Khatri, ERIC EATON, Dinesh Jayaraman, Yang Liu, Deva Ramanan, and James Hays · 2024
Later among the works it cites.
ivideogpt: Interactive videogpts are scalable world models
Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, and Mingsheng Long · 2024
Later among the works it cites.
Spatialtracker: Tracking any 2d pixels in 3d space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou · 2024
Later among the works it cites.
Flow as the cross-domain manipulation interface
Mengda Xu, Zhenjia Xu, Yinghao Xu, Cheng Chi, Gordon Wetzstein, Manuela Veloso, and Shuran Song · 2024
Later among the works it cites.
Learning interactive real-world simulators
Sherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson, Leslie Pack Kaelbling, Dale Schuurmans, and Pieter Abbeel · 2024
Later among the works it cites.
3d-vla: A 3d vision-language-action generative world model
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan · 2024
Later among the works it cites.
Robodreamer: Learning compositional world models for robot imagination
Siyuan Zhou, Yilun Du, Jiaben Chen, YANDONG LI, Dit-Yan Yeung, and Chuang Gan · 2024
Later among the works it cites.
Irasim: Learning interactive real-robot action simulators
Fangqi Zhu, Hongtao Wu, Song Guo, Yuxiao Liu, Chilam Cheang, and Tao Kong · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al · 2025
Closest in time.
Video depth anything: Consistent depth estimation for super-long videos
Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zilong Huang, Jiashi Feng, and Bingyi Kang · 2025
Closest in time.
Enerverse: Envisioning embodied future space for robotics manipulation
Siyuan Huang, Liliang Chen, Pengfei Zhou, Shengcong Chen, Zhengkai Jiang, Yue Hu, Peng Gao, Hongsheng Li, Maoqing Yao, and Guanghui Ren · 2025
Closest in time.
World and human action models towards gameplay ideation
Anssi Kanervisto, Dave Bignell, Linda Yilin Wen, Martin Grayson, Raluca Georgescu, Sergio Valcarcel Macua, Shan Zheng Tan, Tabish Rashid, Tim Pearce, Yuhan Cao, Abdelhak Lemkhenter, Chentian Jiang, Gavin Costello, Gunshi Gupta, Marko Tot, Shu Ishida, Tarun Gupta, Udit Arora, Ryen W. White, Sam Devlin, Cecily Morrison, and Katja Hofmann · 2025
Closest in time.
Depth anything v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2025
Closest in time.