Fetching the paper…
Reading the bibliography…
We aim to develop a model-based planning framework for world models that can be scaled with increasing model and data budgets for general-purpose manipulation tasks with only language and vision inputs.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Hill-climbing search
Bart Selman and Carla P Gomes · 2006
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
Alain Hore and Djemel Ziou · 2010
Earlier work this paper cites.
Consensus paper: roles of the cerebellum in motor control—the diversity of ideas on cerebellar involvement in movement
Mario Manto, James M Bower, Adriana Bastos Conforto, José M Delgado-García, Suzete Nascimento Farias Da Guarda, Marcus Gerwig, Christophe Habas, Nobuhiro Hagura, Richard B Ivry, Peter Mariën, et al · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma · 2013
Earlier work this paper cites.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Earlier work this paper cites.
Learning social affordance grammar from videos: Transferring human interactions to human-robot interactions
Tianmin Shu, Xiaofeng Gao, Michael S Ryoo, and Song-Chun Zhu · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
State representation learning for control: An overview
Timothée Lesort, Natalia Díaz-Rodríguez, Jean-Franois Goudou, and David Filliat · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine, and Google Brain · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly · 2018
Earlier work this paper cites.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Earlier work this paper cites.
Planning with goal-conditioned policies
Soroush Nasiriany, Vitchyr Pong, Steven Lin, and Sergey Levine · 2019
Earlier work this paper cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2020
Earlier work this paper cites.
Irasim: Learning interactive real-robot action simulators
Fangqi Zhu, Hongtao Wu, Song Guo, Yuxiao Liu, Chilam Cheang, and Tao Kong · 2020
Earlier work this paper cites.
Cril: Continual robot imitation learning via generative and prediction model
Chongkai Gao, Haichuan Gao, Shangqi Guo, Tianren Zhang, and Feng Chen · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Earlier work this paper cites.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
Lin Shao, Toki Migimatsu, Qiang Zhang, Karen Yang, and Jeannette Bohg · 2021
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Earlier work this paper cites.
Brains and algorithms partially converge in natural language processing. communications biology, 5 (1), 134, 2022
Charlotte Caucheteux and Jean-Rémi King · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2022
Earlier work this paper cites.
Transformers are sample-efficient world models
Vincent Micheli, Eloi Alonso, and François Fleuret · 2022
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Xirl: Cross-embodiment inverse reinforcement learning
Kevin Zakka, Andy Zeng, Pete Florence, Jonathan Tompson, Jeannette Bohg, and Debidatta Dwibedi · 2022
Cited alongside, same era.
Affordances from human videos as a versatile representation for robotics
Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, and Deepak Pathak · 2023
Cited alongside, same era.
Learning interactive real-world simulators
Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Later among the works it cites.
Aloha 2: An enhanced low-cost hardware for bimanual teleoperation
Jorge Aldaco, Travis Armstrong, Robert Baruch, Jeff Bingham, Sanky Chan, Kenneth Draper, Debidatta Dwibedi, Chelsea Finn, Pete Florence, Spencer Goodrich, et al · 2024
Closest in time.
Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation
Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta, and Shubham Tulsiani · 2024
Closest in time.
Inverse dynamics pretraining learns good representations for multitask imitation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
All are worth words: A vit backbone for diffusion models
Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Cited alongside, same era.
Yilun Du, Mengjiao Yang, Pete Florence, Fei Xia, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Joshua B Tenenbaum, et al · 2023
Cited alongside, same era.
Iterative interactive modeling for knotting plastic bags
Chongkai Gao, Zekun Li, Haichuan Gao, and Feng Chen · 2023
Cited alongside, same era.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Cited alongside, same era.
David Brandfonbrener, Ofir Nachum, and Joan Bruna · 2024
Closest in time.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Closest in time.
Genie: Generative interactive environments
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al · 2024
Closest in time.
Vegetable peeling: A case study in constrained dexterous manipulation
Tao Chen, Eric Cousineau, Naveen Kuppuswamy, and Pulkit Agrawal · 2024
Closest in time.
Keypoint action tokens enable in-context imitation learning in robotics
Norman Di Palo and Edward Johns · 2024
Closest in time.
Learning universal policies via text-guided video generation
Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, and Pieter Abbeel · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation
Wenlong Huang, Chen Wang, Yunzhu Li, and Fei-fei Li · 2024
Closest in time.
Doduo: Learning dense visual correspondence from unsupervised semantic-aware flow
Zhenyu Jiang, Hanwen Jiang, and Yuke Zhu · 2024
Closest in time.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al · 2024
Closest in time.
A path towards autonomous machine intelligence. 2022
Yann LeCun · 2024
Closest in time.
Dreamitate: Real-world visuomotor policy learning via video generation
Junbang Liang, Ruoshi Liu, Ege Ozguroglu, Sruthi Sudhakar, Achal Dave, Pavel Tokmakov, Shuran Song, and Carl Vondrick · 2024
Closest in time.
Latte: Latent diffusion transformer for video generation
Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao · 2024
Closest in time.
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al · 2024
Closest in time.
Generative image as action models
Mohit Shridhar, Yat Long Lo, and Stephen James · 2024
Closest in time.
Diffusion models are real-time game engines
Dani Valevski, Yaniv Leviathan, Moab Arar, and Shlomi Fruchter · 2024
Closest in time.
Lessons from learning to spin” pens”
Jun Wang, Ying Yuan, Haichuan Che, Haozhi Qi, Yi Ma, Jitendra Malik, and Xiaolong Wang · 2024
Closest in time.
ivideogpt: Interactive videogpts are scalable world models
Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, and Mingsheng Long · 2024
Closest in time.
General flow as foundation affordance for scalable robot learning
Chengbo Yuan, Chuan Wen, Tong Zhang, and Yang Gao · 2024
Closest in time.
Adaptigraph: Material-adaptive graph-based neural dynamics for robotic manipulation
Kaifeng Zhang, Baoyu Li, Kris Hauser, and Yunzhu Li · 2024
Closest in time.
Robodreamer: Learning compositional world models for robot imagination
Siyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li, Dit-Yan Yeung, and Chuang Gan · 2024
Closest in time.