Fetching the paper…
Reading the bibliography…
We present LangToMo, a vision-language-action framework structured as a dual-system architecture that uses pixel motion forecasts as intermediate representations.
Flight control in drosophila by visual perception of motion
Karl Georg Götz · 1968
Earlier work this paper cites.
Rheotropism in fishes
Geoff Arnold · 1974
Earlier work this paper cites.
Thinking, fast and slow, 2011
Daniel Kahneman · 2011
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation, 2015
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Dense optical flow prediction from a static image
Jacob Walker, Abhinav Kumar Gupta, and Martial Hebert · 2015
Earlier work this paper cites.
Optic flow stabilizes flight in ruby-throated hummingbirds
Ivo G. Ros and Andrew A. Biewener · 2016
Earlier work this paper cites.
Deep Visual Foresight for Planning Robot Motion
Chelsea Finn and Sergey Levine · 2017
Earlier work this paper cites.
Learning Robot Activities from First-Person Human Videos Using Convolutional Future Regression
Jangwon Lee and Michael S Ryoo · 2017
Earlier work this paper cites.
Universal sentence encoder
Daniel Matthew Cer, Yinfei Yang, Sheng yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil · 2018
Earlier work this paper cites.
Im2flow: Motion hallucination from static images for action recognition
Ruohan Gao, Bo Xiong, and Kristen Grauman · 2018
Earlier work this paper cites.
Learning Plannable Representations with Causal InfoGAN
Thanard Kurutach, Aviv Tamar, Ge Yang, Stuart J Russell, and Pieter Abbeel · 2018
Earlier work this paper cites.
Neural program synthesis from diverse demonstration videos
Shao-Hua Sun, Hyeonwoo Noh, Sriram Somasundaram, and Joseph Lim · 2018
Earlier work this paper cites.
Selflow: Self-supervised learning of optical flow
Pengpeng Liu, Michael Lyu, Irwin King, and Jia Xu · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Third-person visual imitation learning via decoupled hierarchical controller
Pratyusha Sharma, Deepak Pathak, and Abhinav Gupta · 2019
Earlier work this paper cites.
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2019
Earlier work this paper cites.
Flowcontrol: Optical flow based visual servoing
Max Argus, Lukás Hermann, Jon Long, and Thomas Brox · 2020
Earlier work this paper cites.
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Aggressive perception-aware navigation using deep optical flow dynamics and pixelmpc
Keuntaek Lee, Jason Gibson, and Evangelos A. Theodorou · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng · 2020
Earlier work this paper cites.
Learning optical flow from still images
Filippo Aleotti, Matteo Poggi, and Stefano Mattoccia · 2021
Earlier work this paper cites.
The effect of optic flow cues on honeybee flight control in wind
Emily Baird, Norbert Boeddeker, and Mandyam V. Srinivasan · 2021
Earlier work this paper cites.
Learning Generalizable Robotic Reward Functions from ”In-The-Wild” Human Videos
Annie S Chen, Suraj Nair, and Chelsea Finn · 2021
Earlier work this paper cites.
Enhancing optical-flow-based control by learning visual appearance cues for flying robots
Guido C.H.E. de Croon, Christophe de Wagter, and Tobias Seidl · 2021
Earlier work this paper cites.
Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Earlier work this paper cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Oier Mees, Lukás Hermann, Erick Rosete-Beas, and Wolfram Burgard · 2021
Earlier work this paper cites.
Concept2Robot: Learning Manipulation Concepts from Instructions and Human Demonstrations
Lin Shao, Toki Migimatsu, Qiang Zhang, Karen Yang, and Jeannette Bohg · 2021
Earlier work this paper cites.
Human-to-robot imitation in the wild
Shikhar Bahl, Abhinav Gupta, and Deepak Pathak · 2022
Earlier work this paper cites.
Long video generation with time-agnostic vqgan and time-sensitive transformer, 2022
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh · 2022
Cited alongside, same era.
Video Diffusion Models
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet · 2022
Cited alongside, same era.
Planning with Diffusion for Flexible Behavior Synthesis
Michael Janner, Yilun Du, Joshua B. Tenenbaum, and Sergey Levine · 2022
Cited alongside, same era.
R3M: A Universal Visual Representation for Robot Manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Cited alongside, same era.
The Surprising Effectiveness of Representation Learning for Visual Imitation
Jyothish Pari, Nur Muhammad Shafiullah, Sridhar Pandian Arunachalam, and Lerrel Pinto · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Rt-h: Action hierarchies using language
Suneel Belkhale, Tianli Ding, Ted Xiao, Pierre Sermanet, Quon Vuong, Jonathan Tompson, Yevgen Chebotar, Debidatta Dwibedi, and Dorsa Sadigh · 2024
Later among the works it cites.
π \pi 0: A vision-language-action flow model for general robot control
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky · 2024
Later among the works it cites.
Seeing through pixel motion: Learning obstacle avoidance from optical flow with one camera
Yu Hu, Yuang Zhang, Yunlong Song, Yang Deng, Feng Yu, Linzuo Zhang, Weiyao Lin, Danping Zou, and Wenxian Yu · 2024
Later among the works it cites.
Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation
Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, and Fei-Fei Li · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Cited alongside, same era.
A generalist agent
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Freitas de Nando · 2022
Cited alongside, same era.
Diffusion motion: Generate text-guided 3d human motion by diffusion model
Zhiyuan Ren, Zhihong Pan, Xin Zhou, and Le Kang · 2022
Cited alongside, same era.
Pixel-level correspondence for self-supervised learning from video
Yash Sharma, Yi Zhu, Chris Russell, and Thomas Brox · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data, 2022
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman · 2022
Cited alongside, same era.
Robotic Telekinesis: Learning a Robotic Hand Imitator by Watching Humans on Youtube
Aravind Sivakumar, Kenneth Shaw, and Deepak Pathak · 2022
Cited alongside, same era.
Phenaki: Variable length video generation from open domain textual description, 2022
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan · 2022
Cited alongside, same era.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al · 2024
Later among the works it cites.
Onlyflow: Optical flow based motion conditioning for video diffusion models, 2024
Mathis Koroglu, Hugo Caselles-Dupr’e, Guillaume Jeanneret Sanmiguel, and Matthieu Cord · 2024
Later among the works it cites.
Llara: Supercharging robot learning data for vision-language policy
Xiang Li, Cristina Mata, Jong Sung Park, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan Burgert, Mu Cai, Yong Jae Lee, and Michael S. Ryoo · 2024
Later among the works it cites.
Movideo: Motion-aware video generation with diffusion model
Jingyun Liang, Yuchen Fan, Kai Zhang, Radu Timofte, Luc van Gool, and Rakesh Ranjan · 2024
Later among the works it cites.
Flowdiffuser: Advancing optical flow estimation with diffusion models
Ao Luo, Xin Li, Fan Yang, Jiangyu Liu, Haoqiang Fan, and Shuaicheng Liu · 2024
Later among the works it cites.
Llarva: Vision-action instruction tuning enhances robot learning
Dantong Niu, Yuvan Sharma, Giscard Biamby, Jerome Quenum, Yutong Bai, Baifeng Shi, Trevor Darrell, and Roei Herzig · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Charles Xu, Jianlan Luo, Tobias Kreiman, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine · 2024
Later among the works it cites.
Generative image as action models
Mohit Shridhar, Yat Long Lo, and Stephen James · 2024
Later among the works it cites.
Controlling the world by sleight of hand
Sruthi Sudhakar, Ruoshi Liu, Basile Van Hoorick, Carl Vondrick, and Richard Zemel · 2024
Later among the works it cites.
Predictive inverse dynamics models are scalable learners for robotic manipulation
Yang Tian, Sizhe Yang, Jia Zeng, Ping Wang, Dahua Lin, Hao Dong, and Jiangmiao Pang · 2024
Later among the works it cites.
Flow as the cross-domain manipulation interface
Mengda Xu, Zhenjia Xu, Yinghao Xu, Cheng Chi, Gordon Wetzstein, Manuela Veloso, and Shuran Song · 2024
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning
Michal Zawalski, William Chen, Karl Pertsch, Oier Mees, Chelsea Finn, and Sergey Levine · 2024
Later among the works it cites.
Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao, Hal Daumé III, Andrey Kolobov, Furong Huang, and Jianwei Yang · 2024
Later among the works it cites.
Chi-Lam Cheang, Sijin Chen, Zhongren Cui, Yingdong Hu, Liqun Huang, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Xiao Ma, Hao Niu, Wenxuan Ou, Wanli Peng, Zeyu Ren, Haixin Shi, Jiawen Tian, Hongtao Wu, Xin Xiao, Yuyang Xiao, Jiafeng Xu, and Yichu Yang · 2025
Closest in time.
Flip: Flow-centric generative planning as general-purpose manipulation world model
Chongkai Gao, Haozhuo Zhang, Zhixuan Xu, Zhehao Cai, and Lin Shao · 2025
Closest in time.
Video prediction policy: A generalist robot policy with predictive visual representations
Yucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, and Jianyu Chen · 2025
Closest in time.
π \pi 0.5: a vision-language-action model with open-world generalization, 2025
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, James Tanner, Quan Vuong, Homer Rich Walke, Anna Walling, Haohuan Wang, Lili Yu, and Ury Zhilinsky · 2025
Closest in time.
Object-centric world model for language-guided manipulation
Youngjoon Jeong, Junha Chun, Soonwoo Cha, and Taesup Kim · 2025
Closest in time.
Molmoact: Action reasoning models that can reason in space
Jason Lee, Jiafei Duan, Haoquan Fang, Yuquan Deng, Shuo Liu, Boyang Li, Bohan Fang, Jieyu Zhang, Yi Ru Wang, Sangho Lee, Winson Han, Wilbert Pumacay, Angelica Wu, Rose Hendrix, Karen Farley, Eli VanderBilt, Ali Farhadi, Dieter Fox, and Ranjay Krishna · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
Nvidia, Johan Bjorck, et al · 2025
Closest in time.
Motion tracks: A unified representation for human-robot transfer in few-shot imitation learning
Juntao Ren, Priya Sundaresan, Dorsa Sadigh, Sanjiban Choudhury, and Jeannette Bohg · 2025
Closest in time.
History-guided video diffusion, 2025
Kiwhan Song, Boyuan Chen, Max Simchowitz, Yilun Du, Russ Tedrake, and Vincent Sitzmann · 2025
Closest in time.
Dreamvla: A vision-language-action model dreamed with comprehensive world knowledge
Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang, Xinqiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He, He Wang, Zhizheng Zhang, Li Yi, Wenjun Zeng, and Xin Jin · 2025
Closest in time.
Universal actions for enhanced embodied foundation models
Jinliang Zheng, Jianxiong Li, Dongxiu Liu, Yinan Zheng, Zhihao Wang, Zhonghong Ou, Yu Liu, Jingjing Liu, Ya-Qin Zhang, and Xianyuan Zhan · 2025
Closest in time.