Fetching the paper…
Reading the bibliography…
We explore how to enhance next-token prediction models to perform in-context imitation learning on a real robot, where the robot executes new tasks by interpreting contextual information provided during the input phase, without updating its underlying policy parameters.
“ALVINN: An Autonomous Land Vehicle in a Neural Network”
Dean. Pomerleau · 1988
Earlier work this paper cites.
“A survey of robot learning from demonstration”
Brenna Argall, Sonia Chernova, Manuela Veloso and Brett Browning · 2009
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database”
Jia Deng et al · 2009
Earlier work this paper cites.
“End-to-end training of deep visuomotor policies”
Sergey Levine, Chelsea Finn, Trevor Darrell and Pieter Abbeel · 2016
Earlier work this paper cites.
“Model-agnostic meta-learning for fast adaptation of deep networks”
Chelsea Finn, Pieter Abbeel and Sergey Levine · 2017
Earlier work this paper cites.
“One-shot visual imitation learning via meta-learning”
Chelsea Finn et al · 2017
Earlier work this paper cites.
“One-shot imitation learning”
Yan Duan et al · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Jacquard: A large scale dataset for robotic grasp detection”
Amaury Depierre, Emmanuel Dellandréa and Liming Chen · 2018
Earlier work this paper cites.
“Scalable deep reinforcement learning for vision-based robotic manipulation”
Dmitry Kalashnikov et al · 2018
Earlier work this paper cites.
“Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection”
Sergey Levine et al · 2018
Earlier work this paper cites.
“Set transformer: A framework for attention-based permutation-invariant neural networks”
Juho Lee et al · 2019
Earlier work this paper cites.
“ACRONYM: A Large-Scale Grasp Dataset Based on Simulation”
Clemens Eppner, Arsalan Mousavian and Dieter Fox · 2020
Earlier work this paper cites.
“Language models are few-shot learners”
Tom Brown et al · 2020
Earlier work this paper cites.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”
Alexey Dosovitskiy et al · 2020
Earlier work this paper cites.
“Bridge data: Boosting generalization of robotic skills with cross-domain datasets”
Frederik Ebert et al · 2021
Earlier work this paper cites.
“Implicit Behavioral Cloning”
Peter. Florence et al · 2021
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models”
Edward Hu et al · 2021
Earlier work this paper cites.
Scott Reed et al · 2022
Earlier work this paper cites.
“Rt-1: Robotics transformer for real-world control at scale”
Anthony Brohan et al · 2022
Earlier work this paper cites.
“R3m: A universal visual representation for robot manipulation”
Suraj Nair et al · 2022
Earlier work this paper cites.
“Masked visual pre-training for motor control”
Tete Xiao, Ilija Radosavovic, Trevor Darrell and Jitendra Malik · 2022
Earlier work this paper cites.
“Vip: Towards universal visual reward and representation via value-implicit pre-training”
Yecheng Ma et al · 2022
Cited alongside, same era.
“Real-world robot learning with masked visual pre-training”
Ilija Radosavovic et al · 2022
Cited alongside, same era.
“Bc-z: Zero-shot task generalization with robotic imitation learning”
Eric Jang et al · 2022
Cited alongside, same era.
Scott Reed et al · 2022
Cited alongside, same era.
“Towards More Generalizable One-shot Visual Imitation Learning”, 2022
Zhao Mandi, Fangchen Liu, Kimin Lee and Pieter Abbeel · 2022
Cited alongside, same era.
“Scaling up and distilling down: Language-guided robot skill acquisition”
Huy Ha, Pete Florence and Shuran Song · 2023
Later among the works it cites.
“Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Later among the works it cites.
“Robot Learning with Sensorimotor Pre-training”
Ilija Radosavovic et al · 2023
Later among the works it cites.
“VIMA: General robot manipulation with multimodal prompts”
Yunfan Jiang et al · 2023
Later among the works it cites.
“Hyper-decision transformer for efficient online policy adaptation”
Mengdi Xu et al · 2023
Later among the works it cites.
“Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Demonstrate once, imitate immediately (dome): Learning visual servoing for one-shot imitation learning”
Eugene Valassakis, Georgios Papagiannis, Norman Di and Edward Johns · 2022
Cited alongside, same era.
“Prompting decision transformer for few-shot policy generalization”
Mengdi Xu et al · 2022
Cited alongside, same era.
“Diffusion Policy: Visuomotor Policy Learning via Action Diffusion”
Cheng Chi et al · 2023
Cited alongside, same era.
“Learning fine-grained bimanual manipulation with low-cost hardware”
Tony Zhao, Vikash Kumar, Sergey Levine and Chelsea Finn · 2023
Cited alongside, same era.
“Interactive language: Talking to robots in real time”
Corey Lynch et al · 2023
Cited alongside, same era.
“ViNT: A Foundation Model for Visual Navigation”
Dhruv Shah et al · 2023
Cited alongside, same era.
“RoboAgent: Towards Sample Efficient Robot Manipulation with Semantic Augmentations and Action Chunking”
Homanga Bharadhwaj et al · 2023
Cited alongside, same era.
Arjun Majumdar et al · 2023
Later among the works it cites.
“ImageBind-LLM: Multi-modality Instruction Tuning”, 2023
Jiaming Han et al · 2023
Later among the works it cites.
“Large Language Models as General Pattern Machines”, 2023
Suvir Mirchandani et al · 2023
Later among the works it cites.
“Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality”
Wei-Lin Chiang et al · 2023
Later among the works it cites.
“Improved Baselines with Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Later among the works it cites.
“OpenVLA: An Open-Source Vision-Language-Action Model”, 2024
Moo Kim et al · 2024
Closest in time.
“Octo: An Open-Source Generalist Robot Policy”
Octo Model Team et al · 2024
Closest in time.
“Humanoid Locomotion as Next Token Prediction”, 2024
Ilija Radosavovic et al · 2024
Closest in time.
“Open X-Embodiment: Robotic Learning Datasets and RT-X Models”, 2024
Embodiment Collaboration et al · 2024
Closest in time.
“Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics”
Norman Di and Edward Johns · 2024
Closest in time.
“Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers”
Vidhi Jain et al · 2024
Closest in time.
“DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset”, 2024
Alexander Khazatsky et al · 2024
Closest in time.
“SUGAR: Pre-training 3D Visual Representations for Robotics”
Shizhe Chen, Ricardo Garcia, Ivan Laptev and Cordelia Schmid · 2024
Closest in time.
“Rethinking Patch Dependence for Masked Autoencoders”
Letian Fu et al · 2024
Closest in time.
“A Touch, Vision, and Language Dataset for Multimodal Alignment”
Letian Fu et al · 2024
Closest in time.
“Instructblip: Towards general-purpose vision-language models with instruction tuning”
Wenliang Dai et al · 2024
Closest in time.