Fetching the paper…
Reading the bibliography…
We aim to enable humanoid robots to efficiently solve new manipulation tasks from a few video examples.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Evolutionary principles in self-referential learning”
Jurgen Schmidhuber · 1987
Earlier work this paper cites.
“Meta-neural networks that learn by learning”
Devang Naik and Richard Mammone · 1992
Earlier work this paper cites.
“Learning to learn using gradient descent”
Sepp Hochreiter, A Younger and Peter Conwell · 2001
Earlier work this paper cites.
“Representation and control of the task space in humans and humanoid robots”
Michael Mistry and Stefan Schaal · 2015
Earlier work this paper cites.
“Meta-learning with memory-augmented neural networks”
Adam Santoro et al · 2016
Earlier work this paper cites.
“Rl2: Fast reinforcement learning via slow reinforcement learning”
Yan Duan et al · 2016
Earlier work this paper cites.
“Model-agnostic meta-learning for fast adaptation of deep networks”
Chelsea Finn, Pieter Abbeel and Sergey Levine · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“One-shot imitation from observing humans via domain-adaptive meta-learning”
Tianhe Yu et al · 2018
Earlier work this paper cites.
“Task-embedded control networks for few-shot imitation learning”
Stephen James, Michael Bloesch and Andrew Davison · 2018
Earlier work this paper cites.
“Roboturk: A crowdsourcing platform for robotic skill learning through imitation”
Ajay Mandlekar et al · 2018
Earlier work this paper cites.
“Learning latent plans from play”
Corey Lynch et al · 2020
Earlier work this paper cites.
“robosuite: A Modular Simulation Framework and Benchmark for Robot Learning”
Yuke Zhu et al · 2020
Earlier work this paper cites.
“Metaicl: Learning to learn in context”
Sewon Min, Mike Lewis, Luke Zettlemoyer and Hannaneh Hajishirzi · 2021
Earlier work this paper cites.
“Rrl: Resnet as representation for reinforcement learning”
Rutav Shah and Vikash Kumar · 2021
Earlier work this paper cites.
“Concept2robot: Learning manipulation concepts from instructions and human demonstrations”
Lin Shao et al · 2021
Earlier work this paper cites.
“Flamingo: a visual language model for few-shot learning”
Jean-Baptiste Alayrac et al · 2022
Earlier work this paper cites.
“Prompting decision transformer for few-shot policy generalization”
Mengdi Xu et al · 2022
Earlier work this paper cites.
“In-context reinforcement learning with algorithm distillation”
Michael Laskin et al · 2022
Earlier work this paper cites.
“General-purpose in-context learning by meta-learning transformers”
Louis Kirsch, James Harrison, Jascha Sohl-Dickstein and Luke Metz · 2022
Cited alongside, same era.
“Masked autoencoders are scalable vision learners”
Kaiming He et al · 2022
Cited alongside, same era.
“R3m: A universal visual representation for robot manipulation”
Suraj Nair et al · 2022
Cited alongside, same era.
“Vip: Towards universal visual reward and representation via value-implicit pre-training”
Yecheng Ma et al · 2022
Cited alongside, same era.
“Amago: Scalable in-context reinforcement learning for adaptive agents”
Jake Grigsby, Linxi Fan and Yuke Zhu · 2023
“Benchmarking General-Purpose In-Context Learning”
Fan Wang, Chuan Lin, Yang Cao and Yu Kang · 2024
Later among the works it cites.
“DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset”, 2024
Alexander et al · 2024
Later among the works it cites.
“Okami: Teaching humanoid robots manipulation skills through single video imitation”
Jinhan Li et al · 2024
Later among the works it cites.
“Rethinking patch dependence for masked autoencoders”
Letian Fu et al · 2024
Later among the works it cites.
“Vid2robot: End-to-end video-conditioned policy learning with cross-attention transformers”
Vidhi Jain et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Generalization to new sequential decision making tasks with in-context learning”
Sharath Raparthy et al · 2023
Cited alongside, same era.
Open-Embodiment Collaboration · 2023
Cited alongside, same era.
“Mimicplay: Long-horizon imitation learning by watching human play”
Chen Wang et al · 2023
Cited alongside, same era.
“Zero-shot robot manipulation from passive human videos”
Homanga Bharadhwaj, Abhinav Gupta, Shubham Tulsiani and Vikash Kumar · 2023
Cited alongside, same era.
“Where are we in the search for an artificial visual cortex for embodied intelligence?”
Arjun Majumdar et al · 2023
Cited alongside, same era.
“Liv: Language-image representations and rewards for robotic control”
Yecheng Ma et al · 2023
Cited alongside, same era.
“Dinov2: Learning robust visual features without supervision”
Maxime Oquab et al · 2023
Cited alongside, same era.
“RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots”
Soroush Nasiriany et al · 2024
Later among the works it cites.
“Bridging the human to robot dexterity gap through object-oriented rewards”
Irmak Guzey et al · 2024
Later among the works it cites.
“Screwmimic: Bimanual imitation from human videos with screw space projection”
Arpit Bahety, Priyanka Mandikal, Ben Abbatematteo and Roberto Martín-Martín · 2024
Later among the works it cites.
“Egomimic: Scaling imitation learning via egocentric video”
Simar Kareer et al · 2024
Later among the works it cites.
“WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild”
Rolandos Potamias, Jinglei Zhang, Jiankang Deng and Stefanos Zafeiriou · 2024
Later among the works it cites.
“Reconstructing Hands in 3D with Transformers”
Georgios Pavlakos et al · 2024
Later among the works it cites.
“Vision-based manipulation from single human video with open-world object graphs”
Yifeng Zhu, Arisrei Lim, Peter Stone and Yuke Zhu · 2024
Later among the works it cites.
“RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots”
Soroush Nasiriany et al · 2024
Later among the works it cites.
“Self-distillation bridges distribution gap in language model fine-tuning”
Zhaorui Yang et al · 2024
Later among the works it cites.
“RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models”
Kaustubh Sridhar, Souradeep Dutta, Dinesh Jayaraman and Insup Lee · 2025
Closest in time.
“Humanoid Policy˜ Human Policy”
Ri-Zhao Qiu et al · 2025
Closest in time.
“HAND Me the Data: Fast Robot Adaptation via Hand Path Retrieval”
Matthew Hong et al · 2025
Closest in time.
“Phantom: Training Robots Without Robots Using Only Human Videos”
Marion Lepert, Jiaying Fang and Jeannette Bohg · 2025
Closest in time.
“LEGATO: Cross-Embodiment Imitation Using a Grasping Tool”
Mingyo Seo et al · 2025
Closest in time.