Fetching the paper…
Reading the bibliography…
Hierarchical policies that combine language and low-level control have been shown to perform impressively long-horizon robotic tasks, by leveraging either zero-shot high-level planners like pretrained language and vision-language models (LLMs/VLMs) or models trained on annotated robotic demonstrations.
Interactive task planning through natural language
Y.K. Hwang, P.C. Chen, and P.A. Watterberg · 1996
Earlier work this paper cites.
Corey Lynch and Pierre Sermanet · 2005
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
Corey Lynch and Pierre Sermanet · 2005
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, and action in route instructions
Matt MacMahon, Brian Stankiewicz, and Benjamin Kuipers · 2006
Earlier work this paper cites.
Toward understanding natural language directions
Thomas Kollar, Stefanie Tellex, Deb Roy, and Nicholas Roy · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J. Gordon, and J. Bagnell · 2010
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Stefanie Tellex, Thomas Kollar, Steven Dickerson, Matthew Walter, Ashis Banerjee, Seth Teller, and Nicholas Roy · 2011
Earlier work this paper cites.
Bruno Da Silva, George Konidaris, and Andrew Barto · 2012
Earlier work this paper cites.
Reinforcement learning to adjust parametrized motor primitives to new situations
Jens Kober, Andreas Wilhelm, Erhan Oztop, and Jan Peters · 2012
Earlier work this paper cites.
Multi-task policy search for robotics
Marc Peter Deisenroth, Peter Englert, Jan Peters, and Dieter Fox · 2014
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
Hongyuan Mei, Mohit Bansal, and Matthew Walter · 2016
Earlier work this paper cites.
Tell me dave: Context-sensitive grounding of natural language to manipulation instructions
Dipendra K Misra, Jaeyong Sung, Kevin Lee, and Ashutosh Saxena · 2016
Earlier work this paper cites.
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Earlier work this paper cites.
Real-time natural language corrections for assistive robotic manipulators
Alexander Broad, Jacob Arkin, Nathan Ratliff, Thomas Howard, and Brenna Argall · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville · 2017
Earlier work this paper cites.
Hierarchical and interpretable skill acquisition in multi-task reinforcement learning
Tianmin Shu, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Distral: Robust multitask reinforcement learning
Yee Teh, Victor Bapst, Wojciech M Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Multi-task learning for continuous control
Himani Arora, Rajath Kumar, Jason Krone, and Chong Li · 2018
Earlier work this paper cites.
Tadam: Task dependent adaptive metric for improved few-shot learning
Boris Oreshkin, Pau Rodríguez López, and Alexandre Lacoste · 2018
Earlier work this paper cites.
Learning to sequence multiple tasks with competing constraints
Anqing Duan, R. Camoriano, Diego Ferigo, Yanlong Huang, Daniele Calandriello, L. Rosasco, and D. Pucci · 2019
Earlier work this paper cites.
Language as an abstraction for hierarchical deep reinforcement learning
Yiding Jiang, Shixiang Shane Gu, Kevin P Murphy, and Chelsea Finn · 2019
Earlier work this paper cites.
Hg-dagger: Interactive imitation learning with human experts
Michael Kelly, Chelsea Sidrane, K. Driggs-Campbell, and Mykel J. Kochenderfer · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 2019
Cited alongside, same era.
Neural Semantic Parsing with Anonymization for Command Understanding in General-Purpose Service Robots , pages 337–350
Nick Walker, Yu-Tang Peng, and Maya Cakmak · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, S. Gelly, Jakob Uszkoreit, and N. Houlsby · 2020
Cited alongside, same era.
Robotic telekinesis: Learning a robotic hand imitator by watching humans on youtube
Aravind Sivakumar, Kenneth Shaw, and Deepak Pathak · 2022
Later among the works it cites.
Hierarchical reinforcement learning with natural language subgoals
Arun Ahuja, Kavya Kopparapu, Rob Fergus, and Ishita Dasgupta · 2023
Later among the works it cites.
Affordances from human videos as a versatile representation for robotics
Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, and Deepak Pathak · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al · 2023
Later among the works it cites.
Latte: Language trajectory transformer
Arthur Bucker, Luis Figueredo, Sami Haddadin, Ashish Kapoor, Shuang Ma, Sai Vemprala, and Rogerio Bonatti · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ajay Mandlekar, Danfei Xu, Roberto Martín-Martín, Yuke Zhu, Li Fei-Fei, and Silvio Savarese · 2020
Cited alongside, same era.
Language-conditioned imitation learning for robot manipulation tasks
Simon Stepputtis, Joseph Campbell, Mariano Phielipp, Stefan Lee, Chitta Baral, and Heni Ben Amor · 2020
Cited alongside, same era.
Deep imitation learning for bimanual robotic manipulation
Fan Xie, A. Chowdhury, M. Clara De Paolis Kaluza, Linfeng Zhao, Lawson L. S. Wong, and Rose Yu · 2020
Cited alongside, same era.
Integrated task and motion planning
Caelan Reed Garrett, Rohan Chitnis, Rachel M. Holladay, Beomjoon Kim, Tom Silver, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2021
Cited alongside, same era.
Lazydagger: Reducing context switching in interactive imitation learning
Ryan Hoque, Ashwin Balakrishna, Carl Putterman, Michael Luo, Daniel S Brown, Daniel Seita, Brijen Thananjeyan, Ellen Novoseller, and Ken Goldberg · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, A. Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Skill induction and planning with latent language
Pratyusha Sharma, Antonio Torralba, and Jacob Andreas · 2021
Cited alongside, same era.
Later among the works it cites.
Multimodal error correction with natural language and pointing gestures
Stefan Constantin, Fevziye Irem Eyiokur, Dogucan Yaman, Leonard Bärmann, and Alex Waibel · 2023
Later among the works it cites.
”no, to the right” - online language corrections for robotic manipulation via shared autonomy
Yuchen Cui, Siddharth Karamcheti, Raj Palleti, Nidhya Shivakumar, Percy Liang, and Dorsa Sadigh · 2023
Later among the works it cites.
Task and motion planning with large language models for object rearrangement
Yan Ding, Xiaohan Zhang, Chris Paxton, and Shiqi Zhang · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
Danny Driess, F. Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Q. Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, P. Sermanet, Daniel Duckworth, S. Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Peter R. Florence · 2023
Later among the works it cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, F. Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R. Florence, and Andy Zeng · 2023
Later among the works it cites.
Interactive robot learning from verbal correction
Huihan Liu, Alice Chen, Yuke Zhu, Adith Swaminathan, Andrey Kolobov, and Ching-An Cheng · 2023
Later among the works it cites.
Multi-stage cable routing through hierarchical imitation learning
Jianlan Luo, Charles Xu, Xinyang Geng, Gilbert Feng, Kuan Fang, L. Tan, S. Schaal, and S. Levine · 2023
Later among the works it cites.
Interactive language: Talking to robots in real time
Corey Lynch, Ayzaan Wahid, Jonathan Tompson, Tianli Ding, James Betker, Robert Baruch, Travis Armstrong, and Pete Florence · 2023
Later among the works it cites.
Learning reusable manipulation strategies
Jiayuan Mao, J. B. Tenenbaum, Tom’as Lozano-P’erez, and L. Kaelbling · 2023
Later among the works it cites.
Structured world models from human videos
Russell Mendonca, Shikhar Bahl, and Deepak Pathak · 2023
Later among the works it cites.
Chatgpt for robotics: Design principles and model abilities
Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor · 2023
Later among the works it cites.
Learning fine-grained bimanual manipulation with low-cost hardware
Tony Z. Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn · 2023
Later among the works it cites.
Rt-h: Action hierarchies using language
Suneel Belkhale, Tianli Ding, Ted Xiao, Pierre Sermanet, Quon Vuong, Jonathan Tompson, Yevgen Chebotar, Debidatta Dwibedi, and Dorsa Sadigh · 2024
Closest in time.
Learning to learn faster from human feedback with language model predictive control
Jacky Liang, Fei Xia, Wenhao Yu, Andy Zeng, Montserrat Gonzalez Arenas, Maria Attarian, Maria Bauza, Matthew Bennice, Alex Bewley, Adil Dostmohamed, Chuyuan Kelly Fu, Nimrod Gileadi, Marissa Giustina, Keerthana Gopalakrishnan, Leonard Hasenclever, Jan Humplik, Jasmine Hsu, Nikhil Joshi, Ben Jyenis, Chase Kew, Sean Kirmani, Tsang-Wei Edward Lee, Kuang-Huei Lee, Assaf Hurwitz Michaely, Joss Moore, Ken Oslund, Dushyant Rao, Allen Ren, Baruch Tabanpour, Quan Vuong, Ayzaan Wahid, Ted Xiao, Ying Xu, Vincent Zhuang, Peng Xu, Erik Frey, Ken Caluwaerts, Tingnan Zhang, Brian Ichter, Jonathan Tompson, Leila Takayama, Vincent Vanhoucke, Izhak Shafran, Maja Mataric, Dorsa Sadigh, Nicolas Heess, Kanishka Rao, Nik Stewart, Jie Tan, and Carolina Parada · 2024
Closest in time.
Aloha 2: An enhanced low-cost hardware for bimanual teleoperation, 2024
ALOHA 2 Team · 2024
Closest in time.
Mosaic: A modular system for assistive and interactive cooking
Huaxiaoyue Wang, Kushal Kedia, Juntao Ren, Rahma Abdullah, Atiksh Bhardwaj, Angela Chao, Kelly Y Chen, Nathaniel Chin, Prithwish Dan, Xinyi Fan, Gonzalo Gonzalez-Pumariega, Aditya Kompella, Maximus Adrian Pace, Yash Sharma, Xiangwan Sun, Neha Sunkara, and Sanjiban Choudhury · 2024
Closest in time.