Fetching the paper…
Reading the bibliography…
Language-conditioned robot manipulation is an emerging field aimed at enabling seamless communication and cooperation between humans and robotic agents by teaching robots to comprehend and execute instructions conveyed in natural language.
Advances in neural information processing systems 33: 1877–1901
Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al. (2020) Language models are few-shot learners · 1901
Earlier work this paper cites.
arXiv preprint arXiv:1907.11692
Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, Levy O, Lewis M, Zettlemoyer L and Stoyanov V (2019) Roberta: A robustly optimized bert pretraining approach · 1907
Earlier work this paper cites.
In: Proceedings of The 9th Conference on Robot Learning , Proceedings of Machine Learning Research , volume 305. PMLR, pp. 1898–1913
Li P, Wu Y, Xi Z, Li W, Huang Y, Zhang Z, Chen Y, Wang J, Zhu SC, Liu T and Huang S (2025c) Controlvla: Few-shot object-centric adaptation for pre-trained vision-language-action models · 1913
Earlier work this paper cites.
IEEE Robotics and Automation Letters 10(2): 1912–1919
Li P, Wu H, Huang Y, Cheang C, Wang L and Kong T (2025b) Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy · 1919
Earlier work this paper cites.
In: International Conference on Machine Learning . PMLR, pp. 1931–1942
Cho J, Lei J, Tan H and Bansal M (2021) Unifying vision-and-language tasks via text generation · 1942
Earlier work this paper cites.
Conference on Robot Learning : 1930–1942
Das N, Bechtle S, Davchev T, Jayaraman D, Rai A and Meier F (2021) Model-based inverse reinforcement learning from visual demonstrations · 1942
Earlier work this paper cites.
In: 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 1963–1969
Chen H, Tan H, Kuntz A, Bansal M and Alterovitz R (2020a) Enabling robots to understand incomplete natural language instructions using commonsense reasoning · 1969
Earlier work this paper cites.
Artificial intelligence 2(3-4): 189–208
Fikes RE and Nilsson NJ (1971) Strips: A new approach to the application of theorem proving to problem solving · 1971
Earlier work this paper cites.
Conference on Robot Learning : 1949–1974
Ke TW, Gkanatsios N and Fragkiadaki K (2025) 3d diffuser actor: Policy diffusion with 3d scene representations · 1974
Earlier work this paper cites.
Neural computation 3(1): 79–87
Jacobs RA, Jordan MI, Nowlan SJ and Hinton GE (1991) Adaptive mixtures of local experts · 1991
Earlier work this paper cites.
In: Machine Intelligence 15 . pp. 103–129
Bain M and Sammut C (1995) A framework for behavioural cloning · 1995
Earlier work this paper cites.
Cambridge university press
Clark HH (1996) Using language · 1996
Earlier work this paper cites.
In: Proceedings of International Conference on Robotics and Automation , volume 1. pp. 888–894 vol.1
Knoll A, Hildenbrandt B and Zhang J (1997) Instructing cooperating assembly robots through situated dialogues in natural language · 1997
Earlier work this paper cites.
MIT Press google schola 2: 678–686
Fellbaum C (1998) Wordnet: An electronic lexical database · 1998
Earlier work this paper cites.
Cambridge university press
Kintsch W (1998) Comprehension: A paradigm for cognition · 1998
Earlier work this paper cites.
IEEE Transactions on Neural Networks 9(5): 1054–1054
Sutton R and Barto A (1998) Reinforcement learning: An introduction · 1998
Earlier work this paper cites.
MIT press Cambridge
Sutton RS, Barto AG et al. (1998) Introduction to reinforcement learning , volume 135 · 1998
Earlier work this paper cites.
In: Icml , volume 99. Citeseer, pp. 278–287
Ng AY, Harada D and Russell S (1999) Policy invariance under reward transformations: Theory and application to reward shaping · 1999
Earlier work this paper cites.
Robotics and Autonomous Systems 29(1): 91–100
Zhang J, von Collani Y and Knoll A (1999) Interactive assembly by a two-arm robot agent · 1999
Earlier work this paper cites.
Journal of machine learning research 3(Feb): 1137–1155
Bengio Y, Ducharme R, Vincent P and Jauvin C (2003) A neural probabilistic language model · 2003
Earlier work this paper cites.
IEEE Transactions on Robotics and Automation 19(2): 223–237
Yamashita A, Arai T, Ota J and Asama H (2003) Motion planning of multiple mobile robots for cooperative manipulation and transportation · 2003
Earlier work this paper cites.
In: Proceedings of the twenty-first international conference on Machine learning . p. 1
Abbeel P and Ng AY (2004) Apprenticeship learning via inverse reinforcement learning · 2004
Earlier work this paper cites.
Harvard University Press
Goldin-Meadow S (2005) Hearing gesture: How our hands help us think · 2005
Earlier work this paper cites.
Colas C, Akakzia A, Oudeyer PY, Chetouani M and Sigaud O (2020) Language-conditioned goal generation: a new approach to language grounding for rl · 2006
Earlier work this paper cites.
In: AAAI Spring Symposium: Formalizing and Compiling Background Knowledge and Its Applications to Knowledge Representation and Question Answering . AAAI, pp. 44–49
Matuszek C, Cabral J, Witbrock M and DeOliveira J (2006) An introduction to the syntax and content of cyc · 2006
Earlier work this paper cites.
In: RO-MAN 2007-The 16th IEEE International Symposium on Robot and Human Interactive Communication . IEEE, pp. 702–707
Calinon S and Billard A (2007) Active teaching in robot programming by demonstration · 2007
Earlier work this paper cites.
Symbols and embodiment: Debates on meaning and cognition : 309–326
Louwerse MM and Jeuniaux P (2008) Language comprehension is both embodied and symbolic · 2008
Earlier work this paper cites.
In: Aaai , volume 8. Chicago, IL, USA, pp. 1433–1438
Ziebart BD, Maas AL, Bagnell JA, Dey AK et al. (2008) Maximum entropy inverse reinforcement learning · 2008
Earlier work this paper cites.
Robotics and Autonomous Systems 57(5): 469–483
Argall BD, Chernova S, Veloso M and Browning B (2009) A survey of robot learning from demonstration · 2009
Earlier work this paper cites.
The International Journal of Robotics Research 28(1): 104–126
Cambon S, Alami R and Gravot F (2009) A hybrid approach to intricate motion, manipulation and task planning · 2009
Earlier work this paper cites.
IEEE transactions on robotics 25(6): 1370–1381
Kress-Gazit H, Fainekos GE and Pappas GJ (2009) Temporal-logic-based reactive mission and motion planning · 2009
Earlier work this paper cites.
Cognitive processing 11: 133–142
Funke J (2010) Complex problem solving: A case for complex cognition? · 2010
Earlier work this paper cites.
Interspeech 2(3): 1045–1048
Mikolov T, Karafiát M, Burget L, Cernockỳ J and Khudanpur S (2010) Recurrent neural network based language model · 2010
Earlier work this paper cites.
In: 2010 ieee international conference on robotics and automation . IEEE, pp. 1486–1491
Tenorth M, Nyga D and Beetz M (2010) Understanding and executing instructions for everyday manipulation tasks from the world wide web · 2010
Earlier work this paper cites.
In: 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) . pp. 1999–2010
Yang J, Tan W, Jin C, Yao K, Liu B, Fu J, Song R, Wu G and Wang L (2025c) Transferring foundation models for generalizable robotic manipulation · 2010
Earlier work this paper cites.
arXiv preprint arXiv:2011.01975
Batra D, Chang AX, Chernova S, Davison AJ, Deng J, Koltun V, Levine S, Malik J, Mordatch I, Mottaghi R et al. (2020) Rearrangement: A challenge for embodied ai · 2011
Earlier work this paper cites.
Journal of machine learning research 12(ARTICLE): 2493–2537
Collobert R, Weston J, Bottou L, Karlen M, Kavukcuoglu K and Kuksa P (2011) Natural language processing (almost) from scratch · 2011
Earlier work this paper cites.
In: European workshop on reinforcement learning . Springer, pp. 273–284
Dimitrakakis C and Rothkopf CA (2011) Bayesian multitask inverse reinforcement learning · 2011
Earlier work this paper cites.
The international journal of robotics research 30(7): 846–894
Karaman S and Frazzoli E (2011) Sampling-based algorithms for optimal motion planning · 2011
Earlier work this paper cites.
In: Interspeech , volume 11. pp. 2877–2880
Kombrink S, Mikolov T, Karafiát M and Burget L (2011) Recurrent neural network based language modeling in meeting recognition · 2011
Earlier work this paper cites.
IEEE Robotics & Automation Magazine 18(2): 108–118
La Valle SM (2011) Motion planning · 2011
Earlier work this paper cites.
Robotics and Automation Magazine, IEEE 18: 69 – 82
Waibel M, Beetz M, Civera J, D’Andrea R, Elfring J, Gálvez-López D, Häussermann K, Janssen R, Montiel J, Perzylo A, Schießle B, Tenorth M, Zweigle O and Molengraft M (2011) Roboearth - a world wide web for robots · 2011
Earlier work this paper cites.
In: Proceedings of the 28th international conference on machine learning (ICML-11) . pp. 681–688
Welling M and Teh YW (2011) Bayesian learning via stochastic gradient langevin dynamics · 2011
Earlier work this paper cites.
In: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, pp. 251–256
Raman V, Finucane C and Kress-Gazit H (2012) Temporal logic robot mission planning for slow and fast actions · 2012
Earlier work this paper cites.
Foundations and Trends® in Machine Learning 4(4): 267–373
Sutton C, McCallum A et al. (2012) An introduction to conditional random fields · 2012
Earlier work this paper cites.
In: 2012 IEEE/RSJ international conference on intelligent robots and systems . IEEE, pp. 5026–5033
Todorov E, Erez T and Tassa Y (2012) Mujoco: A physics engine for model-based control · 2012
Earlier work this paper cites.
Advances in neural information processing systems 26
Bordes A, Usunier N, Garcia-Duran A, Weston J and Yakhnenko O (2013) Translating embeddings for modeling multi-relational data · 2013
Earlier work this paper cites.
In: 2013 IEEE/RSJ international conference on intelligent robots and systems . IEEE, pp. 1321–1326
Rohmer E, Singh SP and Freese M (2013) V-rep: A versatile and scalable robot simulation framework · 2013
Earlier work this paper cites.
In: Proceedings of the 2014 ACM/IEEE international conference on Human-robot interaction . pp. 33–40
Chai JY, She L, Fang R, Ottarson S, Littley C, Liu C and Hanson K (2014) Collaborative effort towards common ground in situated human-robot dialogue · 2014
Earlier work this paper cites.
In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) . pp. 1532–1543
Pennington J, Socher R and Manning CD (2014) Glove: Global vectors for word representation · 2014
Earlier work this paper cites.
In: The 23rd IEEE International Symposium on Robot and Human Interactive Communication . IEEE, pp. 868–873
She L, Cheng Y, Chai JY, Jia Y, Yang S and Xi N (2014) Teaching robots new actions through natural language instructions · 2014
Earlier work this paper cites.
Ai Magazine 36(3): 90–98
Vallati M, Chrpa L, Grześ M, McCluskey TL, Roberts M, Sanner S et al. (2015) The 2014 international planning competition: Progress and trends · 2014
Earlier work this paper cites.
Proceedings of the AAAI Conference on Artificial Intelligence 28(1)
Wang Z, Zhang J, Feng J and Chen Z (2014) Knowledge graph embedding by translating on hyperplanes · 2014
Earlier work this paper cites.
In: Proceedings of the IEEE international conference on computer vision . pp. 1440–1448
Girshick R (2015) Fast r-cnn · 2015
Earlier work this paper cites.
Proceedings of the AAAI Conference on Artificial Intelligence 29(1)
Lin Y, Liu Z, Sun M, Liu Y and Zhu X (2015) Learning entity and relation embeddings for knowledge graph completion · 2015
Earlier work this paper cites.
In: Robotics: Science and Systems
MacGlashan J, Babes-Vroman M, desJardins M, Littman ML, Muresan S, Squire S, Tellex S, Arumugam D and Yang L (2015) Grounding english commands to reward functions · 2015
Earlier work this paper cites.
In: International conference on machine learning . PMLR, pp. 49–58
Finn C, Levine S and Abbeel P (2016) Guided cost learning: Deep inverse optimal control via policy optimization · 2016
Earlier work this paper cites.
In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . pp. 1814–1824
Gao Q, Doering M, Yang S and Chai J (2016) Physical causality of action verbs in grounded language understanding · 2016
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition . pp. 770–778
He K, Zhang X, Ren S and Sun J (2016) Deep residual learning for image recognition · 2016
Earlier work this paper cites.
In: 2016 IEEE/RSJ international conference on intelligent robots and systems (IROS) . IEEE, pp. 151–156
Mees O, Eitel A and Burgard W (2016) Choosing smartly: Adaptive multimodal fusion for object detection in changing environments · 2016
Earlier work this paper cites.
The International Journal of Robotics Research 35(1-3): 281–300
Misra DK, Sung J, Lee K and Saxena A (2016) Tell me dave: Context-sensitive grounding of natural language to manipulation instructions · 2016
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition . pp. 779–788
Redmon J, Divvala S, Girshick R and Farhadi A (2016) You only look once: Unified, real-time object detection · 2016
Earlier work this paper cites.
In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . pp. 108–117
She L and Chai J (2016) Incremental acquisition of verb hypothesis space towards physical world interaction · 2016
Earlier work this paper cites.
International conference on machine learning : 166–175
Andreas J, Klein D and Levine S (2017) Modular multitask reinforcement learning with policy sketches · 2017
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems , volume 30. Curran Associates, Inc., pp. 5055–5065
Andrychowicz M, Wolski F, Ray A, Schneider J, Fong R, Welinder P, McGrew B, Tobin J, Pieter Abbeel O and Zaremba W (2017) Hindsight experience replay · 2017
Earlier work this paper cites.
arXiv preprint arXiv:1706.06551
Hermann KM, Hill F, Green S, Wang F, Faulkner R, Soyer H, Szepesvari D, Czarnecki WM, Jaderberg M, Teplyashin D et al. (2017) Grounded language learning in a simulated 3d world · 2017
Earlier work this paper cites.
arXiv preprint arXiv:1704.05539
Kaplan R, Sauer C and Sosa A (2017) Beating atari with natural language guided reinforcement learning · 2017
Earlier work this paper cites.
Proceedings of the national academy of sciences 114(13): 3521–3526
Kirkpatrick J, Pascanu R, Rabinowitz N, Veness J, Desjardins G, Rusu AA, Milan K, Quan J, Ramalho T, Grabska-Barwinska A et al. (2017) Overcoming catastrophic forgetting in neural networks · 2017
Earlier work this paper cites.
arXiv preprint arXiv:1712.05474
Kolve E, Mottaghi R, Han W, VanderBilt E, Weihs L, Herrasti A, Deitke M, Ehsani K, Gordon D, Zhu Y et al. (2017) Ai2-thor: An interactive 3d environment for visual ai · 2017
Earlier work this paper cites.
In: 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 3175–3182
Mees O, Abdo N, Mazuran M and Burgard W (2017) Metric learning for generalizing spatial relations to new objects · 2017
Earlier work this paper cites.
In: Proceedings of the 2017 conference on empirical methods in natural language processing . pp. 1004–1015
Misra D, Langford J and Artzi Y (2017) Mapping instructions and visual observations to actions with reinforcement learning · 2017
Earlier work this paper cites.
In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . pp. 1634–1644
She L and Chai J (2017) Interactive learning of grounded verb semantics towards human-robot communication · 2017
Earlier work this paper cites.
the Thirty-First AAAI Conference on Artificial Intelligence : 4444–4451
Speer R, Chin J and Havasi C (2017) Conceptnet 5.5: An open multilingual graph of general knowledge · 2017
Earlier work this paper cites.
Advances in neural information processing systems 30
Srivastava A, Valkov L, Russell C, Gutmann MU and Sutton C (2017) Veegan: Reducing mode collapse in gans using implicit variational learning · 2017
Earlier work this paper cites.
Advances in neural information processing systems 30
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł and Polosukhin I (2017) Attention is all you need · 2017
Earlier work this paper cites.
International Conference on Learning Representations
Bahdanau D, Hill F, Leike J, Hughes E, Hosseini A, Kohli P and Grefenstette E (2018) Learning to understand goal specifications by modelling reward · 2018
Earlier work this paper cites.
IEEE transactions on pattern analysis and machine intelligence 41(2): 423–443
Baltrušaitis T, Ahuja C and Morency LP (2018) Multimodal machine learning: A survey and taxonomy · 2018
Earlier work this paper cites.
In: 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 512–519
Beetz M, Beßler D, Haidu A, Pomarlan M, Bozcuoğlu AK and Bartels G (2018) Know rob 2.0—a 2nd generation knowledge processing framework for cognition-enabled robotic agents · 2018
Earlier work this paper cites.
In: Lang J (ed.) Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden . ijcai.org, pp. 2–9
Chai J, Gao Q, She L, Yang S, Saba-Sadiya S and Xu G (2018) Language to action: Towards interactive task learning with physical agents · 2018
Earlier work this paper cites.
In: Proceedings of the AAAI Conference on Artificial Intelligence , volume 32. pp. 2819–2826
Chaplot DS, Sathyendra KM, Pasumarthi RK, Rajagopal D and Salakhutdinov R (2018) Gated-attention architectures for task-oriented language grounding · 2018
Earlier work this paper cites.
arXiv preprint arXiv:1812.00568
Ebert F, Finn C, Dasari S, Xie A, Lee A and Levine S (2018) Visual foresight: Model-based deep reinforcement learning for vision-based robotic control · 2018
Earlier work this paper cites.
In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1) . pp. 934–945
Gao Q, Yang S, Chai J and Vanderwende L (2018) What action causes this? towards naive physical action-effect prediction · 2018
Earlier work this paper cites.
In: 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 3774–3781
Hatori J, Kikuchi Y, Kobayashi S, Takahashi K, Tsuboi Y, Unno Y, Ko W and Tan J (2018) Interactively picking real-world objects with unconstrained spoken language instructions · 2018
Earlier work this paper cites.
Transactions of the Association for Computational Linguistics 6: 49–61
Janner M, Narasimhan K and Barzilay R (2018) Representation learning for grounded spatial reasoning · 2018
Earlier work this paper cites.
Knowledge-Based Systems 140: 15–26
Liu R and Zhang X (2018) Generating machine-executable plans from end-user’s natural-language instructions · 2018
Earlier work this paper cites.
In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition . pp. 7765–7773
Mallya A and Lazebnik S (2018) Packnet: Adding multiple tasks to a single network by iterative pruning · 2018
Earlier work this paper cites.
In: Conference on Robot Learning . PMLR, pp. 879–893
Mandlekar A, Zhu Y, Garg A, Booher J, Spero M, Tung A, Gao J, Emmons J, Gupta A, Orbay E et al. (2018) Roboturk: A crowdsourcing platform for robotic skill learning through imitation · 2018
Earlier work this paper cites.
In: 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 527–534
Munawar A, De Magistris G, Pham TH, Kimura D, Tatsubori M, Moriyama T, Tachibana R and Booch G (2018) Maestrob: A robotics framework for integrated orchestration of low-level control and high-level reasoning · 2018
Earlier work this paper cites.
In: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) . pp. 2227–2237
Peters ME, Neumann M, Iyyer M, Gardner M, Clark C, Lee K and Zettlemoyer L (2018) Deep contextualized word representations · 2018
Earlier work this paper cites.
In: 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, pp. 3758–3765
Rahmatizadeh R, Abolghasemi P, Bölöni L and Levine S (2018) Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration · 2018
Earlier work this paper cites.
In: International conference on machine learning . PMLR, pp. 4344–4353
Riedmiller M, Hafner R, Lampe T, Neunert M, Degrave J, Wiele T, Mnih V, Heess N and Springenberg JT (2018) Learning by playing solving sparse reward tasks from scratch · 2018
Earlier work this paper cites.
The International Journal of Robotics Research 37(7): 688–716
Sanchez J, Corrales JA, Bouzgarrou BC and Mezouar Y (2018) Robotic manipulation and sensing of deformable objects in domestic and industrial applications: a survey · 2018
Earlier work this paper cites.
In: European semantic web conference . Springer, pp. 593–607
Schlichtkrull M, Kipf TN, Bloem P, Van Den Berg R, Titov I and Welling M (2018) Modeling relational data with graph convolutional networks · 2018
Earlier work this paper cites.
In: Proceedings of the 27th International Joint Conference on Artificial Intelligence . pp. 4950–4957
Torabi F, Warnell G and Stone P (2018) Behavioral cloning from observation · 2018
Earlier work this paper cites.
IEEE Transactions on Automation Science and Engineering 16(2): 640–653
Wang W, Li R, Chen Y, Diekel ZM and Jia Y (2018) Facilitating human–robot collaborative tasks by teaching-learning-collaboration from human demonstrations · 2018
Earlier work this paper cites.
In: 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 5628–5635
Zhang T, McCarthy Z, Jow O, Lee D, Chen X, Goldberg K and Abbeel P (2018) Deep imitation learning for complex manipulation tasks from virtual reality teleoperation · 2018
Earlier work this paper cites.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . pp. 4254–4262
Abolghasemi P, Mazaheri A, Shah M and Boloni L (2019) Pay attention!-robustifying a deep visuomotor policy through task-focused visual attention · 2019
Earlier work this paper cites.
Science 364(6446): eaat8414
Billard A and Kragic D (2019) Trends and challenges in robot manipulation · 2019
Earlier work this paper cites.
In: 2019 IEEE International Conference on Systems, Man and Cybernetics (SMC) . IEEE, pp. 486–492
Ceola F, Tosello E, Tagliapietra L, Nicola G and Ghidoni S (2019) Robot task planning via deep reinforcement learning: a tabletop object sorting application · 2019
Earlier work this paper cites.
7th International Conference on Learning Representations
Chaudhry A, Marc’Aurelio R, Rohrbach M and Elhoseiny M (2019) Efficient lifelong learning with a-gem · 2019
Earlier work this paper cites.
In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) . pp. 4171–4186
Devlin J, Chang MW, Lee K and Toutanova K (2019) Bert: Pre-training of deep bidirectional transformers for language understanding · 2019
Earlier work this paper cites.
Sensors 19(5): 1166
Diab M, Akbari A, Ud Din M and Rosell J (2019) Pmk—a knowledge processing framework for autonomous robotics perception and manipulation · 2019
Earlier work this paper cites.
Advances in neural information processing systems 32
Ding Y, Florensa C, Abbeel P and Phielipp M (2019) Goal-conditioned imitation learning · 2019
Earlier work this paper cites.
International Conference on Learning Representations
Fu J, Korattikara A, Levine S and Guadarrama S (2019) From language to goals: Inverse reinforcement learning for vision-based instruction following · 2019
Earlier work this paper cites.
In: Proceedings of the 28th International Joint Conference on Artificial Intelligence . pp. 2385–2391
Goyal P, Niekum S and Mooney RJ (2019) Using natural language for reward shaping in reinforcement learning · 2019
Earlier work this paper cites.
Advances in neural information processing systems 32
Lu J, Batra D, Parikh D and Lee S (2019) Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks · 2019
Earlier work this paper cites.
In: 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 6083–6089
Mees O, Tatarchenko M, Brox T and Burgard W (2019) Self-supervised 3d shape and viewpoint estimation from single images for robotics · 2019
Earlier work this paper cites.
In: Proceedings of the 57th annual meeting of the association for computational linguistics . pp. 4710–4723
Nathani D, Chauhan J, Sharma C and Kaul M (2019) Learning attention-based embeddings for relation prediction in knowledge graphs · 2019
Earlier work this paper cites.
In: 2019 11th International Conference on Knowledge and Systems Engineering (KSE) . IEEE, pp. 1–7
Nguyen TL, Nguyen DV and Le TH (2019) Reinforcement learning based navigation with semantic knowledge of indoor environments · 2019
Earlier work this paper cites.
In: 2019 International Conference on Robotics and Automation (ICRA) . IEEE, pp. 6964–6971
Park JS, Jia B, Bansal M and Manocha D (2019) Efficient generation of motion plans from attribute-based natural language instructions using dynamic constraint mapping · 2019
Earlier work this paper cites.
In: 2019 International Conference on Robotics and Automation (ICRA) . pp. 6926–6933
Patki S, Daniele AF, Walter MR and Howard TM (2019) Inferring compact representations for efficient natural language understanding of robot instructions · 2019
Earlier work this paper cites.
In: 2019 International Conference on Robotics and Automation (ICRA) . IEEE, pp. 6942–6948
Paxton C, Bisk Y, Thomason J, Byravan A and Foxl D (2019) Prospection: Interpretable plans from language by predicting the future · 2019
Earlier work this paper cites.
OpenAI blog 1(8): 9
Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I et al. (2019) Language models are unsupervised multitask learners · 2019
Earlier work this paper cites.
Journal of Ambient Intelligence and Smart Environments 11(3): 261–275
Rasch R, Sprute D, Pörtner A, Battermann S and König M (2019) Tidy up my room: Multi-agent cooperation for service tasks in smart environments · 2019
Earlier work this paper cites.
In: Proceedings of Thirty-third Conference on Neural Information Processing Systems (NIPS2019)
Sanh V, Debut L, Chaumond J and Wolf T (2019) Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter · 2019
Earlier work this paper cites.
In: 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 5284–5290
Sriram N, Maniar T, Kalyanasundaram J, Gandhi V, Bhowmick B and Krishna KM (2019) Talk to the vehicle: Language conditioned autonomous navigation of self driving cars · 2019
Earlier work this paper cites.
The International Journal of Robotics Research 38(10-11): 1179–1207
Tamosiunaite M, Aein MJ, Braun JM et al. (2019) Cut & recombine: reuse of robot action components based on simple language instructions · 2019
Earlier work this paper cites.
In: 2019 International Conference on Robotics and Automation (ICRA) . IEEE, pp. 6934–6941
Thomason J, Padmakumar A, Sinapov J, Walker N, Jiang Y, Yedidsion H, Hart J, Stone P and Mooney RJ (2019) Improving grounded natural language understanding through human-robot dialog · 2019
Earlier work this paper cites.
IEEE Robotics and Automation Letters 4(4): 3593–3600
Wu Y, Balatti P, Lorenzini M, Zhao F, Kim W and Ajoudani A (2019) A teleoperation interface for loco-manipulation control of mobile collaborative robotic assistant · 2019
Earlier work this paper cites.
Robotics: Science and Systems
Xie A, Ebert F, Levine S and Finn C (2019) Improvisation through physical understanding: Using novel objects as tools with visual foresight · 2019
Earlier work this paper cites.
Advances in neural information processing systems 32: 1319–1329
Xu D, Martín-Martín R, Huang DA, Zhu Y, Savarese S and Fei-Fei LF (2019) Regression planning networks · 2019
Earlier work this paper cites.
The International Journal of Robotics Research 39(10-11): 1279–1304
Arkin J, Park D, Roy S, Walter MR, Roy N, Howard TM and Paul R (2020) Multimodal estimation and communication of latent semantic knowledge for robust execution of robot instructions · 2020
Earlier work this paper cites.
In: Proceedings of the international conference on automated planning and scheduling , volume 30. pp. 440–448
Garrett CR, Lozano-Pérez T and Kaelbling LP (2020) Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning · 2020
Earlier work this paper cites.
Advances in neural information processing systems 33: 6840–6851
Ho J, Jain A and Abbeel P (2020) Denoising diffusion probabilistic models · 2020
Earlier work this paper cites.
IEEE Robotics and Automation Letters 5(2): 3019–3026
James S, Ma Z, Arrojo DR and Davison AJ (2020) Rlbench: The robot learning benchmark & learning environment · 2020
Earlier work this paper cites.
In: 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 4630–4636
Jangir R, Alenya G and Torras C (2020) Dynamic cloth manipulation with deep reinforcement learning · 2020
Earlier work this paper cites.
In: Conference on robot learning . PMLR, pp. 1113–1132
Lynch C, Khansari M, Xiao T, Kumar V, Tompson J, Levine S and Sermanet P (2020) Learning latent plans from play · 2020
Earlier work this paper cites.
In: 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 94–100
Mees O, Emek A, Vertens J and Burgard W (2020) Learning object placements for relational instructions by hallucinating scene representations · 2020
Earlier work this paper cites.
In: Conference on Robot Learning . PMLR, pp. 530–539
Nair A, Bahl S, Khazatsky A, Pong V, Berseth G and Levine S (2020) Contextual imagined goals for self-supervised robotic learning · 2020
Earlier work this paper cites.
In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 5319–5326
Nematollahi I, Mees O, Hermann L and Burgard W (2020) Hindsight for foresight: Unsupervised structured dynamics models from physical interaction · 2020
Earlier work this paper cites.
In: International conference on machine learning . PMLR, pp. 7487–7498
Parisotto E, Song F, Rae J, Pascanu R, Gulcehre C, Jayakumar S, Jaderberg M, Kaufman RL, Clark A, Noury S et al. (2020) Stabilizing transformers for reinforcement learning · 2020
Earlier work this paper cites.
In: Conference on Robot Learning . PMLR, pp. 540–551
Roh J, Paxton C, Pronobis A, Farhadi A and Fox D (2020) Conditional driving from natural language instructions · 2020
Earlier work this paper cites.
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . pp. 10740–10749
Shridhar M, Thomason J, Gordon D, Bisk Y, Han W, Mottaghi R, Zettlemoyer L and Fox D (2020) Alfred: A benchmark for interpreting grounded instructions for everyday tasks · 2020
Earlier work this paper cites.
Advances in Neural Information Processing Systems 33: 13139–13150
Stepputtis S, Campbell J, Phielipp M, Lee S, Baral C and Ben Amor H (2020) Language-conditioned imitation learning for robot manipulation tasks · 2020
Earlier work this paper cites.
Annual Review of Control, Robotics, and Autonomous Systems 3(1): 25–55
Tellex S, Gopalan N, Kress-Gazit H and Matuszek C (2020) Robots that use language · 2020
Earlier work this paper cites.
Conference on robot learning : 1094–1100
Yu T, Quillen D, He Z, Julian R, Hausman K, Finn C and Levine S (2020) Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning · 2020
Earlier work this paper cites.
Artificial Intelligence 297: 103500
Arora S and Doshi P (2021) A survey of inverse reinforcement learning: Challenges, methods and progress · 2021
Cited alongside, same era.
In: Neuro-Symbolic Artificial Intelligence: The State of the Art , volume 342. IOS Press, pp. 1–51
Besold TR, d’Avila Garcez AS, Bader S, Bowman H, Domingos PM, Hitzler P, Kühnberger K, Lamb LC, Lima PMV, de Penning L, Pinkas G, Poon H and Zaverucha G (2021) Neural-symbolic learning and reasoning: A survey and interpretation · 2021
Cited alongside, same era.
arXiv preprint arXiv:2107.03374
Chen M, Tworek J, Jun H, Yuan Q, Pinto HPdO, Kaplan J, Edwards H, Burda Y, Joseph N, Brockman G et al. (2021) Evaluating large language models trained on code · 2021
Cited alongside, same era.
arXiv preprint arXiv:2110.14168
Cobbe K, Kosaraju V, Bavarian M, Chen M, Jun H, Kaiser L, Plappert M, Tworek J, Hilton J, Nakano R et al. (2021) Training verifiers to solve math word problems · 2021
Cited alongside, same era.
http://pybullet.org
Coumans E and Bai Y (2016–2021) Pybullet, a python module for physics simulation for games, robotics and machine learning · 2021
Proceedings of Robotics: Science and Systems (RSS)
Di Palo N and Johns E (2024) Keypoint action tokens enable in-context imitation learning in robotics · 2024
Closest in time.
arXiv preprint arXiv:2410.00371
Duan J, Pumacay W, Kumar N, Wang YR, Tian S, Yuan W, Krishna R, Fox D, Mandlekar A and Guo Y (2024) Aha: A vision-language-model for detecting and reasoning over failures in robotic manipulation · 2024
Closest in time.
In: 2024 IEEE International Conference on Systems, Man, and Cybernetics (SMC) . IEEE, pp. 2864–2869
Gbagbe KF, Cabrera MA, Alabbas A, Alyunes O, Lykov A and Tsetserukou D (2024) Bi-vla: Vision-language-action model-based system for bimanual robotic dexterous manipulations · 2024
Closest in time.
arXiv preprint arXiv:2402.11367
Glazer N, Navon A, Shamsian A and Fetaya E (2024) Multi task inverse reinforcement learning for common sense reward · 2024
Closest in time.
In: International Conference on Intelligent Robotics and Applications . Springer, pp. 3–17
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
International Conference on Learning Representations
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, Uszkoreit J and Houlsby N (2021) An image is worth 16x16 words: Transformers for image recognition at scale · 2021
Cited alongside, same era.
Reinforcement learning algorithms: Analysis and Applications : 25–33
Eschmann J (2021) Reward function design in reinforcement learning · 2021
Cited alongside, same era.
International Journal of Computer Vision 129(6): 1789–1819
Gou J, Yu B, Maybank SJ and Tao D (2021) Knowledge distillation: A survey · 2021
Cited alongside, same era.
In: Conference on Robot Learning . PMLR, pp. 485–497
Goyal P, Niekum S and Mooney R (2021) Pixl2r: Guiding reinforcement learning using natural language by mapping pixels to rewards · 2021
Cited alongside, same era.
The International Journal of Robotics Research 40(4-5): 698–721
Ibarz J, Tan J, Finn C, Kalakrishnan M, Pastor P and Levine S (2021) How to train your robot with deep reinforcement learning: lessons we have learned · 2021
Cited alongside, same era.
In: Proceedings of the IEEE/CVF international conference on computer vision . pp. 1780–1790
Kamath A, Singh M, LeCun Y, Synnaeve G, Misra I and Carion N (2021) Mdetr-modulated detection for end-to-end multi-modal understanding · 2021
Cited alongside, same era.
Advances in neural information processing systems 34: 9694–9705
Li J, Selvaraju R, Gotmare A, Joty S, Xiong C and Hoi SCH (2021) Align before fuse: Vision and language representation learning with momentum distillation · 2021
Cited alongside, same era.
Guo Q, Liu X, Hui J, Liu Z and Huang P (2024) Utilizing large language models for robot skill reward shaping in reinforcement learning · 2024
Closest in time.
In: 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 7315–7322
Habekost JG, Gäde C, Allgeuer P and Wermter S (2024) Inverse kinematics for neuro-robotic grasping with humanoid embodied agents · 2024
Closest in time.
arXiv preprint arXiv:2410.21276
Hurst A, Lerer A, Goucher AP, Perelman A, Ramesh A, Clark A, Ostrow A, Welihinda A, Hayes A, Radford A et al. (2024) Gpt-4o system card · 2024
Closest in time.
In: Proceedings of the International Conference on Automated Planning and Scheduling , volume 34. pp. 301–309
Ju Z, Yang C, Sun F, Wang H and Qiao Y (2024) Rethinking mutual information for language conditioned skill discovery on imitation learning · 2024
Closest in time.
Kang G, Kim J, Shim K, Lee JK and Zhang B (2024) CLIP-RT: learning language-conditioned robotic policies from natural language supervision · 2024
Closest in time.
In: 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 4796–4803
Kapelyukh I, Ren Y, Alzugaray I and Johns E (2024) Dream2real: Zero-shot 3d object rearrangement with vision-language models · 2024
Closest in time.
In: Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction . pp. 361–370
Karli UB, Chen JT, Antony VN and Huang CM (2024) Alchemist: Llm-aided end-user development of robot applications · 2024
Closest in time.
Robotics: Science and Systems
Khazatsky A, Pertsch K, Nair S, Balakrishna A, Dasari S, Karamcheti S, Nasiriany S, Srirama MK, Chen LY, Ellis K et al. (2024) Droid: A large-scale in-the-wild robot manipulation dataset · 2024
Closest in time.
In: 2024 The Twelfth International Conference on Learning Representations
Ko PC, Mao J, Du Y, Sun SH and Tenenbaum JB (2024) Learning to act from actionless videos through dense correspondences · 2024
Closest in time.
IEEE Robotics and Automation Letters 9(7)
Kwon T, Di Palo N and johns E (2024) Language models as zero-shot trajectory generators · 2024
Closest in time.
arXiv e-prints : arXiv–2405
Liu J, Li C, Wang G, Lee L, Zhou K, Chen S, Xiong C, Ge J, Zhang R and Zhang S (2024) Self-corrected multimodal large language model for end-to-end robot manipulation · 2024
Closest in time.
arXiv preprint arXiv:2406.01586
Lu G, Gao Z, Chen T, Dai W, Wang Z and Tang Y (2024) Manicm: Real-time 3d diffusion policy via consistency model for robotic manipulation · 2024
Closest in time.
In: The Twelfth International Conference on Learning Representations
Ma YJ, Liang W, Wang G, Huang DA, Bastani O, Jayaraman D, Zhu Y, Fan L and Anandkumar A (2024) Eureka: Human-level reward design via coding large language models · 2024
Closest in time.
Actuators 13(1)
Miao R, Jia Q, Sun F, Chen G and Huang H (2024) Hierarchical understanding in robotic manipulation: A knowledge-based framework · 2024
Closest in time.
In: 2024 China Automation Congress (CAC) . IEEE, pp. 4310–4315
Ming C, Lin J, Fong P, Wang H, Duan X and He J (2024) Hicrisp: An llm-based hierarchical closed-loop robotic intelligent self-correction planner · 2024
Closest in time.
In: International Conference on Machine Learning . PMLR, pp. 36434–36454
Mu Y, Chen J, Zhang QL, Chen S, Yu Q, Ge C, Chen R, Liang Z, Hu M, Tao C et al. (2024) Robocodex: Multimodal code generation for robotic behavior synthesis · 2024
Closest in time.
In: 2024 Proceedings of the 41st International Conference on Machine Learning . pp. 37321–37341
Nasiriany S, Xia F, Yu W, Xiao T, Liang J, Dasgupta I, Xie A, Driess D, Wahid A, Xu Z et al. (2024) Pivot: iterative visual prompting elicits actionable knowledge for vlms · 2024
Closest in time.
In: Proceedings of Robotics: Science and Systems . Delft, Netherlands
Octo Model Team, Ghosh D, Walke H, Pertsch K, Black K, Mees O et al. (2024) Octo: An open-source generalist robot policy · 2024
Closest in time.
In: ICRA . IEEE, pp. 6892–6903
O’Neill A et al. (2024) Open x-embodiment: Robotic learning datasets and RT-X models : Open x-embodiment collaboration · 2024
Closest in time.
Transactions on Machine Learning Research
Oquab M, Darcet T, Moutakanni T, Vo HV, Szafraniec M et al. (2024) DINOv2: Learning robust visual features without supervision · 2024
Closest in time.
In: Robotics: Science and Systems
Prasad A, Lin K, Wu J, Zhou L and Bohg J (2024) Consistency policy: Accelerated visuomotor policies via consistency distillation · 2024
Closest in time.
In: Robotics: Science and Systems
Reuss M, Yağmurlu ÖE, Wenzel F and Lioutikov R (2024) Multimodal diffusion transformer: Learning versatile behavior from multimodal goals · 2024
Closest in time.
In: 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 2296–2302
Schaldenbrand P, Parmar G, Zhu JY, McCann J and Oh J (2024) Cofrida: Self-supervised fine-tuning for human-robot co-painting · 2024
Closest in time.
IEEE Robotics and Automation Letters
Scheikl PM, Schreiber N, Haas C, Freymuth N, Neumann G, Lioutikov R and Mathis-Ullrich F (2024) Movement primitive diffusion: Learning gentle robotic manipulation of deformable objects · 2024
Closest in time.
Robotics: Science and Systems
Shi LX, Hu Z, Zhao TZ, Sharma A, Pertsch K, Luo J, Levine S and Finn C (2024) Yell at your robot: Improving on-the-fly from language corrections · 2024
Closest in time.
SN Computer Science 5(1): 189
Silvera-Tawil D (2024) Robotics in healthcare: a survey · 2024
Closest in time.
International Conference on Multimedia Modeling : 340–353
Tan W, Liu B, Zhang J, Song R and Fu J (2024) Rold: Robot latent diffusion for multi-task policy modeling · 2024
Closest in time.
In: Intelligent Robotics: 5th China Annual Conference, CIRAC 2024, Dalian, China, September 20–22, 2024, Proceedings , volume 2355. Springer Nature, p. 167
Tang J, Cheng B and Wen M (2025) Hierarchical language-conditioned robot learning with vision-language models · 2024
Closest in time.
In: 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . pp. 11478–11484
Tarakli I, Vinanzi S and Nuovo AD (2024) Interactive reinforcement learning from natural language feedback · 2024
Closest in time.
arXiv preprint arXiv:2408.04380
Urain J, Mandlekar A, Du Y, Shafiullah M, Xu D, Fragkiadaki K, Chalvatzaki G and Peters J (2024) Deep generative models in robotics: A survey on learning from multimodal demonstrations · 2024
Closest in time.
In: Robotics: Science and Systems
Vosylius V, Seo Y, Uruç J and James S (2024) Render and diffuse: Aligning image and action spaces for diffusion-based behaviour cloning · 2024
Closest in time.
IEEE Robotics and Automation Letters 9(11): 10567–10574
Wake N, Kanehira A, Sasabuchi K, Takamatsu J and Ikeuchi K (2024) Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration · 2024
Closest in time.
In: 2024 International Conference on Machine Learning . PMLR, pp. 51936–51983
Wang Y, Xian Z, Chen F, Wang TH, Wang Y, Fragkiadaki K, Erickson Z, Held D and Gan C (2024d) Robogen: Towards unleashing infinite data for automated robot learning via generative simulation · 2024
Closest in time.
In: The Twelfth International Conference on Learning Representations
Xie T, Zhao S, Wu CH, Liu Y, Luo Q, Zhong V, Yang Y and Yu T (2024) Text2reward: Reward shaping with language models for reinforcement learning · 2024
Closest in time.
In: 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 10905–10912
Xu Z, Gao C, Liu Z, Yang G, Tie C et al. (2024) Manifoundation model for general-purpose robotic manipulation of contact synthesis with arbitrary objects and robots · 2024
Closest in time.
arXiv preprint arXiv:2403.04115
Yan G, Wu YH and Wang X (2024) Dnact: Diffusion guided multi-task 3d policy learning · 2024
Closest in time.
In: 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 7694–7701
Yang J, Chen X, Qian S, Madaan N, Iyengar M, Fouhey DF and Chai J (2024a) Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent · 2024
Closest in time.
In: 6th International Conference on Robotics, Intelligent Control and Artificial Intelligence (RICAI) . pp. 16–22
Ye F, Lin Y, Lai Y and Huang Q (2024) Task-oriented language grounding for robot via learning object mask · 2024
Closest in time.
Robotics: Science and Systems
Ze Y, Zhang G, Zhang K, Hu C, Wang M and Xu H (2024) 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations · 2024
Closest in time.
In: International Conference on Machine Learning . PMLR, pp. 58366–58386
Zeng Y, Mu Y and Shao L (2024) Learning reward for robot skills using large language models via self-alignment · 2024
Closest in time.
In: 2024 International Conference on Machine Learning . PMLR, pp. 61229–61245
Zhen H, Qiu X, Chen P, Yang J, Yan X, Du Y, Hong Y and Gan C (2024) 3d-vla: A 3d vision-language-action generative world model · 2024
Closest in time.
In: 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 16655–16661
Zhou H, Lin Y, Yan L, Zhu J and Min H (2024b) Llm-bt: Performing robotic adaptive tasks based on large language models and behavior trees · 2024
Closest in time.
In: 12th International Conference on Learning Representations, ICLR 2024
Zhou Y, Cui C, Yoon J, Zhang L, Deng Z, Finn C, Bansal M and Yao H (2024c) Analyzing and mitigating object hallucination in large vision-language models · 2024
Closest in time.
arXiv preprint arXiv:2510.03342
Abdolmaleki A, Abeyruwan S, Ainslie J, Alayrac JB, Arenas MG, Balakrishna A, Batchelor N, Bewley A, Bingham J, Bloesch M et al. (2025) Gemini robotics 1.5: Pushing the frontier of generalist robots with advanced embodied reasoning, thinking, and motion transfer · 2025
Closest in time.
In: 2025 Conference on Robot Learning . PMLR, pp. 689–723
Agia C, Sinha R, Yang J, Cao Z, Antonova R, Pavone M and Bohg J (2025) Unpacking failure modes of generative policies: Runtime monitoring of consistency and progress · 2025
Closest in time.
Science Robotics 10(106): eadt1497
Ai B, Tian S, Shi H, Wang Y, Pfaff T, Tan C, Christensen HI, Su H, Wu J and Li Y (2025) A review of learning-based dynamics models for robotic manipulation · 2025
Closest in time.
Transactions on Machine Learning Research
Alakuijala M, McLean R, Woungang I, Farsad N, Kaski S, Marttinen P and Yuan K (2025) Video-language critic: Transferable reward functions for language-conditioned robotics · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 8226–8233
Aljalbout E, Sotirakis N, van der Smagt P, Karl M and Chen N (2025) Limt: Language-informed multi-task visual world models · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 1233–1239
Ao J, Wu F, Wu Y, Swiki A and Haddadin S (2025) Llm-as-bt-planner: Leveraging llms for behavior tree generation in robot task planning · 2025
Closest in time.
Intelligent Service Robotics 18(2): 261–277
Asuzu K, Singh H and Idrissi M (2025) Human–robot interaction through joint robot planning with large language models · 2025
Closest in time.
In: Proceedings of Robotics: Science and Systems . LosAngeles, CA, USA
Black K, Brown N, Driess D, Esmail A, Equi MR, Finn C, Fusai N, Groom L, Hausman K, Ichter B, Jakubczak S, Jones T, Ke L, Levine S, Li-Bell A, Mothukuri M, Nair S, Pertsch K, Shi LX, Smith L, Tanner J, Vuong Q, Walling A, Wang H and Zhilinsky U (2025b) π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control · 2025
Closest in time.
In: 2025 Conference on Robot Learning . PMLR, pp. 4158–4187
Blank N, Reuss M, Rühle M, Yağmurlu ÖE, Wenzel F, Mees O and Lioutikov R (2025) Scaling robot policy learning via zero-shot labeling with foundation models · 2025
Closest in time.
IEEE Robotics and Automation Letters 10(5): 4810–4817
Brunke L, Zhang Y, Römer R, Naimer J, Staykov N, Zhou S and Schoellig AP (2025) Semantically safe robot manipulation: From semantic scene understanding to motion safeguards · 2025
Closest in time.
arXiv preprint arXiv:2506.21539
Cen J, Yu C, Yuan H, Jiang Y, Huang S, Guo J, Li X, Song Y, Luo H, Wang F et al. (2025) Worldvla: Towards autoregressive action world model · 2025
Closest in time.
arXiv preprint arXiv:2507.15493
Cheang C, Chen S, Cui Z, Hu Y, Huang L, Kong T, Li H, Li Y, Liu Y, Ma X et al. (2025) Gr-3 technical report · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 11972–11978
Chen K, Shen Z, Zhang Y, Chen L, Wu F, Bing Z, Haddadin S and Knoll A (2025b) Lemmo-plan: Llm-enhanced learning from multi-modal demonstration for planning sequential contact-rich manipulation tasks · 2025
Closest in time.
arXiv preprint arXiv:2508.08706
Cheng Z, Zhang Y, Zhang W, Li H, Wang K, Song L and Zhang H (2025) Omnivtla: Vision-tactile-language-action model with semantic-aligned tactile sensing · 2025
Closest in time.
arXiv preprint arXiv:2502.12861
Cleveston I, Santana AC, Costa PD, Gudwin RR, Simões AS and Colombini EL (2025) Instructrobot: A model-free framework for mapping natural language instructions into robot motion · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 15657–15664
Dai Y, Lee J, Fazeli N and Chai J (2025) Racer: Rich language-guided failure recovery policies for imitation learning · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 15617–15625
Dasari S, Mees O, Zhao S, Srirama MK and Levine S (2025) The ingredients for robotic diffusion transformers · 2025
Closest in time.
Neural Computing and Applications : 1–53
Deshpande S, Walambe R, Kotecha K, Selvachandran G and Abraham A (2025) Advances and applications in inverse reinforcement learning: a comprehensive review · 2025
Closest in time.
IEEE Robotics and Automation Letters 10(6): 5401–5408
Dong Q, Wu T, Zeng P, zang C, Wan G and Cui S (2025) Enhancing robot learning through cognitive reasoning trajectory optimization under unknown dynamics · 2025
Closest in time.
In: 2025 Conference on Robot Learning . PMLR, pp. 496–512
Doshi R, Walke HR, Mees O, Dasari S and Levine S (2025) Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation · 2025
Closest in time.
IEEE Access 13: 184071–184109
Eze C and Crick C (2025) Learning by watching: A review of video-based learning approaches for robot manipulation · 2025
Closest in time.
The International Journal of Robotics Research 44(5): 701–739
Firoozi R, Tucker J, Tian S, Majumdar A et al. (2025) Foundation models in robotics: Applications, challenges, and the future · 2025
Closest in time.
Nature 645(8081): 633–638
Guo D, Yang D, Zhang H, Song J, Wang P, Zhu Q, Xu R, Zhang R, Ma S, Bi X et al. (2025) Deepseek-r1 incentivizes reasoning in llms through reinforcement learning · 2025
Closest in time.
The International Journal of Robotics Research 44(4): 665–698
Habibian S, Alvarez Valdivia A, Blumenschein LH and Losey DP (2025) A survey of communicating robot learning during human-robot interaction · 2025
Closest in time.
Nature : 1–7
Hafner D, Pasukonis J, Ba J and Lillicrap T (2025) Mastering diverse control tasks through world models · 2025
Closest in time.
IEEE Robotics and Automation Letters 10(10): 9726–9733
Hao C, Lin K, Xue Z, Luo S and Soh H (2025) Disco: Language-guided manipulation with diffusion policies and constrained inpainting · 2025
Closest in time.
IEEE Robotics and Automation Letters 10(3): 2614–2621
Hatanaka W, Yamashina R and Matsubara T (2025) Reinforcement learning of flexible policies for symbolic instructions with adjustable mapping specifications · 2025
Closest in time.
In: Conference on Robot Learning . PMLR, pp. 666–688
Hirose N, Glossop C, Sridhar A, Mees O and Levine S (2025) Lelan: Learning a language-conditioned navigation policy from in-the-wild video · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 3617–3624
Hu J, Hendrix R, Farhadi A, Kembhavi A, Martín-Martín R, Stone P, Zeng KH and Ehsani K (2025a) Flare: Achieving masterful and adaptive robot policies with large-scale reinforcement learning fine-tuning · 2025
Closest in time.
In: 2025 Conference on Robot Learning . PMLR, pp. 4573–4602
Huang W, Wang C, Li Y, Zhang R and Fei-Fei L (2025b) Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation · 2025
Closest in time.
In: Human-to-Robot (H2R) Workshop at the 9th Conference on Robot Learning (CoRL 2025)
Hwang M, Forsey-Smerek A and Bobu A (2025) Masked Inverse Reinforcement Learning for Language Conditioned Reward Learning · 2025
Closest in time.
Robotics: Science and Systems (RSS)
Jenamani RK, Silver T, Dodson B, Tong S, Song A, Yang Y, Liu Z, Howe B, Whitneck A and Bhattacharjee T (2025) Feast: A flexible mealtime-assistance system towards in-the-wild personalization · 2025
Closest in time.
Artificial Life and Robotics : 1–10
Jin K, Tian G, Huang B, Cui Y and Zheng X (2025) Task-oriented adaptive learning of robot manipulation skills · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . pp. 5961–5968
Jones J, Mees O, Sferrazza C, Stachowicz K, Abbeel P and Levine S (2025) Beyond sight: Finetuning generalist robot policies with heterogeneous sensors via language grounding · 2025
Closest in time.
IEEE Access 13: 162467–162504
Kawaharazuka K, Oh J, Yamada J, Posner I and Zhu Y (2025) Vision-language-action models for robotics: A review towards real-world applications · 2025
Closest in time.
In: Conference on Robot Learning . PMLR, pp. 3647–3664
Kerr J, Hari K, Weber E, Kim CM, Yi B, Goldberg K, Kanazawa A et al. (2025) Eye, robot: Learning to look to act with a bc-rl perception-action loop · 2025
Closest in time.
In: Proceedings of Robotics: Science and Systems . LosAngeles, CA, USA
Kim MJ, Finn C and Liang P (2025b) Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success · 2025
Closest in time.
In: 2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) . pp. 1609–1616
Kobayashi T, Kobayashi M, Buamanee T and Uranishi Y (2025) Bi-lat: Bilateral control-based imitation learning via natural language and action chunking with transformers · 2025
Closest in time.
In: Proceedings of the Computer Vision and Pattern Recognition Conference . pp. 12166–12175
Lee S, Park S and Kim H (2025) Dynscene: Scalable generation of dynamic robotic manipulation scenes for embodied ai · 2025
Closest in time.
arXiv preprint arXiv:2508.20072
Liang Z, Li Y, Yang T, Wu C, Mao S, Pei L, Yang X, Pang J, Mu Y and Luo P (2025) Discrete diffusion vla: Bringing discrete diffusion to action decoding in vision-language-action policies · 2025
Closest in time.
In: The Thirteenth International Conference on Learning Representations
Lin F, Hu Y, Sheng P, Wen C, You J and Gao Y (2025) Data scaling laws in imitation learning for robotic manipulation · 2025
Closest in time.
Ling Y, Owalekar K, Adesanya O, Biyik E and Seita D (2025) IMPACT: intelligent motion planning with acceptable contact trajectories via vision-language models · 2025
Closest in time.
Conference on Robot Learning : 4651–4669
Ma T, Zhou J, Wang Z, Qiu R and Liang J (2025) Contrastive imitation learning for language-guided multi-task robotic manipulation · 2025
Closest in time.
Journal of Artificial Intelligence Research 83
McCarthy R, Tan DC, Schmidt D, Acero F, Herr N, Du Y, Thuruthel TG and Li Z (2025) Towards generalist robot learning from internet video: A survey · 2025
Closest in time.
Nature Machine Intelligence : 1–10
Mon-Williams R, Li G, Long R, Du W and Lucas CG (2025) Embodied large language models enable robots to complete complex tasks in unpredictable environments · 2025
Closest in time.
In: Conference on Robot Learning . PMLR, pp. 4996–5013
Nakamoto M, Mees O, Kumar A and Levine S (2025) Steering your generalists: Improving robotic foundation models via value guidance · 2025
Closest in time.
ACM Transactions on Intelligent Systems and Technology 16(5): 1–72
Naveed H, Khan AU, Qiu S, Saqib M, Anwar S, Usman M, Akhtar N, Barnes N and Mian A (2025) A comprehensive overview of large language models · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 8284–8290
Papagiannis G, Di Palo N, Vitiello P and Johns E (2025) R+ x: Retrieval and execution from everyday human videos · 2025
Closest in time.
In: Conference on Robot Learning . PMLR, pp. 4684–4696
Parimi V and Williams BC (2025) Diffusion-guided multi-arm motion planning · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 16233–16239
Paulius D, Agostini A, Quartey B and Konidaris G (2025) Bootstrapping object-level planning with large language models · 2025
Closest in time.
In: Proceedings of Robotics: Science and Systems
Pertsch K, Stachowicz K, Ichter B, Driess D, Nair S, Vuong Q, Mees O, Finn C and Levine S (2025) Fast: Efficient action tokenization for vision-language-action models · 2025
Closest in time.
Robotics: Science and Systems
Qu D, Song H, Chen Q, Yao Y, Ye X, Ding Y, Wang Z, Gu J, Zhao B, Wang D et al. (2025) Spatialvla: Exploring spatial representations for visual-language-action model · 2025
Closest in time.
In: 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) . IEEE, pp. 7277–7286
Singh CK, Kumar D, Sanap V and Sinha R (2025) Llm-rspf: Large language model-based robotic system planning framework for domain specific use-cases · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 1225–1232
Styrud J, Iovino M, Norrlöf M, Björkman M and Smith C (2025) Automatic behavior tree expansion with llms for robotic manipulation · 2025
Closest in time.
arXiv preprint arXiv:2508.09071
Sun L, Xie B, Liu Y, Shi H, Wang T and Cao J (2025) Geovla: Empowering 3d representations in vision-language-action models · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 10772–10778
Takeshita K, Yamazaki T, Ono T and Yamamoto T (2025) Robot local planner: A periodic sampling-based motion planner with minimal waypoints for home environments · 2025
Closest in time.
In: Workshop on Foundation Models Meet Embodied Agents at CVPR 2025
Tan S, Dou K, Zhao Y and Kraehenbuehl P (2025) Interactive post-training for vision-language-action models · 2025
Closest in time.
arXiv preprint arXiv:2503.20020
Team GR, Abeyruwan S, Ainslie J, Alayrac JB, Arenas MG, Armstrong T, Balakrishna A, Baruch R, Bauza M, Blokzijl M et al. (2025) Gemini robotics: Bringing ai into the physical world · 2025
Closest in time.
In: The Thirteenth International Conference on Learning Representations
Tian Y, Yang S, Zeng J, Wang P, Lin D, Dong H and Pang J (2025) Predictive inverse dynamics models are scalable learners for robotic manipulation · 2025
Closest in time.
In: 2025 11th International Conference on Automation, Robotics, and Applications (ICARA) . pp. 18–22
Tripathi D, Liu C, Pudasaini N, Hao Y, Tzes A and Fang Y (2025) Lav-act: Language-augmented visual action chunking with transformers for bimanual robotic manipulation · 2025
Closest in time.
IEEE Access 13: 69941–69949
Tsuji T (2025) Mamba as a motion encoder for robotic imitation learning · 2025
Closest in time.
IEEE Robotics and Automation Letters 10(9): 8618–8625
Tu Y, Wang Y, Zhang H, Chen W and Zhang J (2025) Language-embedded 6d pose estimation for tool manipulation · 2025
Closest in time.
IEEE Robotics and Automation Letters 10(9): 8850–8857
Turcato N, Iovino M, Synodinos A, Dalla Libera A, Carli R and Falco P (2025) Towards autonomous reinforcement learning for real-world robotic manipulation with large language models · 2025
Closest in time.
In: Conference on Robot Learning . PMLR, pp. 258–282
Wagenmaker A, Zhang Y, Nakamoto M, Park S, Yagoub W, Nagabandi A, Gupta A and Levine S (2025) Steering your diffusion policy with latent space reinforcement learning · 2025
Closest in time.
In: ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . pp. 1–5
Wang H, Qi L and Sun Y (2025b) Integrating failures in robot skill acquisition with offline action-sequence diffusion rl · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 15626–15633
Wang Y, Wang L, Du Y, Sundaralingam B, Yang X, Chao YW, Pérez-D’Arpino C, Fox D and Shah J (2025d) Inference-time policy steering through human interactions · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 8811–8818
Wu K, Zhu Y, Li J, Wen J, Liu N, Xu Z and Tang J (2025a) Discrete policy: Learning disentangled action space for multi-task robotic manipulation · 2025
Closest in time.
arXiv preprint arXiv:2506.07127
Xia W, Yang Y, Wu H, Ma X, Kong T and Hu D (2025) Robotic policy learning via human-assisted action preference optimization · 2025
Closest in time.
Neurocomputing : 129963
Xiao X, Liu J, Wang Z, Zhou Y, Qi Y, Jiang S, He B and Cheng Q (2025) Robot learning in the era of foundation models: A survey · 2025
Closest in time.
In: Conference on Robot Learning . PMLR, pp. 3139–3156
Xiong C, Shen C, Li X, Zhou K, Liu J, Wang R and Dong H (2025) Autonomous interactive correction mllm for robust robotic manipulation · 2025
Closest in time.
In: Proceedings of Robotics: Science and Systems (RSS)
Xue H, Ren J, Chen W, Zhang G, Fang Y, Gu G, Xu H and Lu C (2025) Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation · 2025
Closest in time.
International Journal of Computer Vision 133(2): 825–843
Zang Y, Li W, Han J, Zhou K and Loy CC (2025) Contextual object detection with multimodal large language models · 2025
Closest in time.
In: Conference on Robot Learning . PMLR, pp. 3157–3181
Zawalski M, Chen W, Pertsch K, Mees O, Finn C and Levine S (2025) Robotic control via embodied chain-of-thought reasoning · 2025
Closest in time.
In: 2025 Conference on Robot Learning . PMLR, pp. 933–946
Zhang J, Guo Y, Chen X, Wang YJ, Hu Y, Shi C and Chen J (2025a) Hirt: Enhancing robotic control with hierarchical robot transformers · 2025
Closest in time.
IEEE Access 13: 161194–161205
Zhang Y and Xue S (2025) A new trend in the warehousing: A review of human robot collaboration · 2025
Closest in time.
In: Proceedings of the Computer Vision and Pattern Recognition Conference . pp. 1702–1713
Zhao Q, Lu Y, Kim MJ, Fu Z, Zhang Z, Wu Y, Li Z, Ma Q, Han S, Finn C et al. (2025) Cot-vla: Visual chain-of-thought reasoning for vision-language-action models · 2025
Closest in time.
In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing . pp. 5377–5395
Zhou Z, Zhu Y, Zhu M, Wen J, Liu N, Xu Z, Meng W, Peng Y, Shen C, Feng F et al. (2025b) Chatvla: Unified multimodal understanding and robot control with vision-language-action model · 2025
Closest in time.
In: Proceedings of the AAAI Conference on Artificial Intelligence , volume 40. pp. 18135–18143
Bi H, Wu L, Lin T, Tan H, Su Z, Su H and Zhu J (2026) H-rdt: Human manipulation enhanced bimanual robotic manipulation · 2026
Closest in time.
IEEE Transactions on Robotics 42: 1158–1177
Chen G, Wang M, Cui T, Yang C, Hu M, Lu H, Peng Z, Zhou T, Jiang X, Yang Y and Yue Y (2026) Learning from videos through graph-to-graphs generative modeling for robotic manipulation · 2026
Closest in time.
Advances in Neural Information Processing Systems 38: 102867–102888
Driess D, Springenberg J, Ichter B, Yu L, Li-Bell A, Pertsch K, Ren A, Walke H, Vuong Q, Shi LX et al. (2026) Knowledge insulating vision-language-action models: Train fast, run fast, generalize better · 2026
Closest in time.
Advances in Neural Information Processing Systems 38: 136705–136736
Gao C, Liu Z, Chi Z, Huang J, Fei X, Hou Y, Zhang Y, Lin Y, Fang Z, Jiang Z and Lin S (2026) Vla-os: Structuring and dissecting planning representations and paradigms in vision-language-action models · 2026
Closest in time.
URL https://arxiv.org/abs/2603.21017
Hu Z, Yao X, Meng Y, Bing Z and Knoll A (2026) Dreaming the unseen: World model-regularized diffusion policy for out-of-distribution robustness · 2026
Closest in time.
In: The Fourteenth International Conference on Learning Representations
Shi H, Xie B, Liu Y, Sun L, Liu F, Wang T, Zhou E, Fan H, Zhang X and Huang G (2026) MemoryVLA: Perceptual-cognitive memory in vision-language-action models for robotic manipulation · 2026
Closest in time.
arXiv preprint arXiv:2603.03596
Torne M, Pertsch K, Walke H, Vedder K, Nair S, Ichter B, Ren AZ, Wang H, Tang J, Stachowicz K et al. (2026) Mem: Multi-scale embodied memory for vision language action models · 2026
Closest in time.
Advances in Neural Information Processing Systems 38: 93409–93439
Yu J, Liu H, Yu Q, Ren J, Hao C, Ding H, Huang G, Huang G, Song Y, Cai P et al. (2026) Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation · 2026
Closest in time.
In: 2025 Conference on Robot Learning . PMLR, pp. 1992–2028
Liu W, Nie N, Zhang R, Mao J and Wu J (2025b) Learning compositional behaviors from demonstration and language · 2028
Closest in time.
In: Conference on Robot Learning . PMLR, pp. 2012–2029
Chen L, Bahl S and Pathak D (2023b) Playfusion: Skill acquisition via diffusion from language-annotated play · 2029
Closest in time.
In: Proceedings of The 9th Conference on Robot Learning , volume 305. pp. 2018–2037
Fan Y, Bai S, Tong X, Ding P, Zhu Y, Lu H, Dai F, Zhao W, Liu Y, Huang S, Fan Z, Chen B and Wang D (2025) Long-vla: Unleashing long-horizon capability of vision language action model for robot manipulation · 2037
Closest in time.
In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, pp. 2086–2092
Ding Y, Zhang X, Paxton C and Zhang S (2023) Task and motion planning with large language models for object rearrangement · 2092
Closest in time.