Fetching the paper…
Reading the bibliography…
Assistive agents should be able to perform under-specified long-horizon tasks while respecting user preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
The RobotSlang Benchmark: Dialog-guided Robot Localization and Navigation
Shurjo Banerjee, Jesse Thomason, and Jason J. Corso · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca Dragan, Shankar Sastry, and Sanjit Seshia · 2017
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba · 2018
Earlier work this paper cites.
Collaborative dialogue in Minecraft
Anjali Narayan-Chen, Prashant Jayannavar, and Julia Hockenmaier · 2019
Earlier work this paper cites.
Continuous control for high-dimensional state spaces: An interactive learning approach
Rodrigo Pérez-Dattari, Carlos Celemin, Javier Ruiz-del Solar, and Jens Kober · 2019
Earlier work this paper cites.
Vision-and-dialog navigation
Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
MindCraft: Theory of mind modeling for situated dialogue in collaborative tasks
Cristian-Paul Bara, Sky CH-Wang, and Joyce Chai · 2021
Earlier work this paper cites.
Learning reward functions from diverse sources of human feedback: Optimally integrating demonstrations and preferences
Erdem Bıyık, Dylan P Losey, Malayandi Palan, Nicholas C Landolfi, Gleb Shevchuk, and Dorsa Sadigh · 2022
Earlier work this paper cites.
Training language models with language feedback
Jon Ander Campos and Jun Shern · 2022
Earlier work this paper cites.
Dialog Acts for Task-Driven Embodied Agents
Spandana Gella, Aishwarya Padmakumar, Patrick Lange, and Dilek Hakkani-Tur · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Earlier work this paper cites.
My house, my rules: Learning tidying preferences with graph neural networks
Ivan Kapelyukh and Edward Johns · 2022
Earlier work this paper cites.
Robot learning on the job: Human-in-the-loop autonomy and learning during deployment
Huihan Liu, Soroush Nasiriany, Lance Zhang, Zhiyao Bao, and Yuke Zhu · 2022
Earlier work this paper cites.
TEACh: Task-driven Embodied Agents that Chat
Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gokhan Tur, and Dilek Hakkani-Tur · 2022
Cited alongside, same era.
Correcting Robot Plans with Natural Language Feedback
Pratyusha Sharma, Balakumar Sundaralingam, Valts Blukis, Chris Paxton, Tucker Hermans, Antonio Torralba, Jacob Andreas, and Dieter Fox · 2022
Cited alongside, same era.
Ask4help: Learning to leverage an expert for embodied tasks
Kunal Pratap Singh, Luca Weihs, Alvaro Herrasti, Jonghyun Choi, Aniruddha Kembhavi, and Roozbeh Mottaghi · 2022
Cited alongside, same era.
Chai: A chatbot ai for task-oriented dialogue with offline reinforcement learning
Siddharth Verma, Justin Fu, Mengjiao Yang, and Sergey Levine · 2022
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et al · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao · 2023
Later among the works it cites.
Homerobot: Open-vocabulary mobile manipulation
Sriram Yenamandra, Arun Ramachandran, Karmesh Yadav, Austin S Wang, Mukul Khanna, Theophile Gervet, Tsung-Yen Yang, Vidhi Jain, Alexander Clegg, John M Turner, Zsolt Kira, Manolis Savva, Angel X Chang, Devendra Singh Chaplot, Dhruv Batra, Roozbeh Mottaghi, Yonatan Bisk, and Chris Paxton · 2023
Later among the works it cites.
Good time to ask: A learning framework for asking for help in embodied visual navigation
Jenny Zhang, Samson Yu, Jiafei Duan, and Cheston Tan · 2023
Later among the works it cites.
STar-GATE: Teaching language models to ask clarifying questions
Chinmaya Andukuri, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah Goodman · 2024
Later among the works it cites.
Incremental learning of humanoid robot behavior from natural interaction and large language models
Leonard Bärmann, Rainer Kartmann, Fabian Peller-Konrad, Jan Niehues, Alex Waibel, and Tamim Asfour · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grammar-constrained decoding for structured NLP tasks without finetuning
Saibo Geng, Martin Josifoski, Maxime Peyrard, and Robert West · 2023
Cited alongside, same era.
Zero-shot goal-directed dialogue via rl on imagined conversations
Joey Hong, Sergey Levine, and Anca Dragan · 2023
Cited alongside, same era.
Transformers are adaptable task planners
Vidhi Jain, Yixin Lin, Eric Undersander, Yonatan Bisk, and Akshara Rai · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2023
Cited alongside, same era.
Llm+ p: Empowering large language models with optimal planning proficiency
Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone · 2023
Cited alongside, same era.
Clara: classifying and disambiguating user commands for reliable interactive robotic agents
Jeongeun Park, Seungwon Lim, Joonhyung Lee, Sangbeom Park, Minsuk Chang, Youngjae Yu, and Sungjoon Choi · 2023
Cited alongside, same era.
Active preference inference using language models and probabilistic reasoning
Wasu Top Piriyakulkij, Volodymyr Kuleshov, and Kevin Ellis · 2023
Cited alongside, same era.
Bayesian preference elicitation with language models
Kunal Handa, Yarin Gal, Ellie Pavlick, Noah Goodman, Jacob Andreas, Alex Tamkin, and Belinda Z Li · 2024
Later among the works it cites.
Learning to learn faster from human feedback with language model predictive control
Jacky Liang, Fei Xia, Wenhao Yu, Andy Zeng, Montserrat Gonzalez Arenas, Maria Attarian, Maria Bauza, Matthew Bennice, Alex Bewley, Adil Dostmohamed, Chuyuan Kelly Fu, Nimrod Gileadi, Marissa Giustina, Keerthana Gopalakrishnan, Leonard Hasenclever, Jan Humplik, Jasmine Hsu, Nikhil Joshi, Ben Jyenis, Chase Kew, Sean Kirmani, Tsang-Wei Edward Lee, Kuang-Huei Lee, Assaf Hurwitz Michaely, Joss Moore, Ken Oslund, Dushyant Rao, Allen Ren, Baruch Tabanpour, Quan Vuong, Ayzaan Wahid, Ted Xiao, Ying Xu, Vincent Zhuang, Peng Xu, Erik Frey, Ken Caluwaerts, Tingnan Zhang, Brian Ichter, Jonathan Tompson, Leila Takayama, Vincent Vanhoucke, Izhak Shafran, Maja Mataric, Dorsa Sadigh, Nicolas Heess, Kanishka Rao, Nik Stewart, Jie Tan, and Carolina Parada · 2024
Later among the works it cites.
Decision-oriented dialogue for human-AI collaboration
Jessy Lin, Nicholas Tomlin, Jacob Andreas, and Jason Eisner · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
APRICOT: Active preference learning and constraint-aware task planning with LLMs
Huaxiaoyue Wang, Nathaniel Chin, Gonzalo Gonzalez-Pumariega, Xiangwan Sun, Neha Sunkara, Maximus Adrian Pace, Jeannette Bohg, and Sanjiban Choudhury · 2024
Later among the works it cites.
Trajectory improvement and reward learning from comparative language feedback
Zhaojing Yang, Miru Jun, Jeremy Tien, Stuart J. Russell, Anca Dragan, and Erdem Biyik · 2024
Later among the works it cites.
PARTNR: A benchmark for planning and reasoning in embodied multi-agent tasks
Matthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote, Ruta Desai, Michal Hlavac, Vladimir Karashchuk, Jacob Krantz, Roozbeh Mottaghi, Priyam Parashar, Siddharth Patki, Ishita Prasad, Xavier Puig, Akshara Rai, Ram Ramrakhya, Daniel Tran, Joanne Truong, John M Turner, Eric Undersander, and Tsung-Yen Yang · 2025
Closest in time.
LLM-personalize: Aligning LLM planners with human preferences via reinforced self-training for housekeeping robots
Dongge Han, Trevor McInroe, Adam Jelley, Stefano V. Albrecht, Peter Bell, and Amos Storkey · 2025
Closest in time.
Mile: Model-based intervention learning
Yigit Korkmaz and Erdem Bıyık · 2025
Closest in time.
Beyond behavior cloning: Robustness through interactive imitation and contrastive learning
Zhaoting Li, Rodrigo Pérez-Dattari, Robert Babuska, Cosimo Della Santina, and Jens Kober · 2025
Closest in time.