Fetching the paper…
Reading the bibliography…
Recent years have seen an increasing amount of work on embodied AI agents that can perform tasks by following human language instructions.
MMDetection: Open mmlab detection toolbox and benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. 2019 · 1906
Earlier work this paper cites.
Conceptualising and developing agents
M Wooldridge. 1995 · 1995
Earlier work this paper cites.
Probabilistic roadmaps for path planning in high-dimensional configuration spaces
L.E. Kavraki, P. Svestka, J.-C. Latombe, and M.H. Overmars. 1996 · 1996
Earlier work this paper cites.
Pddl| the planning domain definition language
Constructions Aeronautiques, Adele Howe, Craig Knoblock, ISI Drew McDermott, Ashwin Ram, Manuela Veloso, Daniel Weld, David Wilkins SRI, Anthony Barrett, Dave Christianson, et al. 1998 · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh. 1999 · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G. Dietterich. 2000 · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy McGovern and Andrew G. Barto. 2001 · 2001
Earlier work this paper cites.
Improving efficiency in mobile robot task planning through world abstraction
C. Galindo, J.-A. Fernandez-Madrigal, and J. Gonzalez. 2004 · 2004
Earlier work this paper cites.
The fast downward planning system
M. Helmert. 2006 · 2006
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L Littman. 2008 · 2008
Earlier work this paper cites.
Robot task planning using semantic maps
Cipriano Galindo, Juan-Antonio Fernández-Madrigal, Javier González, and Alessandro Saffiotti. 2008 · 2008
Earlier work this paper cites.
Robot learning from demonstration by constructing skill trees
George Konidaris, Scott Kuindersma, Roderic Grupen, and Andrew Barto. 2012 · 2012
Earlier work this paper cites.
Back to the blocks world: Learning new actions through situated human-robot dialogue
Lanbo She, Shaohua Yang, Yu Cheng, Yunyi Jia, Joyce Chai, and Ning Xi. 2014 · 2014
Earlier work this paper cites.
Grounding english commands to reward functions
James MacGlashan, Monica Babes-Vroman, Marie desJardins, Michael L. Littman, Smaranda Muresan, S. Squire, Stefanie Tellex, Dilip Arumugam, and Lei Yang. 2015 · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Hierarchical relative entropy policy search
Christian Daniel, Gerhard Neumann, Oliver Kroemer, and Jan Peters. 2016 · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum. 2016 · 2016
Earlier work this paper cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup. 2017 · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. 2017 · 2017
Cited alongside, same era.
Interactive learning of grounded verb semantics towards human-robot communication
Lanbo She and Joyce Chai. 2017 · 2017
Cited alongside, same era.
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. 2018 · 2018
Cited alongside, same era.
Language to action: Towards interactive task learning with physical agents
Joyce Chai, Qiaozi Gao, Lanbo She, Shaohua Yang, Sari Saba-Sadiya, and Guangyue Xu. 2018 · 2018
Cited alongside, same era.
3d scene graph: A structure for unified semantics, 3d space, and camera
Iro Armeni, Zhi-Yang He, Amir Zamir, Junyoung Gwak, Jitendra Malik, Martin Fischer, and Silvio Savarese. 2019 · 2019
Cited alongside, same era.
Learning symbolic operators for task and motion planning
Tom Silver, Rohan Chitnis, Joshua Tenenbaum, Leslie Pack Kaelbling, and Tomás Lozano-Pérez. 2021 · 2021
Later among the works it cites.
Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments
Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, et al. 2021 · 2021
Later among the works it cites.
Look wide and interpret twice: Improving performance on interactive instruction-following tasks
Masanori Suganuma, Takayuki Okatani, et al. 2021 · 2021
Later among the works it cites.
Embodied bert: A transformer model for embodied, language-guided visual task completion
Alessandro Suglia, Qiaozi Gao, Jesse Thomason, Govind Thattai, and Gaurav S. Sukhatme. 2021 · 2021
Later among the works it cites.
Hierarchical task learning from language instructions with unified transformers and self-monitoring
Yichi Zhang and Joyce Chai. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Panoptic segmentation
Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollár. 2019 · 2019
Cited alongside, same era.
Vision-and-dialog navigation
Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
Object goal navigation using goal-oriented semantic exploration
Devendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, and Ruslan Salakhutdinov. 2020 · 2020
Cited alongside, same era.
3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents
Ue-Hwan Kim, Jin-Man Park, Taek-jin Song, and Jong-Hwan Kim. 2020 · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Reverie: Remote embodied visual referring expression in real indoor environments
Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. 2020 · 2020
Cited alongside, same era.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Taskography: Evaluating robot task planning over large 3d scene graphs
Christopher Agia, Krishna Murthy Jatavallabhula, Mohamed Khodeir, Ondrej Miksik, Vibhav Vineet, Mustafa Mukadam, Liam Paull, and Florian Shkurti. 2022 · 2022
Closest in time.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, et al. 2022 · 2022
Closest in time.
A persistent spatial semantic representation for high-level natural language instruction execution
Valts Blukis, Chris Paxton, Dieter Fox, Animesh Garg, and Yoav Artzi. 2022 · 2022
Closest in time.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. 2022 · 2022
Closest in time.
Simple but effective: Clip embeddings for embodied ai
Apoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, and Aniruddha Kembhavi. 2022 · 2022
Closest in time.
Embodied semantic scene graph generation
Xinghang Li, Di Guo, Huaping Liu, and Fuchun Sun. 2022 · 2022
Closest in time.
FILM: Following instructions in language with modular methods
So Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk, and Ruslan Salakhutdinov. 2022 · 2022
Closest in time.
Skill induction and planning with latent language
Pratyusha Sharma, Antonio Torralba, and Jacob Andreas. 2022 · 2022
Closest in time.
Generalizable task planning through representation pretraining
Chen Wang, Danfei Xu, and Li Fei-Fei. 2022 · 2022
Closest in time.
Sgl: Symbolic goal learning in a hybrid, modular framework for human instruction following
Ruinian Xu, Hongyi Chen, Yunzhi Lin, and Patricio A Vela. 2022 · 2022
Closest in time.
Jarvis: A neuro-symbolic commonsense reasoning framework for conversational embodied agents
Kaizhi Zheng, Kaiwen Zhou, Jing Gu, Yue Fan, Jialu Wang, Zonglin Li, Xuehai He, and Xin Eric Wang. 2022 · 2022
Closest in time.
Piglet: Language grounding through neuro-symbolic interaction in a 3d world
Rowan Zellers, Ari Holtzman, Matthew E Peters, Roozbeh Mottaghi, Aniruddha Kembhavi, Ali Farhadi, and Yejin Choi. 2021 · 2050
Closest in time.