Fetching the paper…
Reading the bibliography…
AI systems make decisions in physical environments through primitive actions or affordances that are accessed via API calls.
VRKitchen: an Interactive 3D Virtual Environment for Task-oriented Learning
Gao, X.; Gong, R.; Shu, T.; Xie, X.; Wang, S.; and Zhu, S.-C. 2019 · 1903
Earlier work this paper cites.
Thomason, J.; Murray, M.; Cakmak, M.; and Zettlemoyer, L. 2019 · 1907
Earlier work this paper cites.
On the Utility of Learning about Humans for Human-AI Coordination
Carroll, M.; Shah, R.; Ho, M. K.; Griffiths, T. L.; Seshia, S. A.; Abbeel, P.; and Dragan, A. D. 2019 · 1910
Earlier work this paper cites.
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
Ku, A.; Anderson, P.; Patel, R.; Ie, E.; and Baldridge, J. 2020 · 2010
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Deitke, M.; Ehsani, K.; Gordon, D.; Zhu, Y.; Kembhavi, A.; Gupta, A. K.; and Farhadi, A. 2017 · 2017
Earlier work this paper cites.
Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; and van den Hengel, A. 2018 · 2018
Earlier work this paper cites.
VirtualHome: Simulating Household Activities Via Programs
Puig, X.; Ra, K. K.; Boben, M.; Li, J.; Wang, T.; Fidler, S.; and Torralba, A. 2018 · 2018
Earlier work this paper cites.
Building Generalizable Agents with a Realistic and Rich 3D Environment
Wu, Y.; Wu, Y.; Gkioxari, G.; and Tian, Y. 2018 · 2018
Earlier work this paper cites.
Gibson Env: Real-World Perception for Embodied Agents
Xia, F.; Zamir, A.; He, Z.-Y.; Sax, A.; Malik, J.; and Savarese, S. 2018 · 2018
Earlier work this paper cites.
Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction
Misra, D.; Bennett, A.; Blukis, V.; Niklasson, E.; Shatkhin, M.; and Artzi, Y. 2019 · 2019
Earlier work this paper cites.
Shifting the Baseline: Single Modality Performance on Visual Navigation & QA
Thomason, J.; Gordon, D.; and Bisk, Y. 2019 · 2019
Earlier work this paper cites.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Shridhar, M.; Thomason, J.; Gordon, D.; Bisk, Y.; Han, W.; Mottaghi, R.; Zettlemoyer, L.; and Fox, D. 2020 · 2020
Earlier work this paper cites.
Reasoning about Goals, Steps, and Temporal Ordering with WikiHow
Zhang, L.; Lyu, Q.; and Callison-Burch, C. 2020 · 2020
Earlier work this paper cites.
Goal-Oriented Script Construction
Lyu, Q.; Zhang, L.; and Callison-Burch, C. 2021 · 2021
Earlier work this paper cites.
CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
Mees, O.; Hermann, L.; Rosete-Beas, E.; and Burgard, W. 2021 · 2021
Earlier work this paper cites.
CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge
Onoe, Y.; Zhang, M. J. Q.; Choi, E.; and Durrett, G. 2021 · 2021
Earlier work this paper cites.
TEACh: Task-driven Embodied Agents that Chat
Padmakumar, A.; Thomason, J.; Shrivastava, A.; Lange, P.; Narayan-Chen, A.; Gella, S.; Piramuthu, R.; Tur, G.; and Hakkani-Tur, D. 2021 · 2021
Cited alongside, same era.
Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
Ramakrishnan, S. K.; Gokaslan, A.; Wijmans, E.; Maksymets, O.; Clegg, A.; Turner, J.; Undersander, E.; Galuba, W.; Westbury, A.; Chang, A. X.; Savva, M.; Zhao, Y.; and Batra, D. 2021 · 2021
Cited alongside, same era.
proScript: Partially Ordered Scripts Generation
Sakaguchi, K.; Bhagavatula, C.; Le Bras, R.; Tandon, N.; Clark, P.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments
Srivastava, S.; Li, C.; Lingelbach, M.; Mart’in-Mart’in, R.; Xia, F.; Vainio, K.; Lian, Z.; Gokmen, C.; Buch, S.; Liu, C. K.; Savarese, S.; Gweon, H.; Wu, J.; and Fei-Fei, L. 2021 · 2021
Cited alongside, same era.
MERLOT: Multimodal Neural Script Knowledge Models
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.-T.; Lin, Y.; Cong, X.; Tang, X.; Qian, B.; Zhao, S.; Tian, R.; Xie, R.; Zhou, J.; Gerstein, M. H.; Li, D.; Liu, Z.; and Sun, M. 2023 · 2023
Later among the works it cites.
ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
Singh, I.; Blukis, V.; Mousavian, A.; Goyal, A.; Xu, D.; Tremblay, J.; Fox, D.; Thomason, J.; and Garg, A. 2022 · 2023
Later among the works it cites.
LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models
Song, C. H.; Wu, J.; Washington, C.; Sadler, B. M.; Chao, W.-L.; and Su, Y. 2023 · 2023
Later among the works it cites.
ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
Tang, Q.; Deng, Z.; Lin, H.; Han, X.; Liang, Q.; Cao, B.; and Sun, L. 2023 · 2023
Later among the works it cites.
Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zellers, R.; Lu, X.; Hessel, J.; Yu, Y.; Park, J. S.; Cao, J.; Farhadi, A.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Ahn, M.; Brohan, A.; Brown, N.; Chebotar, Y.; Cortes, O.; David, B.; Finn, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; Ho, D.; Hsu, J.; Ibarz, J.; Ichter, B.; Irpan, A.; Jang, E.; Ruano, R. J.; Jeffrey, K.; Jesmonth, S.; Joshi, N. J.; Julian, R. C.; Kalashnikov, D.; Kuang, Y.; Lee, K.-H.; Levine, S.; Lu, Y.; Luu, L.; Parada, C.; Pastor, P.; Quiambao, J.; Rao, K.; Rettinghouse, J.; Reyes, D. M.; Sermanet, P.; Sievers, N.; Tan, C.; Toshev, A.; Vanhoucke, V.; Xia, F.; Xiao, T.; Xu, P.; Xu, S.; and Yan, M. 2022 · 2022
Cited alongside, same era.
The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark Towards Physically Realistic Embodied AI
Gan, C.; Zhou, S.; Schwartz, J.; Alter, S.; Bhandwaldar, A.; Gutfreund, D.; Yamins, D. L. K.; DiCarlo, J. J.; McDermott, J. H.; Torralba, A.; and Tenenbaum, J. B. 2021 · 2022
Cited alongside, same era.
Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
Huang, W.; Abbeel, P.; Pathak, D.; and Mordatch, I. 2022 · 2022
Cited alongside, same era.
Show Me More Details: Discovering Hierarchies of Procedures from Semi-structured Web Data
Zhou, S.; Zhang, L.; Yang, Y.; Lyu, Q.; Yin, P.; Callison-Burch, C.; and Neubig, G. 2022 · 2022
Cited alongside, same era.
Learning Universal Policies via Text-Guided Video Generation
Du, Y.; Yang, M.; Dai, B.; Dai, H.; Nachum, O.; Tenenbaum, J. B.; Schuurmans, D.; and Abbeel, P. 2023 · 2023
Cited alongside, same era.
Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied Agents
Huang, W.; Xia, F.; Shah, D.; Driess, D.; Zeng, A.; Lu, Y.; Florence, P.; Mordatch, I.; Levine, S.; Hausman, K.; and Ichter, B. 2023 · 2023
Cited alongside, same era.
Mini-BEHAVIOR: A Procedurally Generated Benchmark for Long-horizon Decision-Making in Embodied AI
Jin, E.; Hu, J.; Huang, Z.; Zhang, R.; Wu, J.; Li, F.-F.; and Mart’in-Mart’in, R. 2023 · 2023
Cited alongside, same era.
Uzunoglu, A.; and Şahin, G. 2023 · 2023
Later among the works it cites.
SmartPlay : A Benchmark for LLMs as Intelligent Agents
Wu, Y.; Tang, X.; Mitchell, T. M.; and Li, Y. 2023 · 2023
Later among the works it cites.
On the Tool Manipulation Capability of Open-source Large Language Models
Xu, Q.; Hong, F.; Li, B.; Hu, C.; Chen, Z.; and Zhang, J. 2023 · 2023
Later among the works it cites.
LACMA: Language-Aligning Contrastive Learning with Meta-Actions for Embodied Instruction Following
Yang, C.; Chen, Y.-C.; Yang, J.; Dai, X.; Yuan, L.; Wang, Y.-C. F.; and Chang, K.-W. 2023 · 2023
Later among the works it cites.
Procedure-Aware Pretraining for Instructional Video Understanding
Zhou, H.; Martín-Martín, R.; Kapadia, M.; Savarese, S.; and Niebles, J. C. 2023 · 2023
Later among the works it cites.
API Pack: A Massive Multilingual Dataset for API Call Generation
Guo, Z.; Soria, A. M.; Sun, W.; Shen, Y.; and Panda, R. 2024 · 2024
Closest in time.
SELF-[IN]CORRECT: LLMs Struggle with Refining Self-Generated Responses
Jiang, D.; Zhang, J.; Weller, O.; Weir, N.; Durme, B. V.; and Khashabi, D. 2024 · 2024
Closest in time.
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
Li, C.; Zhang, R.; Wong, J.; Gokmen, C.; Srivastava, S.; Martín-Martín, R.; Wang, C.; Levine, G.; Ai, W.; Martinez, B.; Yin, H.; Lingelbach, M.; Hwang, M.; Hiranaka, A.; Garlanka, S. S.; Aydin, A.; Lee, S.; Sun, J.; Anvari, M.; Sharma, M.; Bansal, D.; Hunter, S.; Kim, K.-Y.; Lou, A.; Matthews, C. R.; Villa-Renteria, I.; Tang, J. H.; Tang, C.; Xia, F.; Li, Y.; Savarese, S.; Gweon, H.; Liu, C. K.; Wu, J.; Li, F.-F.; and Research, S. 2024 · 2024
Closest in time.
RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
Nasiriany, S.; Maddukuri, A.; Zhang, L.; Parikh, A.; Lo, A.; Joshi, A.; Mandlekar, A.; and Zhu, Y. 2024 · 2024
Closest in time.
Uzunoglu, A.; Safa, A. R.; and Şahin, G. G. 2024 · 2024
Closest in time.
KITchen: A Real-World Benchmark and Dataset for 6D Object Pose Estimation in Kitchen Environments
Younes, A.; and Asfour, T. 2024 · 2024
Closest in time.