Fetching the paper…
Reading the bibliography…
The ability to detect and analyze failed executions automatically is crucial for an explainable and robust robotic system.
Ai2-thor: An interactive 3d environment for visual ai
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, M. Deitke, K. Ehsani, D. Gordon, Y. Zhu, et al · 2017
Earlier work this paper cites.
Human trust after robot mistakes: Study of the effects of different forms of robot communication
S. Ye, G. Neville, M. Schrum, M. Gombolay, S. Chernova, and A. Howard · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
End-to-end dense video captioning with parallel decoding
T. Wang, R. Zhang, Z. Lu, F. Zheng, R. Cheng, and P. Luo · 2021
Earlier work this paper cites.
Sketch, ground, and refine: Top-down dense video captioning
C. Deng, S. Chen, D. Chen, Y. He, and Q. Wu · 2021
Earlier work this paper cites.
Explainable ai for robot failures: Generating explanations that improve user assistance in fault recovery
D. Das, S. Banerjee, and S. Chernova · 2021
Earlier work this paper cites.
Semantic-based explainable ai: Leveraging semantic scene graphs and pairwise ranking to explain robot failures
D. Das and S. Chernova · 2021
Earlier work this paper cites.
Fino-net: A deep multimodal sensor fusion framework for manipulation failure detection
A. Inceoglu, E. E. Aksoy, A. C. Ak, and S. Sariel · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Earlier work this paper cites.
Do as i can and not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter · 2022
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2022
Cited alongside, same era.
Unifying event detection and captioning as sequence generation via pre-training
Q. Zhang, Y. Song, and Q. Jin · 2022
Cited alongside, same era.
End-to-end dense video captioning as sequence generation
W. Zhu, B. Pang, A. V. Thapliyal, W. Y. Wang, and R. Soricut · 2022
Cited alongside, same era.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, M. Attarian, B. Ichter, K. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, and P. Florence · 2022
Wav2clip: Learning robust audio representations from clip
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello · 2022
Later among the works it cites.
Embodied semantic scene graph generation
X. Li, D. Guo, H. Liu, and F. Sun · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models with image descriptors are strong few-shot video-language learners
Z. Wang, M. Li, R. Xu, L. Zhou, J. Lei, X. Lin, S. Wang, Z. Yang, C. Zhu, D. Hoiem, et al · 2022
Cited alongside, same era.
Summarizing a virtual robot’s past actions in natural language
C. DeChant and D. Bauer · 2022
Cited alongside, same era.
Why did i fail? a causal-based method to find explanations for robot failures
M. Diehl and K. Ramirez-Amaro · 2022
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
C. H. Song, J. Wu, C. Washington, B. M. Sadler, W.-L. Chao, and Y. Su · 2022
Cited alongside, same era.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2022
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Cited alongside, same era.
Planning with large language models via corrective re-prompting
S. S. Raman, V. Cohen, E. Rosen, I. Idrees, D. Paulius, and S. Tellex · 2022
Cited alongside, same era.
R.-G. Pasca, A. Gavryushin, Y.-L. Kuo, O. Hilliges, and X. Wang · 2023
Closest in time.
P. Khanna, E. Yadollahi, M. Björkman, I. Leite, and C. Smith · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
N. Shinn, B. Labash, and A. Gopinath · 2023
Closest in time.
M. Skreta, N. Yoshikawa, S. Arellano-Rubach, Z. Ji, L. B. Kristensen, K. Darvish, A. Aspuru-Guzik, F. Shkurti, and A. Garg · 2023
Closest in time.
Generalized planning in pddl domains with pretrained large language models
T. Silver, S. Dan, K. Srinivas, J. B. Tenenbaum, L. P. Kaelbling, and M. Katz · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
Visual spatial reasoning
F. Liu, G. Emerson, and N. Collier · 2023
Closest in time.
Modeling dynamic environments with scene graph memory
A. Kurenkov, M. Lingelbach, T. Agarwal, E. Jin, C. Li, R. Zhang, L. Fei-Fei, J. Wu, S. Savarese, and R. Martın-Martın · 2023
Closest in time.