Fetching the paper…
Reading the bibliography…
Equipping embodied agents with commonsense is important for robots to successfully complete complex human instructions in general environments.
The site selection of distribution center based on linear programming transportation method
X. Liu · 2012
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, M. Deitke, K. Ehsani, D. Gordon, Y. Zhu, et al · 2017
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
X. Puig, K. Ra, M. Boben, J. Li, T. Wang, S. Fidler, and A. Torralba · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. D. M.-W. C. Kenton and L. K. Toutanova · 2019
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Gpt-3: Its nature, scope, limits, and consequences
L. Floridi and M. Chiriatti · 2020
Earlier work this paper cites.
Grounding language to autonomously-acquired skills via goal generation
A. Akakzia, C. Colas, P.-Y. Oudeyer, M. Chetouani, and O. Sigaud · 2020
Earlier work this paper cites.
Site selection and layout of earthquake rescue center based on k-means clustering and fruit fly optimization algorithm
X.-Y. Jiang, N.-Y. Pa, W.-C. Wang, T.-T. Yang, and W.-T. Pan · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Open-vocabulary image segmentation
G. Ghiasi, X. Gu, Y. Cui, and T.-Y. Lin · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Piglet: Language grounding through neuro-symbolic interaction in a 3d world
R. Zellers, A. Holtzman, M. Peters, R. Mottaghi, A. Kembhavi, A. Farhadi, and Y. Choi · 2021
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Earlier work this paper cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
C. H. Song, J. Wu, C. Washington, B. M. Sadler, W.-L. Chao, and Y. Su · 2022
Earlier work this paper cites.
Bridging the gap between object and image-level representations for open-vocabulary detection
H. Bangalath, M. Maaz, M. U. Khattak, S. H. Khan, and F. Shahbaz Khan · 2022
Cited alongside, same era.
Detecting twenty-thousand classes using image-level supervision
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Cited alongside, same era.
Nlx-gpt: A model for natural language explanations in vision and vision-language tasks
F. Sammani, T. Mukherjee, and N. Deligiannis · 2022
Cited alongside, same era.
Solving math word problems concerning systems of equations with gpt-3
M. Zong and B. Krishnamachari · 2022
Cited alongside, same era.
Smart explorer: Recognizing objects in dense clutter via interactive exploration
Z. Wu, Z. Wang, Z. Wei, Y. Wei, and H. Yan · 2022
Cited alongside, same era.
Ge-grasp: Efficient target-oriented grasping in dense clutter
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
R. Zhang, J. Han, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, P. Gao, and Y. Qiao · 2023
Closest in time.
B. Peng, C. Li, P. He, M. Galley, and J. Gao · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Closest in time.
Do as i can, not as i say: Grounding language in robotic affordances
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al · 2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Liu, Z. Wang, S. Huang, J. Zhou, and J. Lu · 2022
Cited alongside, same era.
Back to reality: Weakly-supervised 3d object detection with shape-guided label enhancement
X. Xu, Y. Wang, Y. Zheng, Y. Rao, J. Zhou, and J. Lu · 2022
Cited alongside, same era.
A persistent spatial semantic representation for high-level natural language instruction execution
V. Blukis, C. Paxton, D. Fox, A. Garg, and Y. Artzi · 2022
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Cited alongside, same era.
Pre-trained language models for interactive decision-making
S. Li, X. Puig, C. Paxton, Y. Du, C. Wang, L. Fan, T. Chen, D.-A. Huang, E. Akyürek, A. Anandkumar, et al · 2022
Cited alongside, same era.
Tidybot: Personalized robot assistance with large language models
J. Wu, R. Antonova, A. Kan, M. Lepert, A. Zeng, S. Song, J. Bohg, S. Rusinkiewicz, and T. Funkhouser · 2023
Cited alongside, same era.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao · 2023
Cited alongside, same era.
Closest in time.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
Contextual object detection with multimodal large language models
Y. Zang, W. Li, J. Han, K. Zhou, and C. C. Loy · 2023
Closest in time.
Dspdet3d: Dynamic spatial pruning for 3d small object detection
X. Xu, Z. Sun, Z. Wang, H. Liu, J. Zhou, and J. Lu · 2023
Closest in time.
Otter: A multi-modal model with in-context instruction tuning
B. Li, Y. Zhang, L. Chen, J. Wang, J. Yang, and Z. Liu · 2023
Closest in time.
Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration
C. Lyu, M. Wu, L. Wang, X. Huang, B. Liu, Z. Du, S. Shi, and Z. Tu · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, et al · 2023
Closest in time.
Assistgpt: A general multi-modal assistant that can plan, execute, inspect, and learn
D. Gao, L. Ji, L. Zhou, K. Q. Lin, J. Chen, Z. Fan, and M. Z. Shou · 2023
Closest in time.
Chatbridge: Bridging modalities with large language model as a language catalyst
Z. Zhao, L. Guo, T. Yue, S. Chen, S. Shao, X. Zhu, Z. Yuan, and J. Liu · 2023
Closest in time.
F. Chen, M. Han, H. Zhao, Q. Zhang, J. Shi, S. Xu, and B. Xu · 2023
Closest in time.