Fetching the paper…
Reading the bibliography…
We investigate the challenge of task planning for multi-task embodied agents in open-world environments.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
The malmo platform for artificial intelligence experimentation
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell · 2016
Earlier work this paper cites.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Earlier work this paper cites.
Zero-shot task generalization with multi-task deep reinforcement learning
J. Oh, S. Singh, H. Lee, and P. Kohli · 2017
Earlier work this paper cites.
W. H. Guss, C. Codel, K. Hofmann, B. Houghton, N. Kuno, S. Milani, S. Mohanty, D. P. Liebana, R. Salakhutdinov, N. Topin, et al · 2019
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
W. H. Guss, B. Houghton, N. Topin, P. Wang, C. Codel, M. Veloso, and R. Salakhutdinov · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Align-rudder: Learning from few demonstrations by reward redistribution
V. P. Patil, M. Hofmarcher, M.-C. Dinu, M. Dorfer, P. M. Blies, J. Brandstetter, J. A. Arjona-Medina, and S. Hochreiter · 2020
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
M. Shridhar, X. Yuan, M.-A. Côté, Y. Bisk, A. Trischler, and M. Hausknecht · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
The minerl 2020 competition on sample efficient reinforcement learning using human priors
W. H. Guss, M. Y. Castro, S. Devlin, B. Houghton, N. S. Kuno, C. Loomis, S. Milani, S. P. Mohanty, K. Nakata, R. Salakhutdinov, J. Schulman, S. Shiroshita, N. Topin, A. Ummadisingu, and O. Vinyals · 2021
Earlier work this paper cites.
Juewu-mc: Playing minecraft with sample-efficient hierarchical reinforcement learning
Z. Lin, J. Li, J. Shi, D. Ye, Q. Fu, and W. Yang · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Forgetful experience replay in hierarchical reinforcement learning from expert demonstrations
A. Skrynnik, A. Staroverov, E. Aitygulov, K. Aksenov, V. Davydov, and A. I. Panov · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al · 2022
Cited alongside, same era.
Collaborating with language models for embodied reasoning
I. Dasgupta, C. Kaeser-Chen, K. Marino, A. Ahuja, S. Babayan, F. Hill, and R. Fergus · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, et al · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
S. Cai, Z. Wang, X. Ma, A. Liu, and Y. Liang · 2023
Closest in time.
Groot: Learning to follow instructions by watching gameplay videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2022
Cited alongside, same era.
Minerl diamond 2021 competition: Overview, results, and lessons learned
A. Kanervisto, S. Milani, K. Ramanauskas, N. Topin, Z. Lin, J. Li, J. Shi, D. Ye, Q. Fu, W. Yang, W. Hong, Z. Huang, H. Chen, G. Zeng, Y. Lin, V. Micheli, E. Alonso, F. Fleuret, A. Nikulin, Y. Belousov, O. Svidchenko, and A. Shpilman · 2022
Cited alongside, same era.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2022
Cited alongside, same era.
Sqa3d: Situated question answering in 3d scenes
X. Ma, S. Yong, Z. Zheng, Q. Li, Y. Liang, S.-C. Zhu, and S. Huang · 2022
Cited alongside, same era.
Seihai: A sample-efficient hierarchical ai for the minerl competition
H. Mao, C. Wang, X. Hao, Y. Mao, Y. Lu, C. Wu, J. Hao, D. Li, and P. Tang · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
S. Cai, B. Zhang, Z. Wang, X. Ma, A. Liu, and Y. Liang · 2023
Closest in time.
Mindagent: Emergent gaming interaction
R. Gong, Q. Huang, X. Ma, H. Vo, Z. Durante, Y. Noda, Z. Zheng, S.-C. Zhu, D. Terzopoulos, L. Fei-Fei, et al · 2023
Closest in time.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Closest in time.
Steve-1: A generative model for text-to-behavior in minecraft
S. Lifshitz, K. Paster, H. Chan, J. Ba, and S. McIlraith · 2023
Closest in time.
Mcu: A task-centric framework for open-ended agent evaluation in minecraft
H. Lin, Z. Wang, J. Ma, and Y. Liang · 2023
Closest in time.
Text2motion: From natural language instructions to feasible plans
K. Lin, C. Agia, T. Migimatsu, M. Pavone, and J. Bohg · 2023
Closest in time.
Llm as a robotic brain: Unifying egocentric memory and control
J. Mai, J. Chen, B. Li, G. Qian, M. Elhoseiny, and B. Ghanem · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
N. Shinn, B. Labash, and A. Gopinath · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
Proagent: Building proactive cooperative ai with large language models
C. Zhang, K. Yang, S. Hu, Z. Wang, G. Li, Y. Sun, C. Zhang, Z. Zhang, A. Liu, S.-C. Zhu, et al · 2023
Closest in time.