Fetching the paper…
Reading the bibliography…
Learning skills in open-world environments is essential for developing agents capable of handling a variety of tasks by combining basic skills.
A new algorithm for data compression
P. Gage · 1994
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Event perception: a mind-brain perspective
J. M. Zacks, N. K. Speer, K. M. Swallow, T. S. Braver, and J. R. Reynolds · 2007
Earlier work this paper cites.
Episodic memories
M. A. Conway · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
K. Sohn, H. Lee, and X. Yan · 2015
Earlier work this paper cites.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Earlier work this paper cites.
Taco: Learning task decomposition via temporal alignment for control
K. Shiarlis, M. Wulfmeier, S. Salter, S. Whiteson, and I. Posner · 2018
Earlier work this paper cites.
Option discovery using deep skill chaining
A. Bagaria and G. Konidaris · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov · 2019
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2019
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
W. H. Guss, B. Houghton, N. Topin, P. Wang, C. Codel, M. M. Veloso, and R. Salakhutdinov · 2019
Earlier work this paper cites.
CompILE: Compositional imitation learning and execution
T. Kipf, Y. Li, H. Dai, V. Zambaldi, A. Sanchez-Gonzalez, E. Grefenstette, P. Kohli, and P. Battaglia · 2019
Earlier work this paper cites.
Learning robot skills with temporal variational inference
T. Shankar and A. Gupta · 2020
Earlier work this paper cites.
Vtnet: Visual transformer network for object goal navigation
H. Du, X. Yu, and L. Zheng · 2021
Cited alongside, same era.
Simple but effective: Clip embeddings for embodied ai
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi · 2021
Cited alongside, same era.
Flexible option learning
M. Klissarov and D. Precup · 2021
Cited alongside, same era.
Skid raw: Skill discovery from raw trajectories
D. Tanneberg, K. Ploeger, E. Rueckert, and J. Peters · 2021
Cited alongside, same era.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Towards evaluating generalist agents: An automated benchmark in open world
X. Zheng, H. Lin, K. He, Z. Wang, Z. Zheng, and Y. Liang · 2023
Later among the works it cites.
X. Zhu, Y. Chen, H. Tian, C. Tao, W. Su, C. Yang, G. Huang, B. Li, L. Lu, X. Wang, Y. Qiao, Z. Zhang, and J. Dai · 2023
Later among the works it cites.
Rt-h: Action hierarchies using language
S. Belkhale, T. Ding, T. Xiao, P. Sermanet, Q. Vuong, J. Tompson, Y. Chebotar, D. Dwibedi, and D. Sadigh · 2024
Later among the works it cites.
Rocket-1: Master open-world interaction with visual-temporal context prompting
S. Cai, Z. Wang, K. Lian, Z. Mu, X. Ma, A. Liu, and Y. Liang · 2024
Later among the works it cites.
Exploring large language model based intelligent agents: Definitions, methods, and prospects
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Alma: Hierarchical learning for composite multi-agent tasks
S. Iqbal, R. Costales, and F. Sha · 2022
Cited alongside, same era.
Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation
Y. Zhu, P. Stone, and Y. Zhu · 2022
Cited alongside, same era.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Cited alongside, same era.
Steve-1: A generative model for text-to-behavior in minecraft
S. Lifshitz, K. Paster, H. Chan, J. Ba, and S. McIlraith · 2023
Cited alongside, same era.
Mcu: A task-centric framework for open-ended agent evaluation in minecraft
H. Lin, Z. Wang, J. Ma, and Y. Liang · 2023
Cited alongside, same era.
Future-conditioned unsupervised pretraining for decision transformer
Z. Xie, Z. Lin, D. Ye, Q. Fu, W. Yang, and S. Li · 2023
Cited alongside, same era.
Y. Cheng, C. Zhang, Z. Zhang, X. Meng, S. Hong, W. Li, Z. Wang, Z. Wang, F. Yin, J. Zhao, et al · 2024
Later among the works it cites.
Luban: Building open-ended creative agents via autonomous embodied verification
Y. Guo, S. Peng, J. Guo, D. Huang, X. Zhang, R. Zhang, Y. Hao, L. Li, Z. Tian, M. Gao, Y. Li, Y. Gan, S. Liang, Z. Zhang, Z. Du, Q. Guo, X. Hu, and Y. Chen · 2024
Later among the works it cites.
Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks
Z. Li, Y. Xie, R. Shao, G. Chen, D. Jiang, and L. Nie · 2024
Later among the works it cites.
Mp5: A multi-modal open-ended embodied system in minecraft via active perception
Y. Qin, E. Zhou, Q. Liu, Z. Yin, L. Sheng, R. Zhang, Y. Qiao, and J. Shao · 2024
Later among the works it cites.
Pre-training goal-based models for sample-efficient reinforcement learning
H. Yuan, Z. Mu, F. Xie, and Z. Lu · 2024
Later among the works it cites.
Proagent: building proactive cooperative agents with large language models
C. Zhang, K. Yang, S. Hu, Z. Wang, G. Li, Y. Sun, C. Zhang, Z. Zhang, A. Liu, S.-C. Zhu, et al · 2024
Later among the works it cites.
Minedreamer: Learning to follow instructions via chain-of-imagination for simulated-world control
E. Zhou, Y. Qin, Z. Yin, Y. Huang, R. Zhang, L. Sheng, Y. Qiao, and J. Shao · 2024
Later among the works it cites.
Reinforcement learning friendly vision-language model for minecraft
H. Jiang, J. Yue, H. Luo, Z. Ding, and Z. Lu · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Omnijarvis: Unified vision-language-action tokenization enables open-world instruction following agents
Z. Wang, S. Cai, Z. Mu, H. Lin, C. Zhang, X. Liu, Q. Li, A. Liu, X. S. Ma, and Y. Liang · 2025
Closest in time.