Fetching the paper…
Reading the bibliography…
Recent advancements in Large Language Models(LLMs) have led to the development of LLM-based AI agents.
R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence , vol. 112, no. 1-2, pp. 181–211, 1999
1999
Earlier work this paper cites.
2012
Earlier work this paper cites.
S. Ontanón, “The combinatorial multi-armed bandit problem and its application to real-time strategy games,” in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , vol. 9, no. 1, 2013, pp. 58–64
2013
Earlier work this paper cites.
——, “Active learning of parameterized skills,” in International Conference on Machine Learning . PMLR, 2014, pp. 1737–1745
2014
Earlier work this paper cites.
A. Akbari, J. Rosell et al. , “Task planning using physics-based heuristics on manipulation actions,” in 2016 IEEE 21st International Conference on Emerging Technologies and Factory Automation (ETFA) . IEEE, 2016, pp. 1–8
2016
Earlier work this paper cites.
W. Masson, P. Ranchod, and G. Konidaris, “Reinforcement learning with parameterized actions,” in Proceedings of the AAAI conference on artificial intelligence , vol. 30, no. 1, 2016
2016
Earlier work this paper cites.
S. Krishnan, R. Fox, I. Stoica, and K. Goldberg, “Ddco: Discovery of deep continuous options for robot learning from demonstrations,” in Conference on robot learning . PMLR, 2017, pp. 418–437
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Ontañón, N. A. Barriga, C. R. Silva, R. O. Moraes, and L. H. Lelis, “The first microrts artificial intelligence competition,” AI Magazine , vol. 39, no. 1, pp. 75–83, 2018
2018
Earlier work this paper cites.
A. Amiranashvili, A. Dosovitskiy, V. Koltun, and T. Brox, “Motion perception in reinforcement learning with dynamic objects,” in Conference on Robot Learning . PMLR, 2018, pp. 156–168
2018
Earlier work this paper cites.
G. Konidaris, L. P. Kaelbling, and T. Lozano-Perez, “From skills to symbols: Learning symbolic representations for abstract high-level planning,” Journal of Artificial Intelligence Research , vol. 61, pp. 215–289, 2018
2018
Earlier work this paper cites.
B. Ames, A. Thackston, and G. Konidaris, “Learning symbolic representations for planning with parameterized skills,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 526–533
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, 2019, pp. 4171–4186. [Online]. Available: https://aclanthology.org/N19-1423
2019
Earlier work this paper cites.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al. , “Grandmaster level in starcraft ii using multi-agent reinforcement learning,” nature , vol. 575, no. 7782, pp. 350–354, 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 1877–1901. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
2020
Earlier work this paper cites.
2021
Cited alongside, same era.
S. Huang, S. Ontañón, C. Bamford, and L. Grela, “Gym- μ \mu rts: Toward affordable full game real-time strategy games research with deep reinforcement learning,” in 2021 IEEE Conference on Games (CoG) . IEEE, 2021, pp. 1–8
2021
Cited alongside, same era.
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch, “Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,” in Proceedings of the 39th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, Eds., vol. 162. PMLR, 17–23 Jul 2022, pp. 9118–9147. [Online]. Available: https://proceedings.mlr.press/v162/huang22a.html
2022
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
E. Brooks, L. Walls, R. L. Lewis, and S. Singh, “Large language models can implement policy iteration,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
brian ichter, A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, D. Kalashnikov, S. Levine, Y. Lu, C. Parada, K. Rao, P. Sermanet, A. T. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, M. Yan, N. Brown, M. Ahn, O. Cortes, N. Sievers, C. Tan, S. Xu, D. Reyes, J. Rettinghouse, J. Quiambao, P. Pastor, L. Luu, K.-H. Lee, Y. Kuang, S. Jesmonth, K. Jeffrey, R. J. Ruano, J. Hsu, K. Gopalakrishnan, B. David, A. Zeng, and C. K. Fu, “Do as i can, not as i say: Grounding language in robotic affordances,” in 6th Annual Conference on Robot Learning , 2022. [Online]. Available: https://openreview.net/forum?id=bdHkMjBJG_w
2022
Cited alongside, same era.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. Curran Associates, Inc., 2022, pp. 22 199–22 213. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/file/8bb0d291acd4acf06ef112099c16f326-Paper-Conference.pdf
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Imani, L. Du, and H. Shrivastava, “MathPrompter: Mathematical reasoning using large language models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track) , S. Sitaram, B. Beigman Klebanov, and J. D. Williams, Eds. Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 37–42. [Online]. Available: https://aclanthology.org/2023.acl-industry.4
2023
Cited alongside, same era.
Later among the works it cites.
A. Zhao, D. Huang, Q. Xu, M. Lin, Y.-J. Liu, and G. Huang, “Expel: Llm agents are experiential learners,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp. 19 632–19 642
2024
Later among the works it cites.
R. Hazra, P. Z. Dos Martires, and L. De Raedt, “Saycanpay: Heuristic planning with large language models using learnable domain knowledge,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 18, 2024, pp. 20 123–20 133
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Tang, A. Zou, Z. Zhang, Z. Li, Y. Zhao, X. Zhang, A. Cohan, and M. Gerstein, “MedAgents: Large language models as collaborators for zero-shot medical reasoning,” in Findings of the Association for Computational Linguistics ACL 2024 , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand and virtual meeting: Association for Computational Linguistics, Aug. 2024, pp. 599–621. [Online]. Available: https://aclanthology.org/2024.findings-acl.33
2024
Later among the works it cites.
B. Prakash, T. Oates, and T. Mohsenin, “Using llms for augmenting hierarchical agents with common sense priors,” The International FLAIRS Conference Proceedings , vol. 37, no. 1, May 2024. [Online]. Available: https://journals.flvc.org/FLAIRS/article/view/135602
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Zhou, B. Hu, C. Zhao, P. Zhang, and B. Liu, “Large language model as a policy teacher for training reinforcement learning agents,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24 , K. Larson, Ed. International Joint Conferences on Artificial Intelligence Organization, 8 2024, pp. 5671–5679, main Track. [Online]. Available: https://doi.org/10.24963/ijcai.2024/627
2024
Later among the works it cites.
M. Dalal, T. Chiruvolu, D. S. Chaplot, and R. Salakhutdinov, “Plan-seq-learn: Language model guided RL for solving long horizon robotics tasks,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=hQVCCxQrYN
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.