Fetching the paper…
Reading the bibliography…
We introduce Pok\'eChamp, a minimax agent powered by Large Language Models (LLMs) for Pok\'emon battles.
Deep blue
Campbell, M., Hoane Jr, A. J., and Hsu, F.-h · 2002
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2017
Earlier work this paper cites.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Earlier work this paper cites.
Superhuman ai for multiplayer poker
Brown, N. and Sandholm, T · 2019
Earlier work this paper cites.
A self-play policy optimization approach to battling pokémon
Huang, D. and Lee, S · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Earlier work this paper cites.
The nethack learning environment
Küttler, H., Nardelli, N., Miller, A., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Earlier work this paper cites.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
(FAIR)†, M. F. A. R. D. T., Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., Jacob, A. P., Komeili, M., Konath, K., Kwon, M., Lerer, A., Lewis, M., Miller, A. H., Mitts, S., Renduchintala, A., Roller, S., Rowe, D., Shi, W., Spisak, J., Wei, A., Wu, D., Zhang, H., and Zijlstra, M · 2022
Earlier work this paper cites.
How an a.i. is becoming the world’s best pokemon player, 2022
The-Third-Build · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Outracing champion gran turismo drivers with deep reinforcement learning
Wurman, P. R., Barrett, S., Kawamoto, K., MacGlashan, J., Subramanian, K., Walsh, T. J., Capobianco, R., Devlic, A., Eckert, F., Fuchs, F., et al · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Alphazero-like tree-search can guide large language model decoding and training
Large language models as commonsense knowledge for large-scale task planning
Zhao, Z., Lee, W. S., and Hsu, D · 2023
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Playing nethack with llms: Potential & limitations as zero-shot agents
Jeurissen, D., Perez-Liebana, D., Gow, J., Cakmak, D., and Kwan, J · 2024
Later among the works it cites.
Training language models to self-correct via reinforcement learning
Kumar, A., Zhuang, V., Agarwal, R., Su, Y., Co-Reyes, J. D., Singh, A., Baumli, K., Iqbal, S., Bishop, C., Roelofs, R., et al · 2024
Later among the works it cites.
Fightladder: A benchmark for competitive multi-agent reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Feng, X., Wan, Z., Wen, M., McAleer, S. M., Wen, Y., Zhang, W., and Wang, J · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z · 2023
Cited alongside, same era.
Motif: Intrinsic motivation from artificial intelligence feedback
Klissarov, M., D’Oro, P., Sodhani, S., Raileanu, R., Bacon, P.-L., Vincent, P., Zhang, A., and Henaff, M · 2023
Cited alongside, same era.
Large language models play starcraft ii: Benchmarks and a chain of summarization approach
Ma, W., Mi, Q., Yan, X., Wu, Y., Lin, R., Zhang, H., and Wang, J · 2023
Cited alongside, same era.
Cooperation on the fly: Exploring language agents for ad hoc teamwork in the avalon game
Shi, Z., Fang, M., Zheng, S., Deng, S., Chen, L., and Du, Y · 2023
Cited alongside, same era.
Stepputtis, S., Campbell, J., Xie, Y., Qi, Z., Zhang, W. S., Wang, R., Rangreji, S., Lewis, M., and Sycara, K · 2023
Cited alongside, same era.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Cited alongside, same era.
Training ai to play pokemon with reinforcement learning, 2023
Whidden, P · 2023
Cited alongside, same era.
Li, W., Ding, Z., Karten, S., and Jin, C · 2024
Later among the works it cites.
Pokéllmon trainer: Llm model distillation
Nalty, C. and Rosenthal, S · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
Benchmarking large language model (llm) performance for game playing via tic-tac-toe
Topsakal, O. and Harper, J. B · 2024
Later among the works it cites.
Spring: Studying papers and reasoning to play games
Wu, Y., Min, S. Y., Prabhumoye, S., Bisk, Y., Salakhutdinov, R. R., Azaria, A., Mitchell, T. M., and Li, Y · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Reflect-rl: Two-player online rl fine-tuning for lms
Zhou, R., Du, S. S., and Li, B · 2024
Later among the works it cites.