Fetching the paper…
Reading the bibliography…
Agents built with large language models (LLMs) have shown great potential across a wide range of domains.
Formalizing properties of agents
Goodwin, R · 1995
Earlier work this paper cites.
Intelligent agents: Theory and practice
Wooldridge, M. and Jennings, N. R · 1995
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 2007
Earlier work this paper cites.
Monte carlo sampling for regret minimization in extensive games
Lanctot, M., Waugh, K., Zinkevich, M., and Bowling, M · 2009
Earlier work this paper cites.
Solving large imperfect information games using cfr+
Tammelin, O · 2014
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Heinrich, J., Lanctot, M., and Silver, D · 2015
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Heinrich, J. and Silver, D · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T · 2017
Earlier work this paper cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Earlier work this paper cites.
Application of deep reinforcement learning in werewolf game agents
Wang, T. and Kaneko, T · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Superhuman ai for multiplayer poker
Brown, N. and Sandholm, T · 2019
Cited alongside, same era.
Deep counterfactual regret minimization
Brown, N., Lerer, A., Gross, S., and Sandholm, T · 2019
Cited alongside, same era.
A generalized training approach for multiagent learning
Muller, P., Omidshafiei, S., Rowland, M., Tuyls, K., Perolat, J., Liu, S., Hennes, D., Marris, L., Lanctot, M., Hughes, E., et al · 2019
Cited alongside, same era.
Finding friend and foe in multi-agent games
Serrino, J., Kleiman-Weiner, M., Parkes, D. C., and Tenenbaum, J · 2019
Cited alongside, same era.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Closest in time.
Strategic reasoning with language models
Gandhi, K., Sadigh, D., and Goodman, N. D · 2023
Closest in time.
Camel: Communicative agents for" mind" exploration of large scale language model society
Li, G., Hammoud, H. A. A. K., Itani, H., Khizbullin, D., and Ghanem, B · 2023
Closest in time.
Llm+ p: Empowering large language models with optimal planning proficiency
Liu, B., Jiang, Y., Zhang, X., Liu, Q., Zhang, S., Biswas, J., and Stone, P · 2023
Closest in time.
Large language models play starcraft ii: Benchmarks and a chain of summarization approach
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
Real world games look like spinning tops
Czarnecki, W. M., Gidel, G., Tracey, B., Tuyls, K., Omidshafiei, S., Balduzzi, D., and Jaderberg, M · 2020
Cited alongside, same era.
Neural replicator dynamics: Multiagent learning via hedging policy gradients
Hennes, D., Morrill, D., Omidshafiei, S., Munos, R., Perolat, J., Lanctot, M., Gruslys, A., Lespiau, J.-B., Parmas, P., Duéñez-Guzmán, E., et al · 2020
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Cited alongside, same era.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
Meta, Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Ma, W., Mi, Q., Yan, X., Wu, Y., Lin, R., Zhang, H., and Wang, J · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Closest in time.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Shah, D., Osiński, B., Levine, S., et al · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Shinn, N., Labash, B., and Gopinath, A · 2023
Closest in time.
Chatgpt for robotics: Design principles and model abilities
Vemprala, S., Bonatti, R., Bucker, A., and Kapoor, A · 2023
Closest in time.
Epidemic modeling with generative agents
Williams, R., Hosseinichimeh, N., Majumdar, A., and Ghaffarzadegan, N · 2023
Closest in time.
Fictitious cross-play: Learning global nash equilibrium in mixed cooperative-competitive games
Xu, Z., Liang, Y., Yu, C., Wang, Y., and Wu, Y · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Closest in time.
Gpt-4v (ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Closest in time.