Fetching the paper…
Reading the bibliography…
We introduce PokerBench - a benchmark for evaluating the poker-playing abilities of large language models (LLMs).
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Games with incomplete information
Harsanyi, J. C. 1995 · 1995
Earlier work this paper cites.
Deep blue
Campbell, M.; Hoane Jr, A. J.; and Hsu, F.-h. 2002 · 2002
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2009
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. 2016 · 2016
Earlier work this paper cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N.; and Sandholm, T. 2018 · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Talmor, A.; Herzig, J.; Lourie, N.; and Berant, J. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Earlier work this paper cites.
Deep counterfactual regret minimization
Brown, N.; Lerer, A.; Gross, S.; and Sandholm, T. 2019 · 2019
Earlier work this paper cites.
Superhuman AI for multiplayer poker
Brown, N.; and Sandholm, T. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A.; Pruksachatkun, Y.; Nangia, N.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. 2019 · 2019
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Cited alongside, same era.
ChatGPT - https://openai.com/blog/chatgpt#OpenAI
OpenAI. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
GPT-4 Technical Report - https://cdn.openai.com/papers/gpt-4.pdf
OpenAI. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
A survey on large language model-based game agents
Hu, S.; Huang, T.; Ilhan, F.; Tekin, S.; Liu, G.; Kompella, R.; and Liu, L. 2024 · 2024
Later among the works it cites.
PokerGPT: An End-to-End Lightweight Solver for Multi-Player Texas Hold’em via Large Language Model
Huang, C.; Cao, Y.; Wen, Y.; Zhou, T.; and Zhang, Y. 2024 · 2024
Later among the works it cites.
Llama 3 - https://ai.meta.com/blog/meta-llama-3/
Meta. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Are ChatGPT and GPT-4 Good Poker Players?–A Pre-Flop Analysis
Gupta, A. 2023 · 2023
Cited alongside, same era.
Evaluating Large Language Models in Theory of Mind Tasks
Kosinski, M. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Grandmaster-level chess without search
Ruoss, A.; Delétang, G.; Medapati, S.; Grau-Moya, J.; Wenliang, L. K.; Catt, E.; Reid, J.; and Genewein, T. 2024 · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
Team, G.; Mesnard, T.; Hardin, C.; Dadashi, R.; Bhupatiraju, S.; Pathak, S.; Sifre, L.; Rivière, M.; Kale, M. S.; Love, J.; et al. 2024 · 2024
Later among the works it cites.
A Survey on Game Playing Agents and Large Models: Methods, Applications, and Challenges
Xu, X.; Wang, Y.; Xu, C.; Ding, Z.; Jiang, J.; Ding, Z.; and Karlsson, B. F. 2024 · 2024
Later among the works it cites.