Fetching the paper…
Reading the bibliography…
Traditional search algorithms have issues when applied to games of imperfect information where the number of possible underlying states and trajectories are very large.
OpenSpiel: A Framework for Reinforcement Learning in Games
Lanctot, M.; Lockhart, E.; Lespiau, J.-B.; Zambaldi, V.; Upadhyay, S.; Pérolat, J.; Srinivasan, S.; Timbers, F.; Tuyls, K.; Omidshafiei, S.; Hennes, D.; Morrill, D.; Muller, P.; Ewalds, T.; Faulkner, R.; Kramár, J.; Vylder, B. D.; Saeta, B.; Bradbury, J.; Ding, D.; Borgeaud, S.; Lai, M.; Schrittwieser, J.; Anthony, T.; Hughes, E.; Danihelka, I.; and Ryan-Davis, J. 2019 · 1908
Earlier work this paper cites.
Huggingface’s Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2019 · 1910
Earlier work this paper cites.
Optimal Control of Markov Processes with Incomplete State Information I
Åström, K. J. 1965 · 1965
Earlier work this paper cites.
The Million Pound Bridge Program, Heuristic Programming in Artificial Intelligence: The First Computer Olympiad
Levy, D. 1989 · 1989
Earlier work this paper cites.
Integrated Architectures for Learning, Planning, and Reacting Based on Approximating Dynamic Programming
Sutton, R. S. 1990 · 1990
Earlier work this paper cites.
Search in Games with Incomplete Information: A Case Study using Bridge Card Play
Frank, I.; and Basin, D. 1998 · 1998
Earlier work this paper cites.
Planning and Acting in Partially Observable Stochastic Domains
Kaelbling, L. P.; Littman, M. L.; and Cassandra, A. R. 1998 · 1998
Earlier work this paper cites.
GIB: Imperfect Information in a Computationally Challenging Game
Ginsberg, M. L. 2001 · 2001
Earlier work this paper cites.
Deep Blue
Campbell, M.; Hoane Jr, A. J.; and Hsu, F.-h. 2002 · 2002
Earlier work this paper cites.
A Sparse Sampling Algorithm for Near-Optimal Planning in Large Markov Decision Processes
Kearns, M.; Mansour, Y.; and Ng, A. Y. 2002 · 2002
Earlier work this paper cites.
Non-Cooperative Games with Many Players
Khan, M. A.; and Sun, Y. 2002 · 2002
Earlier work this paper cites.
Gerhardt, G. 2004
2004
Earlier work this paper cites.
Current challenges in multi-player game search
Sturtevant, N. 2004 · 2004
Earlier work this paper cites.
Single-Agent Optimization Through Policy Iteration using Monte-Carlo Tree Search
Seify, A.; and Buro, M. 2020 · 2005
Earlier work this paper cites.
Bandit Based Monte-Carlo Planning
Kocsis, L.; and Szepesvári, C. 2006 · 2006
Cited alongside, same era.
Prob-maxˆ n: Playing N-player Games with Opponent Models
Sturtevant, N.; Zinkevich, M.; and Bowling, M. 2006 · 2006
Cited alongside, same era.
Feature Construction for Reinforcement Learning in Hearts
Sturtevant, N. R.; and White, A. M. 2007 · 2006
Cited alongside, same era.
Checkers is Solved
Schaeffer, J.; Burch, N.; Bjornsson, Y.; Kishimoto, A.; Muller, M.; Lake, R.; Lu, P.; and Sutphen, S. 2007 · 2007
Cited alongside, same era.
An analysis of UCT in multi-player games
Sturtevant, N. 2008 · 2008
Cited alongside, same era.
Improving State Evaluation, Inference, and Search in Trick-Based Card Games
Buro, M.; Long, J. R.; Furtak, T.; and Sturtevant, N. 2009 · 2009
Cited alongside, same era.
DeepStack: Expert-Level Artificial Intelligence in Heads-Up No-Limit Poker
Moravčík, M.; Schmid, M.; Burch, N.; Lisỳ, V.; Morrill, D.; Bard, N.; Davis, T.; Waugh, K.; Johanson, M.; and Bowling, M. 2017 · 2017
Later among the works it cites.
Surprising Negative Results for Generative Adversarial Tree Search
Azizzadenesheli, K.; Yang, B.; Liu, W.; Lipton, Z. C.; and Anandkumar, A. 2018 · 2018
Later among the works it cites.
Superhuman AI for Heads-up No-limit Poker: Libratus Beats Top Professionals
Brown, N.; and Sandholm, T. 2018 · 2018
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kopf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019 · 2019
Later among the works it cites.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding the Success of Perfect Information Monte Carlo Sampling in Game Tree Search
Long, J. R.; Sturtevant, N. R.; Buro, M.; and Furtak, T. 2010 · 2010
Cited alongside, same era.
Monte-Carlo Planning in Large POMDPs
Silver, D.; and Veness, J. 2010 · 2010
Cited alongside, same era.
Information Set Monte Carlo Tree Search
Cowling, P. I.; Powley, E. J.; and Whitehouse, D. 2012 · 2012
Cited alongside, same era.
A Survey on Policy Search for Robotics
Deisenroth, M. P.; Neumann, G.; Peters, J.; et al. 2013 · 2013
Cited alongside, same era.
Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
Heinrich, J.; and Silver, D. 2016 · 2016
Cited alongside, same era.
Mastering the Game of Go with Deep Neural Networks and Tree Search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Schrittwieser, J.; Antonoglou, I.; Hubert, T.; Simonyan, K.; Sifre, L.; Schmitt, S.; Guez, A.; Lockhart, E.; Hassabis, D.; Graepel, T.; et al. 2020 · 2020
Later among the works it cites.
The Crew: The Quest for Planet Nine
Sing, T. 2020 · 2020
Later among the works it cites.
Planning in Stochastic Environments with a Learned Model
Antonoglou, I.; Schrittwieser, J.; Ozair, S.; Hubert, T. K.; and Silver, D. 2021 · 2021
Later among the works it cites.
Decision transformer: Reinforcement Learning via Sequence Modeling
Chen, L.; Lu, K.; Rajeswaran, A.; Lee, K.; Grover, A.; Laskin, M.; Abbeel, P.; Srinivas, A.; and Mordatch, I. 2021 · 2021
Later among the works it cites.
Offline reinforcement Learning as one Big Sequence Modeling Problem
Janner, M.; Li, Q.; and Levine, S. 2021 · 2021
Later among the works it cites.
Vector Quantized Models for Planning
Ozair, S.; Li, Y.; Razavi, A.; Antonoglou, I.; Van Den Oord, A.; and Vinyals, O. 2021 · 2021
Later among the works it cites.
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
Li, K.; Hopkins, A. K.; Bau, D.; Viégas, F.; Pfister, H.; and Wattenberg, M. 2022 · 2022
Later among the works it cites.
MCTransformer: Combining Transformers and Monte-Carlo Tree Search for Offline Reinforcement Learning
Yaari, G.; Rokach, L.; Puzis, R.; and Katz, G. 2022 · 2022
Later among the works it cites.
History Filtering in Imperfect Information Games: Algorithms and Complexity
Solinas, C.; Rebstock, D.; Sturtevant, N. R.; and Buro, M. 2023 · 2023
Later among the works it cites.