Fetching the paper…
Reading the bibliography…
Monte Carlo Tree Search (MCTS) algorithms such as AlphaGo and MuZero have achieved superhuman performance in many challenging tasks.
State abstraction for programmable reinforcement learning agents
David Andre and Stuart J Russell · 2002
Earlier work this paper cites.
Approximate equivalence of markov decision processes
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Exploration exploitation in go: Uct for monte-carlo go
Sylvain Gelly and Yizao Wang · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Multi-armed bandits with episode context
Christopher D Rosin · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Earlier work this paper cites.
State aggregation in monte carlo tree search
Jesse Hostetler, Alan Fern, and Tom Dietterich · 2014
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Extreme state aggregation beyond markov decision processes
Marcus Hutter · 2016
Earlier work this paper cites.
Markovian state and action abstractions for mdps via hierarchical mcts
Aijun Bai, Siddharth Srivastava, and Stuart Russell · 2016
Earlier work this paper cites.
Near optimal behavior via approximate state abstraction
David Abel, David Hershkowitz, and Michael Littman · 2016
Earlier work this paper cites.
Oga-uct: On-the-go abstractions in uct
Ankit Anand, Ritesh Noothigattu, Parag Singla, et al · 2016
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Cited alongside, same era.
Toward good abstractions for lifelong learning
David Abel, Dilip Arumugam, Lucas Lehnert, and Michael L Littman · 2017
Cited alongside, same era.
State abstractions for lifelong reinforcement learning
David Abel, Dilip Arumugam, Lucas Lehnert, and Michael Littman · 2018
Cited alongside, same era.
Mastering atari games with limited data
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao · 2021
Later among the works it cites.
Research on action strategies and simulations of drl and mcts-based intelligent round game
Yuxiang Sun, Bo Yuan, Yongliang Zhang, Wanwen Zheng, Qingfeng Xia, Bojian Tang, and Xianzhong Zhou · 2021
Later among the works it cites.
Monte carlo tree search with iteratively refining state abstractions
Samuel Sokota, Caleb Y Ho, Zaheen Ahmad, and J Zico Kolter · 2021
Later among the works it cites.
Game state and action abstracting monte carlo tree search for general strategy game-playing
Alexander Dockhorn, Jorge Hurtado-Grueso, Dominik Jeurissen, Linjie Xu, and Diego Perez-Liebana · 2021
Later among the works it cites.
Online and offline reinforcement learning by planning with a learned model
Julian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain, Ioannis Antonoglou, and David Silver · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging { \{ AI } \} applications
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I Jordan, et al · 2018
Cited alongside, same era.
Performance guarantees for homomorphisms beyond markov decision processes
Sultan Javed Majeed and Marcus Hutter · 2019
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Cited alongside, same era.
Driving maneuvers prediction based autonomous driving control by deep monte carlo tree search
Jienan Chen, Cong Zhang, Jinting Luo, Junfei Xie, and Yan Wan · 2020
Cited alongside, same era.
Learning and planning in complex action spaces
Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain, Simon Schmitt, and David Silver · 2021
Later among the works it cites.
An mcts-based recommender system for education complex
Debin Zhao, Zhengyuan Hu, and Yinjian Yang · 2022
Later among the works it cites.
Policy improvement by planning with gumbel
Ivo Danihelka, Arthur Guez, Julian Schrittwieser, and David Silver · 2022
Later among the works it cites.
Are alphazero-like agents robust to adversarial perturbations?
Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai, I Wu, Cho-Jui Hsieh, et al · 2022
Later among the works it cites.
Spending thinking time wisely: Accelerating mcts with virtual expansions
Weirui Ye, Pieter Abbeel, and Yang Gao · 2022
Later among the works it cites.