Fetching the paper…
Reading the bibliography…
We propose Pgx, a suite of board game reinforcement learning (RL) environments written in JAX and optimized for GPU/TPU accelerators.
A simplified two-person poker
Harold W Kuhn · 1950
Earlier work this paper cites.
A Knowledge-Based Approach of Connect-Four
Victor Allis · 1988
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
The game of Go
John Tromp · 1995
Earlier work this paper cites.
Bayes’ Bluff: Opponent Modelling in Poker
Finnegan Southey, Michael P Bowling, Bryce Larson, Carmelo Piccione, Neil Burch, Darse Billings, and Chris Rayner · 2005
Earlier work this paper cites.
Whole-History Rating: A Bayesian Rating System for Players of Time-Varying Strength
Rémi Coulom · 2008
Earlier work this paper cites.
An Analysis of a Board Game “Doubutsu Shogi”
Tetsuro Tanaka · 2009
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
PACHI: State of the Art Open Source Go Program
Petr Baudiš and Jean-loup Gailly · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M G Bellemare, Y Naddaf, J Veness, and M Bowling · 2013
Earlier work this paper cites.
Gardner’s Minichess Variant is Solved
Mehdi Mhalla and Frédéric Prost · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Identity Mappings in Deep Residual Networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Earlier work this paper cites.
Thinking Fast and Slow with Deep Learning and Tree Search
Thomas Anthony, Zheng Tian, and David Barber · 2017
Earlier work this paper cites.
dlshogi
Tadao Yamaoka · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
OpenSpiel: A Framework for Reinforcement Learning in Games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Paul Muller, Timo Ewalds, Ryan Faulkner, János Kramár, Bart De Vylder, Brennan Saeta, James Bradbury, David Ding, Sebastian Borgeaud, Matthew Lai, Julian Schrittwieser, Thomas Anthony, Edward Hughes, Ivo Danihelka, and Jonah Ryan-Davis · 2019
Brax - A Differentiable Physics Engine for Large Scale Rigid Body Simulation
Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem · 2021
Later among the works it cites.
Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, and Gavriel State · 2021
Later among the works it cites.
PettingZoo: Gym for Multi-Agent Reinforcement Learning
Justin K Terry, Benjamin Black, Mario Jayakumar, Ananth Hari, Luis Santos, Clemens Dieffendahl, Niall L Williams, Yashas Lokesh, Ryan Sullivan, Caroline Horsch, and Praveen Ravi · 2021
Later among the works it cites.
Podracer architectures for scalable Reinforcement Learning
Matteo Hessel, Manuel Kroiss, Aidan Clark, Iurii Kemaev, John Quan, Thomas Keck, Fabio Viola, and Hado van Hasselt · 2021
Later among the works it cites.
Revisiting Rainbow: Promoting more Insightful and Inclusive Deep Reinforcement Learning Research
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments
Kenny Young and Tian Tian · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Competitive Bridge Bidding with Deep Neural Networks
Jiang Rong, Tao Qin, and Bo An · 2019
Cited alongside, same era.
Mastering Atari, Go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Cited alongside, same era.
Suphx: Mastering Mahjong with Deep Reinforcement Learning
Junjie Li, Sotetsu Koyamada, Qiwei Ye, Guoqing Liu, Chao Wang, Ruihan Yang, Li Zhao, Tao Qin, Tie-Yan Liu, and Hsiao-Wuen Hon · 2020
Cited alongside, same era.
RLCard: A Platform for Reinforcement Learning in Card Games
Daochen Zha, Kwei-Herng Lai, Songyi Huang, Yuanpu Cao, Keerthana Reddy, Juan Vargas, Alex Nguyen, Ruzhe Wei, Junyu Guo, and Xia Hu · 2020
Cited alongside, same era.
Behaviour Suite for Reinforcement Learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado Van Hasselt · 2020
Cited alongside, same era.
Johan Samir Obando Ceron and Pablo Samuel Castro · 2021
Later among the works it cites.
Spectral Normalisation for Deep Reinforcement Learning: An Optimisation Perspective
Florin Gogianu, Tudor Berariu, Mihaela C Rosca, Claudia Clopath, Lucian Busoniu, and Razvan Pascanu · 2021
Later among the works it cites.
Policy improvement by planning with Gumbel
Ivo Danihelka, Arthur Guez, Julian Schrittwieser, and David Silver · 2022
Later among the works it cites.
EnvPool: A Highly Parallel Reinforcement Learning Environment Execution Engine
Jiayi Weng, Min Lin, Shengyi Huang, Bo Liu, Denys Makoviichuk, Viktor Makoviychuk, Zichen Liu, Yufan Song, Ting Luo, Yukun Jiang, et al · 2022
Later among the works it cites.
gymnax: A JAX-based Reinforcement Learning Environment Library
Robert Tjarko Lange · 2022
Later among the works it cites.
WarpDrive: Fast End-to-End Deep Multi-Agent Reinforcement Learning on a GPU
Tian Lan, Sunil Srinivasa, Huan Wang, and Stephan Zheng · 2022
Later among the works it cites.
Planning in Stochastic Environments with a Learned Model
Ioannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K Hubert, and David Silver · 2022
Later among the works it cites.
You Can’t Count on Luck: Why Decision Transformers and RvS Fail in Stochastic Environments
Keiran Paster, Sheila McIlraith, and Jimmy Ba · 2022
Later among the works it cites.
Efficient Learning for AlphaZero via Path Consistency
Dengwei Zhao, Shikui Tu, and Lei Xu · 2022
Later among the works it cites.
Tianshou: A Highly Modularized Deep Reinforcement Learning Library
Jiayi Weng, Huayu Chen, Dong Yan, Kaichao You, Alexis Duburcq, Minghao Zhang, Yi Su, Hang Su, and Jun Zhu · 2022
Later among the works it cites.
Discovered Policy Optimisation
Chris Lu, Jakub Kuba, Alistair Letcher, Luke Metz, Christian Schroeder de Witt, and Jakob Foerster · 2022
Later among the works it cites.
Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX
Clément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana, Vincent Coyette, Paul Duckworth, Laurence I. Midgley, Tristan Kalloniatis, Sasha Abramowitz, Cemlyn N. Waters, Andries P. Smit, Nathan Grinsztajn, Ulrich A. Mbou Sob, Omayma Mahjoub, Elshadai Tegegn, Mohamed A. Mimouni, Raphael Boige, Ruan de Kock, Daniel Furelos-Blanco, Victor Le, Arnu Pretorius, and Alexandre Laterre · 2023
Closest in time.