Fetching the paper…
Reading the bibliography…
Generative models are trained with the simple objective of imitating the conditional probability distribution induced by the data they are trained on.
XXII. Programming a computer for playing chess
C. E. Shannon · 1941
Earlier work this paper cites.
Bagging predictors
L. Breiman · 1996
Earlier work this paper cites.
A short introduction to boosting, 1999
Y. Freund and R. E. Schapire · 1999
Earlier work this paper cites.
Deep Blue
M. Campbell, A. J. Hoane, and F.-h. Hsu · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. M. Kakade and J. Langford · 2002
Earlier work this paper cites.
Chess (1953)
A. Turing · 2004
Earlier work this paper cites.
Visualizing data using t-sne
L. Van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Example of the glicko-2 system
M. E. Glickman · 2012
Earlier work this paper cites.
Multi-agent team formation: Diversity beats strength?
L. S. Marcolino, A. X. Jiang, and M. Tambe · 2013
Earlier work this paper cites.
Diverse randomized agents vote to win
A. Jiang, L. Soriano Marcolino, A. D. Procaccia, T. Sandholm, N. Shah, and M. Tambe · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Give a Hard Problem to a Diverse Team: Exploring Large Action Spaces
L. Soriano Marcolino, H. Xu, A. Xin Jiang, M. Tambe, and E. Bowring · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Cited alongside, same era.
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm, Dec. 2017
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Efficiently Updatable Neural-Network-based Evaluation Functions for Computer Shogi, 2018
Y. Nasu · 2018
Cited alongside, same era.
AlphaZero Crushes Stockfish In New 1,000-Game Match
Decision transformer: Reinforcement learning via sequence modeling, 2021
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem, 2021
M. Janner, Q. Li, and S. Levine · 2021
Later among the works it cites.
Learning from disagreement: A survey
A. N. Uma, T. Fornaciari, D. Hovy, S. Paun, B. Plank, and M. Poesio · 2021
Later among the works it cites.
Chess as a Testbed for Language Model State Tracking
S. Toshniwal, S. Wiseman, K. Livescu, and K. Gimpel · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pete · 2018
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
Ensemble distillation for robust model fusion in federated learning
T. Lin, L. Kong, S. U. Stich, and M. Jaggi · 2020
Cited alongside, same era.
Aligning Superhuman AI with Human Behavior: Chess as a Model System
R. McIlroy-Young, S. Sen, J. Kleinberg, and A. Anderson · 2020
Cited alongside, same era.
The Chess Transformer: Mastering Play using Generative Language Models, Sept. 2020
D. Noever, M. Ciolino, and J. Kalin · 2020
Cited alongside, same era.
Stable policy optimization via off-policy divergence regularization
A. Touati, A. Zhang, J. Pineau, and P. Vincent · 2020
Cited alongside, same era.
Self-training with noisy student improves imagenet classification
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le · 2020
Cited alongside, same era.
C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, et al · 2023
Later among the works it cites.
ChessGPT: Bridging Policy Learning and Language Modeling
X. Feng, Y. Luo, Z. Wang, H. Tang, M. Yang, K. Shao, D. Mguni, Y. Du, and J. Wang · 2023
Later among the works it cites.
Offline reinforcement learning with closed-form policy improvement operators
J. Li, E. Zhang, M. Yin, Q. Bai, Y.-X. Wang, and W. Y. Wang · 2023
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data, 2023
H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2023
Later among the works it cites.
Emergent world models and latent variable estimation in chess-playing language models
A. Karvonen · 2024
Closest in time.
Leveraging ensemble diversity for robust self-training in the presence of sample selection bias, 2024
A. Odonnat, V. Feofanov, and I. Redko · 2024
Closest in time.
Stockfish, 2024
The Stockfish developers (see AUTHORS file) · 2024
Closest in time.