Fetching the paper…
Reading the bibliography…
Adversarial self-play in two-player games has delivered impressive results when used with reinforcement learning algorithms that combine deep neural networks and tree search.
TD-gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
Heuristics from nature for hard combinatorial optimization problems
Alberto Colorni, Marco Dorigo, Francesco Maffioli, Vittorio Maniezzo, Giovanni Righini, and Marco Trubian · 1996
Earlier work this paper cites.
Heuristics for the 0–1 multidimensional knapsack problem
V. Boyer, M. Elkihel, and D. El Baz · 2009
Earlier work this paper cites.
Traveling salesman problem heuristics: Leading methods, implementations and latest advances
César Rego, Dorabela Gamboa, Fred Glover, and Colin Osterman · 2011
Earlier work this paper cites.
A survey of Monte Carlo tree search methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Cited alongside, same era.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly · 2015
Cited alongside, same era.
Neural combinatorial optimization with reinforcement learning
Irwan Bello, Hieu Pham, Quoc V. Le, Mohammad Norouzi, and Samy Bengio · 2016
Cited alongside, same era.
Thinking fast and slow with deep learning and tree search
Thomas Anthony, Zheng Tian, and David Barber · 2017
Cited alongside, same era.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2017
Cited alongside, same era.
Solving a new 3D bin packing problem with deep reinforcement learning method
Haoyuan Hu, Xiaodong Zhang, Xiaowei Yan, Longfei Wang, and Yinghui Xu · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy P. Lillicrap, Karen Simonyan, and Demis Hassabis · 2017
Later among the works it cites.
Gurobi optimizer reference manual, 2018
LLC Gurobi Optimization · 2018
Closest in time.
A multi-task selected learning approach for solving new type 3D bin packing problem
Haoyuan Hu, Lu Duan, Xiaodong Zhang, Yinghui Xu, and Jiangwen Wei · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thomas M. Moerland, Joost Broekens, Aske Plaat, and Catholijn M. Jonker · 2018
Closest in time.