Fetching the paper…
Reading the bibliography…
This paper introduces the Behaviour Suite for Reinforcement Learning, or bsuite for short.
Psychology and the process of group living
Kurt Lewin · 1943
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Jeannette Kiefer and Jacob Wolfowitz · 1952
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
Some studies oin machine learning using the game of checkers
AL Samuel · 1959
Earlier work this paper cites.
Steps towards artificial intelligence
Marvin Minsky · 1961
Earlier work this paper cites.
The group method of data of handling; a rival of the method of stochastic approximation
Alexey Grigorevich Ivakhnenko · 1968
Earlier work this paper cites.
The hippocampus as a spatial map: preliminary evidence from unit activity in the freely-moving rat
John O’Keefe and Jonathan Dostrovsky · 1971
Earlier work this paper cites.
Neural network model for a mechanism of pattern recognition unaffected by shift in position-neocognitron
Kunihiko Fukushima · 1979
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
John C Gittins · 1979
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R.S. Sutton · 1988
Earlier work this paper cites.
Efficient memory-based learning for robot control
Andrew William Moore · 1990
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Temporal difference learning and TD-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner, et al · 1998
Earlier work this paper cites.
Cognitive neuroscience and the study of memory
Brenda Milner, Larry R Squire, and Eric R Kandel · 1998
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
A collection of definitions of intelligence
Shane Legg, Marcus Hutter, et al · 2007
Earlier work this paper cites.
IPython: a system for interactive scientific computing
Fernando Pérez and Brian E. Granger · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Cited alongside, same era.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Cited alongside, same era.
Rl-glue: Language-independent software for reinforcement-learning experiments
Brian Tanner and Adam White · 2009
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Deep Reinforcement Learning with Double Q-Learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Later among the works it cites.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Later among the works it cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2017
Later among the works it cites.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi et al · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Massively Parallel Methods for Deep Reinforcement Learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, et al · 2015
Cited alongside, same era.
Later among the works it cites.
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2017
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Daniel Russo, Zheng Wen, and Benjamin Van Roy · 2017
Later among the works it cites.
A tutorial on Thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, and Ian Osband · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 2017
Later among the works it cites.
Reconciling modern machine learning and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Later among the works it cites.
Dopamine: A Research Framework for Deep Reinforcement Learning
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare · 2018
Later among the works it cites.
MIPLIB 2017, 2018
miplib2017 · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
Behaviour suite for reinforcement learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, , Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado Van Hasselt · 2019
Closest in time.
Openai five
Jakub Pachocki, David Farhi, Szymon Sidor, Greg Brockman, Filip Wolski, Henrique Pondé, Jie Tang, Jonathan Raiman, Michael Petrov, Christy Dennison, Brooke Chan, Susan Zhang, Rafał Józefowicz, and Przemysław Dębiak · 2019
Closest in time.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Jaderberg, et al · 2019
Closest in time.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Closest in time.