Fetching the paper…
Reading the bibliography…
Reinforcement learning defines the problem facing agents that learn to make good decisions through action and observation alone.
A mathematical theory of communication
Claude E. Shannon · 1948
Earlier work this paper cites.
Some informational aspects of visual perception
Fred Attneave · 1954
Earlier work this paper cites.
Dynamic programming and Lagrange multipliers
Richard Bellman · 1956
Earlier work this paper cites.
A Markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Empirical explorations of the logic theory machine: a case study in heuristic
Allen Newell, John Clark Shaw, and Herbert A. Simon · 1957
Earlier work this paper cites.
Models of man; social and rational
Herbert A. Simon · 1957
Earlier work this paper cites.
Possible principles underlying the transformation of sensory messages
Horace B. Barlow · 1961
Earlier work this paper cites.
Algebraic structure theory of sequential machines
Juris Hartmanis and Richard E. Stearns · 1966
Earlier work this paper cites.
Über ein paradoxon aus der verkehrsplanung
Dietrich Braess · 1968
Earlier work this paper cites.
Rate distortion theory: A mathematical basis for data compression
Toby Berger · 1971
Earlier work this paper cites.
STRIPS: A new approach to the application of theorem proving to problem solving
Richard E. Fikes and Nils J. Nilsson · 1971
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vladimir N. Vapnik and Aleksei Y. Chervonenkis · 1971
Earlier work this paper cites.
An algorithm for computing the capacity of arbitrary discrete memoryless channels
Suguru Arimoto · 1972
Earlier work this paper cites.
Computation of channel capacity and rate-distortion functions
Richard Blahut · 1972
Earlier work this paper cites.
Theories of bounded rationality
Herbert A. Simon · 1972
Earlier work this paper cites.
Algebraic connectivity of graphs
Miroslav Fiedler · 1973
Earlier work this paper cites.
Discretizing dynamic programs
Bennett L. Fox · 1973
Earlier work this paper cites.
Planning in a hierarchy of abstraction spaces
Earl D. Sacerdoti · 1974
Earlier work this paper cites.
Science and statistics
George E.P. Box · 1976
Earlier work this paper cites.
A set of measures of centrality based on betweenness
Linton C. Freeman · 1977
Earlier work this paper cites.
An algorithm for finding best matches in logarithmic expected time
Jerome H. Friedman, Jon Louis Bentley, and Raphael Ari Finkel · 1977
Earlier work this paper cites.
Estimating the dimension of a model
Gideon Schwarz · 1978
Earlier work this paper cites.
Approximations of dynamic programs, i
Ward Whitt · 1978
Earlier work this paper cites.
A greedy heuristic for the set-covering problem
Vasek Chvatal · 1979
Earlier work this paper cites.
Approximations of dynamic programs, ii
Ward Whitt · 1979
Earlier work this paper cites.
State aggregation in dynamic programming — an application to scheduling of independent jobs on parallel processors
Sven Axsäter · 1983
Earlier work this paper cites.
Learning to solve problems by searching for macro-operators
Richard E. Korf · 1983
Earlier work this paper cites.
A theory of the learnable
Leslie G. Valiant · 1984
Earlier work this paper cites.
Macro-operators: A weak method for learning
Richard E. Korf · 1985
Earlier work this paper cites.
The complexity of Markov decision processes
Christos H. Papadimitriou and John N. Tsitsiklis · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Adaptive aggregation methods for infinite horizon dynamic programming
Dimitri P. Bertsekas and David A. Castanon · 1989
Earlier work this paper cites.
Bounds on the cover time
Andrei Z. Broder and Anna R. Karlin · 1989
Earlier work this paper cites.
A heuristic approach to the discovery of macro-operators
Glenn A. Iba · 1989
Earlier work this paper cites.
Markov and Markov reward model transient analysis: An overview of numerical approaches
Andrew Reibman, Roger Smith, and Kishor Trivedi · 1989
Earlier work this paper cites.
Solving H {H} -horizon, stationary Markov decision problems in time proportional to log ( H ) \log({H})
Paul Tseng · 1990
Earlier work this paper cites.
Complexity results for planning
Tom Bylander · 1991
Earlier work this paper cites.
Input generalization in delayed reinforcement learning: An algorithm and performance comparisons
David Chapman and Leslie Pack Kaelbling · 1991
Earlier work this paper cites.
O-plan: the open planning architecture
Ken Currie and Austin Tate · 1991
Earlier work this paper cites.
Bisimulation through probabilistic testing
Kim G. Larsen and Arne Skou · 1991
Earlier work this paper cites.
Reinforcement learning in Markovian and non-Markovian environments
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
A theory of abstraction
Fausto Giunchiglia and Toby Walsh · 1992
Earlier work this paper cites.
Adaptive state space quantisation for reinforcement learning of collision-free navigation
Ben J.A. Krose and Joris W.M. Van Dam · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Scaling reinforcement learning algorithms by learning variable temporal resolution models
Satinder Singh · 1992
Earlier work this paper cites.
Introduction: The challenge of reinforcement learning
Richard S. Sutton · 1992
Earlier work this paper cites.
Q Q -learning
Christopher J.C.H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E. Hinton · 1993
Earlier work this paper cites.
Hierarchical reinforcement learning: Preliminary results
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Tight performance bounds on greedy policies based on imperfect value functions
Ronald J. Williams and Leemon C. Baird · 1993
Earlier work this paper cites.
The computational complexity of propositional STRIPS planning
Tom Bylander · 1994
Earlier work this paper cites.
HTN planning: Complexity and expressivity
Kutluhan Erol, James Hendler, and Dana S. Nau · 1994
Earlier work this paper cites.
Automatically generating abstractions for planning
Craig A. Knoblock · 1994
Earlier work this paper cites.
The parti-game algorithm for variable resolution reinforcement learning in multidimensional state-spaces
Andrew W. Moore · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Gavin A. Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Reversible Markov chains and random walks on graphs
David Aldous and James Fill · 1995
Earlier work this paper cites.
On the complexity of solving Markov decision problems
Michael L. Littman, Thomas L. Dean, and Leslie Pack Kaelbling · 1995
Earlier work this paper cites.
Reinforcement learning with selective perception and hidden state
Andrew McCallum · 1995
Earlier work this paper cites.
Provably Bounded-Optimal Agents
Stuart Russell and Devika Subramanian · 1995
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder Singh, Tommi Jaakkola, and Michael I. Jordan · 1995
Earlier work this paper cites.
Finding structure in reinforcement learning
Sebastian Thrun and Anton Schwartz · 1995
Earlier work this paper cites.
Neuro-dynamic programming , volume 5
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Spectral graph theory
Fan R.K. Chung · 1996
Earlier work this paper cites.
Active learning with statistical models
David A. Cohn, Zoubin Ghahramani, and Michael I. Jordan · 1996
Earlier work this paper cites.
Reasoning the fast and frugal way: models of bounded rationality
Gerd Gigerenzer and Daniel G. Goldstein · 1996
Earlier work this paper cites.
Chattering in SARSA( λ \lambda )
Geoffrey J. Gordon · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L. Littman, and Andrew W. Moore · 1996
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Michael L. Littman and Csaba Szepesvári · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S. Sutton · 1996
Earlier work this paper cites.
Is learning the n n -th thing any easier than learning the first?
Sebastian Thrun · 1996
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
David H. Wolpert · 1996
Earlier work this paper cites.
Robot learning from demonstration
Christopher G. Atkeson and Stefan Schaal · 1997
Earlier work this paper cites.
Model minimization in Markov decision processes
Thomas Dean and Robert Givan · 1997
Earlier work this paper cites.
Abstraction and approximate decision-theoretic planning
Richard Dearden and Craig Boutilier · 1997
Earlier work this paper cites.
Bounded parameter Markov decision processes
Robert Givan, Sonia Leach, and Thomas Dean · 1997
Earlier work this paper cites.
Roles of macro-actions in accelerating reinforcement learning
Amy McGovern, Richard S. Sutton, and Andrew H. Fagg · 1997
Earlier work this paper cites.
Multi-time models for reinforcement learning
Doina Precup and Richard S. Sutton · 1997
Earlier work this paper cites.
A neural substrate of prediction and reward
Wolfram Schultz, Peter Dayan, and P. Read Montague · 1997
Earlier work this paper cites.
HQ-learning
Marco Wiering and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Learning hierarchical control structures for multiple tasks and changing environments
Bruce L. Digney · 1998
Earlier work this paper cites.
Hierarchical solution of Markov decision processes using macro-actions
Milos Hauskrecht, Nicolas Meuleau, Leslie Pack Kaelbling, Thomas Dean, and Craig Boutilier · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
acQuire-macros: An algorithm for automatically learning macro-actions
Amy McGovern · 1998
Earlier work this paper cites.
An O ( log ∗ n ) {O}(\log^{*}n) approximation algorithm for the asymmetric p p -center problem
Rina Panigrahy and Sundar Vishwanathan · 1998
Earlier work this paper cites.
Hierarchical Control and Learning for Markov Decision Processes
Ronald Parr · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart Russell · 1998
Earlier work this paper cites.
Multi-time models for temporally abstract planning
Doina Precup and Richard S. Sutton · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Tree based discretization for continuous state space reinforcement learning
William T.B. Uther and Manuela M. Veloso · 1998
Earlier work this paper cites.
Sequential composition of dynamically dexterous robot behaviors
Robert R. Burridge, Alfred A. Rizzi, and Daniel E. Koditschek · 1999
Earlier work this paper cites.
Computing factored value functions for policies in structured MDPs
Daphne Koller and Ronald Parr · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C. Pereira, and William Bialek · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng and Stuart Russell · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Two O ( log ∗ k ) {O}(\log^{*}k) -approximation algorithms for the asymmetric k k -center problem
Aaron Archer · 2001
Earlier work this paper cites.
Max-norm projections for factored MDPs
Carlos Guestrin, Daphne Koller, and Ronald Parr · 2001
Earlier work this paper cites.
Automated state abstraction for options using the U-tree algorithm
Anders Jonsson and Andrew G. Barto · 2001
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy McGovern and Andrew G. Barto · 2001
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Doina Precup · 2001
Earlier work this paper cites.
State abstraction for programmable reinforcement learning agents
David Andre and Stuart Russell · 2002
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Reward, motivation, and reinforcement learning
Peter Dayan and Bernard W. Balleine · 2002
Earlier work this paper cites.
Discovering hierarchy in reinforcement learning with HEXQ
Bernhard Hengst · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Q-cut - dynamic discovery of sub-goals in reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2002
Cited alongside, same era.
PolicyBlocks: An algorithm for creating useful macro-actions in reinforcement learning
Marc Pickett and Andrew G. Barto · 2002
Cited alongside, same era.
Model minimization in hierarchical reinforcement learning
Balaraman Ravindran and Andrew G. Barto · 2002
Cited alongside, same era.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Cited alongside, same era.
Recent advances in hierarchical reinforcement learning
Andrew G. Barto and Sridhar Mahadevan · 2003
Cited alongside, same era.
Approximate equivalence of Markov decision processes
Eyal Even-Dar and Yishay Mansour · 2003
Cited alongside, same era.
Time-regularized interrupting options
Daniel J. Mankowitz, Timothy A. Mann, and Shie Mannor · 2014
Later among the works it cites.
Scaling up approximate value iteration with options: Better policies with fewer iterations
Timothy A. Mann and Shie Mannor · 2014
Later among the works it cites.
Selecting near-optimal approximate state representations in reinforcement learning
Ronald Ortner, Odalric-Ambrym Maillard, and Daniil Ryabko · 2014
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L. Puterman · 2014
Later among the works it cites.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Later among the works it cites.
Optimal behavioral hierarchy
Alec Solway, Carlos Diuk, Natalia Córdova, Debbie Yee, Andrew G. Barto, Yael Niv, and Matthew Botvinick · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the sample complexity of reinforcement learning
Sham Kakade · 2003
Cited alongside, same era.
BREVE: a 3D environment for the simulation of decentralized systems and artificial life
Jon Klein · 2003
Cited alongside, same era.
On spectral graph drawing
Yehuda Koren · 2003
Cited alongside, same era.
Least-squares policy iteration
Michail G. Lagoudakis and Ronald Parr · 2003
Cited alongside, same era.
Optimal mutual information quantization is NP-complete
Brendan Mumey and Tomáš Gedeon · 2003
Cited alongside, same era.
SMDP homomorphisms: An algebraic approach to abstraction in semi Markov decision processes
Balaraman Ravindran · 2003
Cited alongside, same era.
Later among the works it cites.
Efficient abstraction selection in reinforcement learning
Harm van Seijen, Shimon Whiteson, and Leon Kester · 2014
Later among the works it cites.
ASAP-UCT: abstraction of state-action pairs in UCT
Ankit Anand, Aditya Grover, Mausam, and Parag Singla · 2015
Later among the works it cites.
Strengths, weaknesses, and combinations of model-based and model-free reinforcement learning
Kavosh Asadi · 2015
Later among the works it cites.
Reinforcement learning, efficient coding, and the statistics of natural tasks
Matthew Botvinick, Ari Weinstein, Alec Solway, and Andrew G. Barto · 2015
Later among the works it cites.
The online coupon-collector problem and its application to lifelong reinforcement learning
Emma Brunskill and Lihong Li · 2015
Later among the works it cites.
Value iteration with options and state aggregation
Kamil Ciosek and David Silver · 2015
Later among the works it cites.
Computational rationality: A converging paradigm for intelligence in brains, minds, and machines
Samuel J. Gershman, Eric J. Horvitz, and Joshua Tenenbaum · 2015
Later among the works it cites.
Learning state representations with robotic priors
Rico Jonschkowski and Oliver Brock · 2015
Later among the works it cites.
Symbol acquisition for probabilistic high-level planning
George Konidaris, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Approximate value iteration with temporally extended actions
Timothy A. Mann, Shie Mannor, and Doina Precup · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Nonparametric Bayesian reward segmentation for skill discovery using inverse reinforcement learning
Pravesh Ranchod, Benjamin Rosman, and George Konidaris · 2015
Later among the works it cites.
Efficient approximation of channel capacities
Tobias Sutter, David Sutter, Peyman Mohajerin Esfahani, and John Lygeros · 2015
Later among the works it cites.
Portable option discovery for automated learning transfer in object-oriented Markov decision processes
Nicholay Topin, Nicholas Haltmeyer, Shawn Squire, John Winder, Marie desJardins, and James MacGlashan · 2015
Later among the works it cites.
A deeper look at planning as learning from replay
Harm van Seijen and Richard S. Sutton · 2015
Later among the works it cites.
8-month-old infants spontaneously learn and generalize hierarchical rules
Denise M. Werchan, Anne G.E. Collins, Michael J. Frank, and Dima Amso · 2015
Later among the works it cites.
Near optimal behavior via approximate state abstraction
David Abel, D. Ellis Hershkowitz, and Michael L. Littman · 2016
Later among the works it cites.
OGA-UCT: on-the-go abstractions in UCT
Ankit Anand, Ritesh Noothigattu, Mausam, and Parag Singla · 2016
Later among the works it cites.
Markovian State and Action Abstractions for MDPs via Hierarchical MCTS
Aijun Bai and Stuart Russell · 2016
Later among the works it cites.
OpenAI gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Extreme state aggregation beyond Markov decision processes
Marcus Hutter · 2016
Later among the works it cites.
Using task features for zero-shot knowledge transfer in lifelong learning
David Isele, Mohammad Rostami, and Eric Eaton · 2016
Later among the works it cites.
Constructing abstraction hierarchies using a skill-symbol loop
George Konidaris · 2016
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D. Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Later among the works it cites.
Nonparametric General Reinforcement Learning
Jan Leike · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
State of the art control of atari games using shallow reinforcement learning
Yitao Liang, Marlos C. Machado, Erik Talvitie, and Michael Bowling · 2016
Later among the works it cites.
Efficient Bayesian clustering for reinforcement learning
Travis Mandel, Yun-En Liu, Emma Brunskill, and Zoran Popovic · 2016
Later among the works it cites.
Adaptive skills adaptive partitions (asap)
Daniel J. Mankowitz, Timothy A. Mann, and Shie Mannor · 2016
Later among the works it cites.
Asynchronous Methods for Deep Reinforcement Learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Rate–distortion theory and human perception
Chris R. Sims · 2016
Later among the works it cites.
Toward good abstractions for lifelong learning
David Abel, Dilip Arumugam, Lucas Lehnert, and Michael L. Littman · 2017
Later among the works it cites.
An alternative softmax operator for reinforcement learning
Kavosh Asadi and Michael L. Littman · 2017
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Later among the works it cites.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J. Hunt, Tom Schaul, Hado van Hasselt, and David Silver · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Later among the works it cites.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Later among the works it cites.
Exploration–exploitation in MDPs with options
Ronan Fruit and Alessandro Lazaric · 2017
Later among the works it cites.
Regret minimization in MDPs with options without prior knowledge
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Emma Brunskill · 2017
Later among the works it cites.
Planning with abstract Markov decision processes
Nakul Gopalan, Marie desJardins, Michael L. Littman, James MacGlashan, Shawn Squire, Stefanie Tellex, John Winder, and Lawson L.S. Wong · 2017
Later among the works it cites.
An analysis of Monte Carlo tree search
Steven James, George Konidaris, and Benjamin Rosman · 2017
Later among the works it cites.
Deep variational Bayes filters: Unsupervised learning of state space models from raw data
Maximilian Karl, Maximilian Soelch, Justin Bayer, and Patrick van der Smagt · 2017
Later among the works it cites.
An efficient approach to model-based hierarchical reinforcement learning
Zhuoru Li, Akshay Narayan, and Tze-Yun Leong · 2017
Later among the works it cites.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Later among the works it cites.
Neural bases of action abstraction
Lorna C. Quandt, Yune Sang Lee, and Anjan Chatterjee · 2017
Later among the works it cites.
The deterministic information bottleneck
DJ Strouse and David J. Schwab · 2017
Later among the works it cites.
Self-correcting models for model-based reinforcement learning
Erik Talvitie · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Later among the works it cites.
Regularizing reinforcement learning with state abstraction
Riad Akrour, Filipe Veiga, Jan Peters, and Gerhard Neumann · 2018
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
Kavosh Asadi, Dipendra Misra, and Michael L. Littman · 2018
Later among the works it cites.
Feature-based aggregation and deep reinforcement learning: A survey and some new implementations
Dimitri P. Bertsekas · 2018
Later among the works it cites.
Evidence for hierarchically-structured reinforcement learning in humans
Maria Eckstein and Anne Collins · 2018
Later among the works it cites.
When waiting is not an option: Learning options with a deliberation cost
Jean Harb, Pierre-Luc Bacon, Martin Klissarov, and Doina Precup · 2018
Later among the works it cites.
Learning with options that terminate off-policy
Anna Harutyunyan, Peter Vrancx, Pierre-Luc Bacon, Doina Precup, and Ann Nowé · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver · 2018
Later among the works it cites.
Learning to plan with portable symbols
Steven James, Benjamin Rosman, and George Konidaris · 2018
Later among the works it cites.
Disentangling by factorising
Hyunjik Kim and Andriy Mnih · 2018
Later among the works it cites.
From skills to symbols: Learning symbolic representations for abstract high-level planning
George Konidaris, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2018
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Later among the works it cites.
Hierarchical imitation and reinforcement learning
Hoang M. Le, Nan Jiang, Alekh Agarwal, Miroslav Dudík, Yisong Yue, and Hal Daumé III · 2018
Later among the works it cites.
Transfer with model features in reinforcement learning
Lucas Lehnert and Michael L Littman · 2018
Later among the works it cites.
On value function representation of long horizon problems
Lucas Lehnert, Romain Laroche, and Harm van Seijen · 2018
Later among the works it cites.
State representation learning for control: An overview
Timothée Lesort, Natalia Díaz-Rodríguez, Jean-Franois Goudou, and David Filliat · 2018
Later among the works it cites.
Eigenoption Discovery through the Deep Successor Representation
Marlos C. Machado, Clemens Rosenbaum, Xiaoxiao Guo, Miao Liu, Gerald Tesauro, and Murray Campbell · 2018
Later among the works it cites.
State abstraction synthesis for discrete models of continuous domains
Jacob Menashe and Peter Stone · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Shane Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
Learning abstract options
Matthew Riemer, Miao Liu, and Gerald Tesauro · 2018
Later among the works it cites.
Efficient coding explains the universal law of generalization in human perception
Chris R. Sims · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Approximate exploration through state abstraction
Adrien Ali Taïga, Aaron Courville, and Marc G. Bellemare · 2018
Later among the works it cites.
The option keyboard: Combining skills in reinforcement learning
André Barreto, Diana Borsa, Shaobo Hou, Gheorghe Comanici, Eser Aygün, Philippe Hamel, Daniel Toyama, Jonathan Hunt, Shibl Mourad, David Silver, and Doina Precup · 2019
Later among the works it cites.
Provably efficient RL with rich observations via latent state decoding
Simon S. Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudík, and John Langford · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Later among the works it cites.
Combined reinforcement learning via abstract representations
Vincent François-Lavet, Yoshua Bengio, Doina Precup, and Joelle Pineau · 2019
Later among the works it cites.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Later among the works it cites.
Hindsight credit assignment
Anna Harutyunyan, Will Dabney, Thomas Mesnard, Mohammad Gheshlaghi Azar, Bilal Piot, Nicolas Heess, Hado van Hasselt, Gregory Wayne, Satinder Singh, and Doina Precup · 2019
Later among the works it cites.
The value of abstraction
Mark K. Ho, David Abel, Thomas L. Griffiths, and Michael L. Littman · 2019
Later among the works it cites.
On the necessity of abstraction
George Konidaris · 2019
Later among the works it cites.
Successor features support model-based and model-free reinforcement learning
Lucas Lehnert and Michael L Littman · 2019
Later among the works it cites.
Learning multi-level hierarchies with hindsight
Andrew Levy, George Konidaris, Robert Platt, and Kate Saenko · 2019
Later among the works it cites.
Performance guarantees for homomorphisms beyond Markov decision processes
Sultan Javed Majeed and Marcus Hutter · 2019
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2019
Later among the works it cites.
Near-optimal representation learning for hierarchical reinforcement learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2019
Later among the works it cites.
Regret bounds for learning state representations in reinforcement learning
Ronald Ortner, Matteo Pirotta, Alessandro Lazaric, Ronan Fruit, and Odalric-Ambrym Maillard · 2019
Later among the works it cites.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Later among the works it cites.
Natural option critic
Saket Tiwari and Philip S. Thomas · 2019
Later among the works it cites.
Composing value functions in reinforcement learning
Benjamin Van Niekerk, Steven James, Adam Earle, and Benjamin Rosman · 2019
Later among the works it cites.
DAC: The double actor-critic architecture for learning options
Shangtong Zhang and Shimon Whiteson · 2019
Later among the works it cites.
Value preserving state-action abstractions
David Abel, Nathan Umbanhowar, Khimya Khetarpal, Dilip Arumugam, Doina Precup, and Michael L. Littman · 2020
Later among the works it cites.
Learning state abstractions for transfer in continuous control
Kavosh Asadi, David Abel, and Michael L. Littman · 2020
Later among the works it cites.
Option discovery using deep skill chaining
Akhil Bagaria and George Konidaris · 2020
Later among the works it cites.
Scalable methods for computing state similarity in deterministic Markov decision processes
Pablo Samuel Castro · 2020
Later among the works it cites.
A distributional code for value in dopamine-based reinforcement learning
Will Dabney, Zeb Kurth-Nelson, Naoshige Uchida, Clara Kwon Starkweather, Demis Hassabis, Rémi Munos, and Matthew Botvinick · 2020
Later among the works it cites.
The efficiency of human cognition reflects planned use of information processing
Mark K. Ho, David Abel, Jonathan D. Cohen, Michael L. Littman, and Thomas L. Griffiths · 2020
Later among the works it cites.
Exploration in reinforcement learning with deep covering options
Yuu Jinnai, Jee Won Park, Marlos C. Machado, and George Konidaris · 2020
Later among the works it cites.
Options of interest: Temporal abstraction with interest functions
Khimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon, and Doina Precup · 2020
Later among the works it cites.
Learning skill hierarchies from predicate descriptions and self-supervision
Tom Silver, Rohan Chitnis, Anurag Ajay, Josh Tenenbaum, and Leslie Pack Kaelbling · 2020
Later among the works it cites.
Planning with abstract learned models while learning transferable subtasks
John Winder, Stephanie Milani, Matthew Landen, Erebus Oh, Shane Parr, Shawn Squire, Marie desJardins, and Cynthia Matuszek · 2020
Later among the works it cites.