Fetching the paper…
Reading the bibliography…
Goal-conditioned reinforcement learning (RL) is a promising direction for training agents that are capable of solving multiple tasks and reach a diverse set of objectives.
Asymptotic properties of non-linear least squares estimators
Robert I Jennrich · 1969
Earlier work this paper cites.
Logic and conversation
Herbert P Grice · 1975
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Richard Stuart Sutton · 1984
Earlier work this paper cites.
An architecture for intelligent reactive systems
Leslie Pack Kaelbling et al · 1987
Earlier work this paper cites.
Practical issues in temporal difference learning
Gerald Tesauro · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1992
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E. Hinton · 1992
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W Moore and Christopher G Atkeson · 1993
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Satinder P Singh and Richard S Sutton · 1996
Earlier work this paper cites.
Weak Convergence and Empirical Processes
Aad W. van der Vaart and Jon A. Wellner · 1996
Earlier work this paper cites.
Hq-learning
Marco A. Wiering and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Predictive reward signal of dopamine neurons
Wolfram Schultz · 1998
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G. Dietterich · 2000
Earlier work this paper cites.
Planning with a language for extended goals
Ugo Dal Lago, Marco Pistore, and Paolo Traverso · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew G Barto and Sridhar Mahadevan · 2003
Earlier work this paper cites.
Towards grounding concepts for transfer in goal learning from demonstration
Crystal Chao, Maya Cakmak, and Andrea L Thomaz · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Transfer reinforcement learning with shared dynamics
Romain Laroche and Merwan Barlier · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Learning to understand goal specifications by modelling reward
Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes, Arian Hosseini, Pushmeet Kohli, and Edward Grefenstette · 2018
Cited alongside, same era.
Babyai: A platform to study the sample efficiency of grounded language learning
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Grounding language to autonomously-acquired skills via goal generation
Ahmed Akakzia, Cédric Colas, Pierre-Yves Oudeyer, Mohamed Chetouani, and Olivier Sigaud · 2020
Later among the works it cites.
What can i do here? a theory of affordances in reinforcement learning
Khimya Khetarpal, Zafarali Ahmed, Gheorghe Comanici, David Abel, and Doina Precup · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Later among the works it cites.
Deep reinforcement and infomax learning
Bogdan Mazoure, Remi Tachet des Combes, Thang Long Doan, Philip Bachman, and R Devon Hjelm · 2020
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2018
Cited alongside, same era.
Near-optimal representation learning for hierarchical reinforcement learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2018
Cited alongside, same era.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, et al · 2018
Cited alongside, same era.
Temporal difference models: Model-free deep rl for model-based control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Cited alongside, same era.
Temporal difference models: Model-free deep RL for model-based control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Cited alongside, same era.
Learning goal embeddings via self-play for hierarchical reinforcement learning
Sainbayar Sukhbaatar, Emily Denton, Arthur Szlam, and Rob Fergus · 2018
Cited alongside, same era.
Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation
Suraj Nair and Chelsea Finn · 2020
Later among the works it cites.
RIDE: rewarding impact-driven exploration for procedurally-generated environments
Roberta Raileanu and Tim Rocktäschel · 2020
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2020
Later among the works it cites.
Generating adjacency-constrained subgoals in hierarchical reinforcement learning
Tianren Zhang, Shangqi Guo, Tian Tan, Xiaolin Hu, and Feng Chen · 2020
Later among the works it cites.
Learning with amigo: Adversarially motivated intrinsic goals
Andres Campero, Roberta Raileanu, Heinrich Küttler, Joshua B. Tenenbaum, Tim Rocktäschel, and Edward Grefenstette · 2021
Later among the works it cites.
Goal-conditioned reinforcement learning with imagined subgoals
Elliot Chane-Sane, Cordelia Schmid, and Ivan Laptev · 2021
Later among the works it cites.
Provable RL with exogenous distractors via multistep inverse dynamics
Yonathan Efroni, Dipendra Misra, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2021
Later among the works it cites.
C-learning: Learning to achieve goals via recursive classification
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2021
Later among the works it cites.
Dr jekyll & mr hyde: the strange case of off-policy policy updates
Romain Laroche and Remi Tachet des Combes · 2021
Later among the works it cites.
Discrete-valued neural communication
Dianbo Liu, Alex M Lamb, Kenji Kawaguchi, Anirudh Goyal, Chen Sun, Michael C Mozer, and Yoshua Bengio · 2021
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm, Aaron C. Courville, and Philip Bachman · 2021
Later among the works it cites.
Pretraining representations for data-efficient reinforcement learning
Max Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand, Laurent Charlin, R. Devon Hjelm, Philip Bachman, and Aaron C. Courville · 2021
Later among the works it cites.
Reward is enough
David Silver, Satinder Singh, Doina Precup, and Richard S Sutton · 2021
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Adam Stooke, Kimin Lee, Pieter Abbeel, and Michael Laskin · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making
Mengjiao Yang and Ofir Nachum · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
World model as a graph: Learning latent landmarks for planning
Lunjun Zhang, Ge Yang, and Bradly C. Stadie · 2021
Later among the works it cites.
Guaranteed discovery of controllable latent states with multi-step inverse models
Alex Lamb, Riashat Islam, Yonathan Efroni, Aniket Didolkar, Dipendra Misra, Dylan Foster, Lekan Molu, Rajan Chari, Akshay Krishnamurthy, and John Langford · 2022
Closest in time.
One-shot learning from a demonstration with hierarchical latent language, 2022
Nathaniel Weir, Xingdi Yuan, Marc-Alexandre Côté, Matthew Hausknecht, Romain Laroche, Ida Momennejad, Harm Van Seijen, and Benjamin Van Durme · 2022
Closest in time.