Fetching the paper…
Reading the bibliography…
Reinforcement learning can train policies that effectively perform complex tasks.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Alex X. Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 1907
Earlier work this paper cites.
Soft actor-critic for discrete action settings
Petros Christodoulou · 1910
Earlier work this paper cites.
Improving sample efficiency in model-free reinforcement learning from images
Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus · 1910
Earlier work this paper cites.
The theory of affordances
James J. Gibson · 1977
Earlier work this paper cites.
Making the world differentiable: On using self-supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
Learning robot control - using control policies as abstract actions
Manfred Huber and Roderic A. Grupen · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G. Dietterich · 2000
Earlier work this paper cites.
Temporal abstraction in reinforcement learning, 2000
Doina Precup · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy McGovern and Andrew G. Barto · 2001
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton · 2002
Earlier work this paper cites.
Q-cut - dynamic discovery of sub-goals in reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2002
Earlier work this paper cites.
Using relative novelty to identify useful temporal abstractions in reinforcement learning
Özgür Şimşek and Andrew G. Barto · 2004
Earlier work this paper cites.
Identifying useful subgoals in reinforcement learning by local graph partitioning
Özgür Şimşek, Alicia P. Wolfe, and Andrew G. Barto · 2005
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Building portable options: Skill transfer in reinforcement learning
George Konidaris and Andrew Barto · 2007
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Linear bellman combination for control of character animation
Marco da Silva, Frédo Durand, and Jovan Popović · 2009
Earlier work this paper cites.
Skill discovery in continuous reinforcement learning domains using skill chaining
George Konidaris and Andrew Barto · 2009
Earlier work this paper cites.
A survey of numerical methods for optimal control
Anil V. Rao · 2009
Earlier work this paper cites.
Compositionality of optimal control laws
Emanuel Todorov · 2009
Earlier work this paper cites.
Predicting future object states using learned affordances
Emre Ugur, Erol Şahin, and Ethan Oztop · 2009
Earlier work this paper cites.
Affordance as general value function: A computational model
Daniel Graves, Johannes Günther, and Jun Luo · 2010
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D. Ziebart · 2010
Earlier work this paper cites.
Hierarchical reinforcement learning with movement primitives
Freek Stulp and Stefan Schaal · 2011
Earlier work this paper cites.
Autonomous reinforcement learning on raw visual input data in a real world application
S. Lange, Martin A. Riedmiller, and Arne Voigtländer · 2012
Earlier work this paper cites.
Motor primitive discovery
Philip S. Thomas and Andrew G. Barto · 2012
Earlier work this paper cites.
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, and Jan Peters · 2014
Earlier work this paper cites.
Constructing symbolic representations for high-level planning
George Konidaris, Leslie Kaelbling, and Tomas Lozano-Perez · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Model predictive path integral control using covariance variable importance sampling
Grady Williams, Andrew Aldrich, and Evangelos A. Theodorou · 2015
Cited alongside, same era.
Hierarchical relative entropy policy search
Christian Daniel, Gerhard Neumann, Oliver Kroemer, and Jan Peters · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Hierarchical reinforcement learning using spatio-temporal abstractions and deep neural networks
Ramnandan Krishnamurthy, Aravind S. Lakshminarayanan, Peeyush Kumar, and Balaraman Ravindran · 2016
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
Ashvin Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Later among the works it cites.
Learning abstract options
Matthew Riemer, Miao Liu, and Gerald Tesauro · 2018
Later among the works it cites.
Learning abstract options
Matthew Riemer, Miao Liu, and Gerald Tesauro · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D. Kulkarni, Karthik R. Narasimhan, Ardavan Saeedi, and Joshua B. Tenenbaum · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Manfred Otto Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
Evan Shelhamer, Parsa Mahmoudieh, Max Argus, and Trevor Darrell · 2016
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Later among the works it cites.
Semi-parametric topological memory for navigation
Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
David Warde-Farley, Tom Van de Wiele, Tejas D. Kulkarni, Catalin Ionescu, Steven Hansen, and Volodymyr Mnih · 2018
Later among the works it cites.
The option keyboard: Combining skills in reinforcement learning
Andre Barreto, Diana Borsa, Shaobo Hou, Gheorghe Comanici, Eser Aygün, Philippe Hamel, Daniel Toyama, Jonathan hunt, Shibl Mourad, David Silver, and Doina Precup · 2019
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Ben Eysenbach, Russ R Salakhutdinov, and Sergey Levine · 2019
Later among the works it cites.
Learning actionable representations with goal conditioned policies
Dibya Ghosh, Abhishek Gupta, and Sergey Levine · 2019
Later among the works it cites.
Transfer and exploration via the information bottleneck
Anirudh Goyal, Riashat Islam, DJ Strouse, Zafarali Ahmed, Hugo Larochelle, Matthew Botvinick, Sergey Levine, and Yoshua Bengio · 2019
Later among the works it cites.
Shaping Belief States with Generative Environment Models for RL
Karol Gregor, Danilo Jimenez Rezende, Frederic Besse, Yan Wu, Hamza Merzic, and Aäron van den Oord · 2019
Later among the works it cites.
Composing entropic policies using divergence correction, 2019
Jonathan J Hunt, Andre Barreto, Timothy P Lillicrap, and Nicolas Heess · 2019
Later among the works it cites.
Robot motion planning in learned latent spaces
Brian Ichter and Marco Pavone · 2019
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
Kendall Lowrey, Aravind Rajeswaran, Sham Kakade, Emanuel Todorov, and Igor Mordatch · 2019
Later among the works it cites.
Why does hierarchy (sometimes) work so well in reinforcement learning?, 2019
Ofir Nachum, Haoran Tang, Xingyu Lu, Shixiang Gu, Honglak Lee, and Sergey Levine · 2019
Later among the works it cites.
Natural option critic
Saket Tiwari and Philip S. Thomas · 2019
Later among the works it cites.
SOLAR: Deep structured representations for model-based reinforcement learning
Marvin Zhang, Sharad Vikram, Laura Smith, Pieter Abbeel, Matthew Johnson, and Sergey Levine · 2019
Later among the works it cites.
Scalable methods for computing state similarity in deterministic markov decision processes
Pablo Samuel Castro · 2020
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Later among the works it cites.
CURL: Contrastive unsupervised representations for reinforcement learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, and Marc G Bellemare · 2021
Closest in time.
Broadly-exploring, local-policy trees for long-horizon task planning
Brian Ichter, Pierre Sermanet, and Corey Lynch · 2021
Closest in time.
BC-z: Zero-shot task generalization with robotic imitation learning
Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn · 2021
Closest in time.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
Dmitry Kalashnikov, Jacob Varley, Yevgen Chebotar, Benjamin Swanson, Rico Jonschkowski, Chelsea Finn, Sergey Levine, and Karol Hausman · 2021
Closest in time.
Learning dexterous grasping with object-centric visual affordances, 2021
Priyanka Mandikal and Kristen Grauman · 2021
Closest in time.
Deep affordance foresight: Planning through what can be done in the future, 2021
Danfei Xu, Ajay Mandlekar, Roberto Martín-Martín, Yuke Zhu, Silvio Savarese, and Li Fei-Fei · 2021
Closest in time.