Fetching the paper…
Reading the bibliography…
We study the problem of predicting and controlling the future state distribution of an autonomous agent.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
Hilbert J Kappen · 2005
Earlier work this paper cites.
Temporal-difference networks
Richard S Sutton and Brian Tanner · 2005
Earlier work this paper cites.
Discriminative learning for differing training and test distributions
Steffen Bickel, Michael Brückner, and Tobias Scheffer · 2007
Earlier work this paper cites.
General duality between optimal control and estimation
Emanuel Todorov · 2008
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
Evangelos Theodorou, Jonas Buchli, and Stefan Schaal · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Thermodynamics as a theory of decision-making with information-processing costs
Pedro A Ortega and Daniel A Braun · 2013
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2013
Earlier work this paper cites.
Universal option models
Csaba Szepesvari, Richard S Sutton, Joseph Modayil, Shalabh Bhatnagar, et al · 2014
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Adversarially learned inference
Vincent Dumoulin, Ishmael Belghazi, Ben Poole, Olivier Mastropietro, Alex Lamb, Martin Arjovsky, and Aaron Courville · 2016
Earlier work this paper cites.
Loss is its own reward: Self-supervision for reinforcement learning
Evan Shelhamer, Parsa Mahmoudieh, Max Argus, and Trevor Darrell · 2016
Earlier work this paper cites.
Amortised map inference for image super-resolution
Casper Kaae Sønderby, Jose Caballero, Lucas Theis, Wenzhe Shi, and Ferenc Huszár · 2016
Earlier work this paper cites.
Generative adversarial nets from a density ratio estimation perspective
Masatoshi Uehara, Issei Sato, Masahiro Suzuki, Kotaro Nakayama, and Yutaka Matsuo · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, Hado P van Hasselt, and David Silver · 2017
Cited alongside, same era.
Variational inference using implicit distributions
Ferenc Huszár · 2017
Cited alongside, same era.
Data-efficient deep reinforcement learning for dexterous manipulation
Ivaylo Popov, Nicolas Heess, Timothy Lillicrap, Roland Hafner, Gabriel Barth-Maron, Matej Vecerik, Thomas Lampe, Yuval Tassa, Tom Erez, and Martin Riedmiller · 2017
Cited alongside, same era.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Frederik Ebert, Chelsea Finn, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Curriculum-guided hindsight experience replay
Meng Fang, Tianyi Zhou, Yali Du, Lei Han, and Zhengyou Zhang · 2019
Later among the works it cites.
Diagnosing bottlenecks in deep q-learning algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Later among the works it cites.
Learning to reach goals without reinforcement learning
Dibya Ghosh, Abhishek Gupta, Justin Fu, Ashwin Reddy, Coline Devin, Benjamin Eysenbach, and Sergey Levine · 2019
Later among the works it cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scott Fujimoto, Herke Van Hoof, and David Meger · 2018
Cited alongside, same era.
Temporal difference variational auto-encoder
Karol Gregor, George Papamakarios, Frederic Besse, Lars Buesing, and Theophane Weber · 2018
Cited alongside, same era.
Tf-agents: A library for reinforcement learning in tensorflow, 2018
Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, Ethan Holly, Sam Fishman, Ke Wang, Ekaterina Gonina, Neal Wu, Chris Harris, et al · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
Near-optimal representation learning for hierarchical reinforcement learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2018
Cited alongside, same era.
Junhyuk Oh, Yijie Guo, Satinder Singh, and Honglak Lee · 2018
Cited alongside, same era.
Xingyu Lin, Harjatin Singh Baweja, and David Held · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Later among the works it cites.
Planning with goal-conditioned policies
Soroush Nasiriany, Vitchyr Pong, Steven Lin, and Sergey Levine · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
Vitchyr H Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine · 2019
Later among the works it cites.
Policy continuation with hindsight inverse dynamics
Hao Sun, Zhizhong Li, Xiaotong Liu, Bolei Zhou, and Dahua Lin · 2019
Later among the works it cites.
Maximum entropy-regularized multi-goal reinforcement learning
Rui Zhao, Xudong Sun, and Volker Tresp · 2019
Later among the works it cites.
Neural topological slam for visual navigation
Devendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, and Saurabh Gupta · 2020
Closest in time.
Rewriting history with inverse rl: Hindsight inference for policy improvement
Benjamin Eysenbach, Xinyang Geng, Sergey Levine, and Ruslan Salakhutdinov · 2020
Closest in time.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Closest in time.
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 2020
Closest in time.
Goal-aware prediction: Learning to model what matters
Suraj Nair, Silvio Savarese, and Chelsea Finn · 2020
Closest in time.
Maximum entropy gain exploration for long horizon multi-goal reinforcement learning
Silviu Pitis, Harris Chan, Stephen Zhao, Bradly Stadie, and Jimmy Ba · 2020
Closest in time.
Yannick Schroecker and Charles Isbell · 2020
Closest in time.
dm_control: Software and tasks for continuous control
Yuval Tassa, Saran Tunyasuvunakool, Alistair Muldal, Yotam Doron, Siqi Liu, Steven Bohez, Josh Merel, Tom Erez, Timothy Lillicrap, and Nicolas Heess · 2020
Closest in time.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Closest in time.