Fetching the paper…
Reading the bibliography…
Goal-Conditioned Reinforcement Learning (RL) problems often have access to sparse rewards where the agent receives a reward signal only when it has achieved the goal, making policy optimization a difficult problem.
Using natural language for reward shaping in reinforcement learning
Prasoon Goyal, Scott Niekum, and Raymond J. Mooney · 1903
Earlier work this paper cites.
Imitation learning as f-divergence minimization
Liyiming Ke, Matt Barnes, Wen Sun, Gilwoo Lee, Sanjiban Choudhury, and Siddhartha S. Srinivasa · 1905
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman · 1910
Earlier work this paper cites.
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard S. Zemel, and Shixiang Gu · 1911
Earlier work this paper cites.
Markov decision processes
Martin L Puterman · 1990
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
D4RL: datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew L. Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Evolved intrinsic reward functions for reinforcement learning
Scott Niekum · 2010
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Satinder Singh, Richard L. Lewis, Andrew G. Barto, and Jonathan Sorg · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D. Ziebart · 2010
Earlier work this paper cites.
C-learning: Learning to achieve goals via recursive classification
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2011
Earlier work this paper cites.
f-irl: Inverse reinforcement learning via state marginal matching
Tianwei Ni, Harshit S. Sikchi, Yufei Wang, Tejus Gupta, Lisa Lee, and Benjamin Eysenbach · 2011
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Hilbert J. Kappen, Vicenç Gómez, and Manfred Opper · 2012
Earlier work this paper cites.
Ving: Learning open-world navigation with visual goals
Dhruv Shah, Benjamin Eysenbach, Gregory Kahn, Nicholas Rhinehart, and Sergey Levine · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Intrinsic Motivation and Reinforcement Learning , pp. 17–47
Andrew G. Barto · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation, 2016
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Faulty reward functions in the wild, 2016
Jack Clark and Dario Amodei · 2016
Cited alongside, same era.
On the expressivity of markov reward
David Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho, Michael L. Littman, Doina Precup, and Satinder Singh · 2021
Later among the works it cites.
Adversarial intrinsic motivation for reinforcement learning
Ishan Durugkar, Mauricio Tec, Scott Niekum, and Peter Stone · 2021
Later among the works it cites.
Replacing rewards with examples: Example-based policy search via recursive classification
Benjamin Eysenbach, Sergey Levine, and Ruslan Salakhutdinov · 2021
Later among the works it cites.
Learning reward machines: A study in partially observable reinforcement learning
Rodrigo Toro Icarte, Ethan Waldie, Toryn Q. Klassen, Richard Anthony Valenzano, Margarita P. Castro, and Sheila A. McIlraith · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy P. Lillicrap, Karen Simonyan, and Demis Hassabis · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Provably efficient maximum entropy exploration
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2018
Cited alongside, same era.
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon Kohl, Andrew Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, and Demis Hassabis · 2021
Later among the works it cites.
Reward (mis)design for autonomous driving
W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone · 2021
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation
OpenAI, Matthias Plappert, Raul Sampedro, Tao Xu, Ilge Akkaya, Vineet Kosaraju, Peter Welinder, Ruben D’Sa, Arthur Petron, Henrique Pondé de Oliveira Pinto, Alex Paino, Hyeonwoo Noh, Lilian Weng, Qiming Yuan, Casey Chu, and Wojciech Zaremba · 2021
Later among the works it cites.
Reward is enough
David Silver, Satinder Singh, Doina Precup, and Richard S. Sutton · 2021
Later among the works it cites.
Discovering faster matrix multiplication algorithms with reinforcement learning
Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J. R. Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, David Silver, Demis Hassabis, and Pushmeet Kohli · 2022
Later among the works it cites.
Robot peels banana with goal-conditioned dual-action deep imitation learning, 2022
Heecheol Kim, Yoshiyuki Ohmura, and Yasuo Kuniyoshi · 2022
Later among the works it cites.
How far i’ll go: Offline goal-conditioned reinforcement learning via f f -advantage regression, 2022
Yecheng Jason Ma, Jason Yan, Dinesh Jayaraman, and Osbert Bastani · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Later among the works it cites.
"Information Theory From Coding to Learning"
Yury Polyanskiy and Yihong Wu · 2022
Later among the works it cites.
Outracing champion gran turismo drivers with deep reinforcement learning
Peter R Wurman, Samuel Barrett, Kenta Kawamoto, James MacGlashan, Kaushik Subramanian, Thomas J Walsh, Roberto Capobianco, Alisa Devlic, Franziska Eckert, Florian Fuchs, et al · 2022
Later among the works it cites.
The perils of trial-and-error reward design: misdesign through overfitting and invalid task specifications
Serena Booth, Julie Shah, Scott Niekum, Peter Stone, and Alessandro Allievi · 2023
Closest in time.
Estimation and control of visitation distributions for reinforcement learning
Ishan Durugkar et al · 2023
Closest in time.
Navigating to objects in the real world
Theophile Gervet, Soumith Chintala, Dhruv Batra, Jitendra Malik, and Devendra Singh Chaplot · 2023
Closest in time.