Fetching the paper…
Reading the bibliography…
Learning with an objective to minimize the mismatch with a reference distribution has been shown to be useful for generative modeling and imitation learning.
Wasserstein Adversarial Imitation Learning
Huang Xiao, Michael Herman, Joerg Wagner, Sebastian Ziesche, Jalal Etesami, and Thai Hong Linh · 1906
Earlier work this paper cites.
What Can Learned Intrinsic Rewards Capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver, and Satinder Singh · 1912
Earlier work this paper cites.
A formal basis for the heuristic determination of minimum cost paths
Peter E Hart, Nils J Nilsson, and Bertram Raphael · 1968
Earlier work this paper cites.
Quasi uniformities: Reconciling domains with metric spaces
M. Smyth · 1987
Earlier work this paper cites.
Markov decision processes
Martin L Puterman · 1990
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Intrinsically motivated learning of hierarchical collections of skills
Andrew G Barto, Satinder Singh, and Nuttapong Chentanez · 2004
Earlier work this paper cites.
Intrinsic motivation for reinforcement learning systems
Andrew G Barto and Ozgür Simsek · 2005
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Satinder Singh, Andrew G Barto, and Nuttapong Chentanez · 2005
Earlier work this paper cites.
Primal Wasserstein Imitation Learning
Robert Dadashi, Léonard Hussenot, Matthieu Geist, and Olivier Pietquin · 2006
Earlier work this paper cites.
An intrinsic reward mechanism for efficient exploration
Özgür Şimşek and Andrew G Barto · 2006
Earlier work this paper cites.
Automatic Curriculum Learning through Value Disagreement
Yunzhi Zhang, Pieter Abbeel, and Lerrel Pinto · 2006
Earlier work this paper cites.
Evolving internal reinforcers for an intrinsically motivated reinforcement-learning robot
Massimiliano Schembri, Marco Mirolli, and Gianluca Baldassarre · 2007
Earlier work this paper cites.
How can we define intrinsic motivation?
Pierre-Yves Oudeyer and Frederic Kaplan · 2008
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Cédric Villani · 2008
Earlier work this paper cites.
R-iac: Robust intrinsically motivated exploration and active learning
Adrien Baranes and Pierre-Yves Oudeyer · 2009
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan · 2009
Earlier work this paper cites.
Near-optimal Regret Bounds for Reinforcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Evolved intrinsic reward functions for reinforcement learning
Scott Niekum · 2010
Earlier work this paper cites.
Intrinsically Motivated Reinforcement Learning: An Evolutionary Perspective
Satinder Singh, Richard L. Lewis, Andrew G. Barto, and Jonathan Sorg · 2010
Earlier work this paper cites.
What are intrinsic motivations? a biological perspective
Gianluca Baldassarre · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Intrinsic Motivation and Reinforcement Learning
Andrew G. Barto · 2013
Cited alongside, same era.
Information-seeking, curiosity, and attention: computational and neural mechanisms
Jacqueline Gottlieb, Pierre-Yves Oudeyer, Manuel Lopes, and Adrien Baranes · 2013
Cited alongside, same era.
Which is the best intrinsic motivation signal for learning multiple skills?
Vieri Giuliano Santucci, Gianluca Baldassarre, and Marco Mirolli · 2013
Cited alongside, same era.
Intrinsic motivations and open-ended development in animals, humans, and robots: an overview
Gianluca Baldassarre, Tom Stafford, Marco Mirolli, Peter Redgrave, Richard M Ryan, and Andrew Barto · 2014
Cited alongside, same era.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Universal value function approximators
Sharpening jensen’s inequality
JG Liao and Arthur Berg · 2018
Later among the works it cites.
Rl baselines zoo
Antonin Raffin · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
On Learning Intrinsic Rewards for Policy Gradient Methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
Jose A Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, and Sepp Hochreiter · 2019
Later among the works it cites.
Smirl: Surprise minimizing reinforcement learning in unstable environments
Glen Berseth, Daniel Geng, Coline Devin, Nicholas Rhinehart, Chelsea Finn, Dinesh Jayaraman, and Sergey Levine · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Schaul, Daniel Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Faulty reward functions in the wild, Dec 2016
Jack Clark and Dario Amodei · 2016
Cited alongside, same era.
Deep reinforcement learning in parameterized action space
Matthew Hausknecht and Peter Stone · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Later among the works it cites.
Goal-conditioned imitation learning
Y. Ding, Carlos Florensa, Mariano Phielipp, and P. Abbeel · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Computational optimal transport
Gabriel Peyré and Marco Cuturi · 2019
Later among the works it cites.
Generative Adversarial Imitation from Observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2019
Later among the works it cites.
Self-supervised learning of distance functions for goal-conditioned reinforcement learning
Srinivas Venkattaramanujam, E. Crawford, T. Doan, and Doina Precup · 2019
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2020
Later among the works it cites.
The stochastic shortest path problem: a polyhedral combinatorics perspective
Matthieu Guillot and Gautier Stauffer · 2020
Later among the works it cites.
Dynamical distance learning for semi-supervised and unsupervised skill discovery
Kristian Hartikainen, Xinyang Geng, Tuomas Haarnoja, and Sergey Levine · 2020
Later among the works it cites.
Exploration in reinforcement learning with deep covering options
Yuu Jinnai, Jee Won Park, Marlos C. Machado, and George Konidaris · 2020
Later among the works it cites.
Lipschitz constrained gans via boundedness and continuity
Kanglin Liu and Guoping Qiu · 2020
Later among the works it cites.
Disco rl: Distribution-conditioned reinforcement learning for general-purpose policies
Soroush Nasiriany · 2020
Later among the works it cites.
F-irl: Inverse reinforcement learning via state marginal matching
Tianwei Ni, Harshit Sikchi, Yufei Wang, Tejus Gupta, Lisa Lee, and Benjamin Eysenbach · 2020
Later among the works it cites.
Is the policy gradient a gradient?
Chris Nota and Philip S Thomas · 2020
Later among the works it cites.
An information-theoretic perspective on credit assignment in reinforcement learning
Dilip Arumugam, P. Henderson, and P. Bacon · 2021
Closest in time.
Reward (mis) design for autonomous driving
W Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone · 2021
Closest in time.
Wasserstein gans work because they fail (to approximate the wasserstein distance)
Jan Stanczuk, Christian Etmann, L. Kreusser, and C. Schönlieb · 2021
Closest in time.