Fetching the paper…
Reading the bibliography…
This paper deals with the problem of learning a skill-conditioned policy that acts meaningfully in the absence of a reward signal.
Wasserstein Adversarial Imitation Learning
Huang Xiao, Michael Herman, Joerg Wagner, Sebastian Ziesche, Jalal Etesami, and Thai Hong Linh · 1906
Earlier work this paper cites.
What Can Learned Intrinsic Rewards Capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver, and Satinder Singh · 1912
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Jack Kiefer and Jacob Wolfowitz · 1952
Earlier work this paper cites.
Probability with martingales
David Williams · 1991
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Intrinsically motivated learning of hierarchical collections of skills
Andrew G Barto, Satinder Singh, and Nuttapong Chentanez · 2004
Earlier work this paper cites.
Intrinsic motivation for reinforcement learning systems
Andrew G Barto and Ozgür Simsek · 2005
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Satinder Singh, Andrew G Barto, and Nuttapong Chentanez · 2005
Earlier work this paper cites.
Primal Wasserstein Imitation Learning
Robert Dadashi, Léonard Hussenot, Matthieu Geist, and Olivier Pietquin · 2006
Earlier work this paper cites.
An intrinsic reward mechanism for efficient exploration
Özgür Şimşek and Andrew G Barto · 2006
Earlier work this paper cites.
Evolving internal reinforcers for an intrinsically motivated reinforcement-learning robot
Massimiliano Schembri, Marco Mirolli, and Gianluca Baldassarre · 2007
Earlier work this paper cites.
How can we define intrinsic motivation?
Pierre-Yves Oudeyer and Frederic Kaplan · 2008
Earlier work this paper cites.
R-iac: Robust intrinsically motivated exploration and active learning
Adrien Baranes and Pierre-Yves Oudeyer · 2009
Cited alongside, same era.
What is intrinsic motivation? a typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan · 2009
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Cited alongside, same era.
Evolved intrinsic reward functions for reinforcement learning
Scott Niekum · 2010
Cited alongside, same era.
Intrinsically Motivated Reinforcement Learning: An Evolutionary Perspective
Satinder Singh, Richard L. Lewis, Andrew G. Barto, and Jonathan Sorg · 2010
Cited alongside, same era.
Intrinsic Motivation and Reinforcement Learning
Andrew G. Barto · 2013
Cited alongside, same era.
Improved Training of Wasserstein GANs
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Later among the works it cites.
Synthesizing programs for images using reinforced adversarial learning
Yaroslav Ganin, Tejas Kulkarni, Igor Babuschkin, SM Ali Eslami, and Oriol Vinyals · 2018
Later among the works it cites.
Combinatorial Structure of Finite Metric Spaces
Filip Jevtić · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Which is the best intrinsic motivation signal for learning multiple skills?
Vieri Giuliano Santucci, Gianluca Baldassarre, and Marco Mirolli · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Wasserstein Generative Adversarial Networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
On Learning Intrinsic Rewards for Policy Gradient Methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Later among the works it cites.
Computational optimal transport
Gabriel Peyré and Marco Cuturi · 2019
Later among the works it cites.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Víctor Campos, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giró-i Nieto, and Jordi Torres · 2020
Later among the works it cites.
Wasserstein distance guided adversarial imitation learning with reward shape exploration
M. Zhang, Y. Wang, X. Ma, L. Xia, J. Yang, Z. Li, and X. Li · 2020
Later among the works it cites.
Relative variational intrinsic control
Kate Baumli, David Warde-Farley, Steven Hansen, and Volodymyr Mnih · 2021
Closest in time.
Adversarial intrinsic motivation for reinforcement learning
Ishan Durugkar, Mauricio Tec, Scott Niekum, and Peter Stone · 2021
Closest in time.