Fetching the paper…
Reading the bibliography…
Finding different solutions to the same problem is a key aspect of intelligence associated with creativity and adaptation to novel situations.
Zur theorie der gesellschaftsspiele
J. von Neumann · 1928
Earlier work this paper cites.
Applied imagination
A. F. Osborn · 1953
Earlier work this paper cites.
An algorithm for quadratic programming
M. Frank and P. Wolfe · 1956
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 1984
Earlier work this paper cites.
Exact sampling with coupled markov chains and applications to statistical mechanics
J. G. Propp and D. B. Wilson · 1996
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
E. Altman · 1999
Earlier work this paper cites.
Elements of information theory
T. M. Cover · 1999
Earlier work this paper cites.
Game 8: Leko wins to take the lead, 2004
R. Behovits · 2004
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
V. S. Borkar · 2005
Earlier work this paper cites.
Transfer in variable-reward hierarchical reinforcement learning
N. Mehta, S. Natarajan, P. Tadepalli, and A. Fern · 2008
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
S. Singh, R. L. Lewis, A. G. Barto, and J. Sorg · 2010
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
S. Bhatnagar and K. Lakshmanan · 2012
Earlier work this paper cites.
Determinantal point processes for machine learning
A. Kulesza, B. Taskar, et al · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
On the global linear convergence of frank-wolfe optimization variants
M. Jaggi and S. Lacoste-Julien · 2015
Cited alongside, same era.
Illuminating search spaces by mapping elites
J.-B. Mouret and J. Clune · 2015
Cited alongside, same era.
Why greatness cannot be planned: The myth of the objective
K. O. Stanley and J. Lehman · 2015
Cited alongside, same era.
Quality diversity: A new frontier for evolutionary computation
J. K. Pugh, L. B. Soros, and K. O. Stanley · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
D. Soudry and Y. Carmon · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Reward constrained policy optimization
C. Tessler, D. J. Mankowitz, and S. Mannor · 2019
Later among the works it cites.
Learning novel policies for tasks
Y. Zhang, W. Yu, and G. Turk · 2019
Later among the works it cites.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
A. Agarwal, M. Henaff, S. Kakade, and W. Sun · 2020
Later among the works it cites.
Relative variational intrinsic control
K. Baumli, D. Warde-Farley, S. Hansen, and V. Mnih · 2020
Later among the works it cites.
Swapping songs with chess grandmaster garry kasparov, 2020
C. da Fonseca-Wollheim · 2020
Later among the works it cites.
Harnessing distribution ratio estimators for learning agents with quality and diversity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. P. van Hasselt, and D. Silver · 2017
Cited alongside, same era.
Variational intrinsic control
K. Gregor, D. J. Rezende, and D. Wierstra · 2017
Cited alongside, same era.
Variational option discovery algorithms
J. Achiam, H. Edwards, D. Amodei, and P. Abbeel · 2018
Cited alongside, same era.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune · 2018
Cited alongside, same era.
Diversity-driven exploration strategy for deep reinforcement learning
Z.-W. Hong, T.-Y. Shann, S.-Y. Su, Y.-H. Chang, T.-J. Fu, and C.-Y. Lee · 2018
Cited alongside, same era.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning
Z. Xu, H. van Hasselt, and D. Silver · 2018
Cited alongside, same era.
T. Gangwani, J. Peng, and Y. Zhou · 2020
Later among the works it cites.
Fast task inference with variational intrinsic successor features
S. Hansen, W. Dabney, A. Barreto, D. Warde-Farley, T. V. de Wiele, and V. Mnih · 2020
Later among the works it cites.
One solution is not all you need: Few-shot extrapolation via structured maxent rl
S. Kumar, A. Kumar, S. Levine, and C. Finn · 2020
Later among the works it cites.
Evaluating agents without rewards
B. Matusch, J. Ba, and D. Hafner · 2020
Later among the works it cites.
Effective diversity in population based reinforcement learning
J. Parker-Holder, A. Pacchiano, K. M. Choromanski, and S. J. Roberts · 2020
Later among the works it cites.
Non-local policy optimization via diversity-regularized collaborative exploration
Z. Peng, H. Sun, and B. Zhou · 2020
Later among the works it cites.
Novel policy seeking with constrained optimization
H. Sun, Z. Peng, B. Dai, J. Guo, D. Lin, and B. Zhou · 2020
Later among the works it cites.
Constrained mdps and the reward hypothesis, 2020
C. Szepesvári · 2020
Later among the works it cites.
Balancing constraints and rewards with meta-gradient d4pg
D. A. Calian, D. J. Mankowitz, T. Zahavy, Z. Xu, J. Oh, N. Levine, and T. Mann · 2021
Closest in time.
Discovering a set of policies for the worst case reward
T. Zahavy, A. Barreto, D. J. Mankowitz, S. Hou, B. O’Donoghue, I. Kemaev, and S. Singh · 2021
Closest in time.