Fetching the paper…
Reading the bibliography…
In problem-solving, we humans can come up with multiple novel solutions to the same problem.
A. Ruszczyński, “Feasible direction methods for stochastic programming problems,” Mathematical Programming
1980
Earlier work this paper cites.
L. Rüschendorf, “The wasserstein distance and approximation theorems,” Probability Theory and Related Fields
1985
Earlier work this paper cites.
Oxford university press, 1990
B. Rogoff, Apprenticeship in thinking: Cognitive development in social context · 1990
Earlier work this paper cites.
A. Conn, N. Gould, and P. Toint, “A globally convergent lagrangian barrier algorithm for optimization with general inequality constraints and simple bounds,” Mathematics of Computation of the American Mathematical Society
1997
Earlier work this paper cites.
MIT press Cambridge, 1998
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
J. Herskovits, “Feasible direction interior-point technique for nonlinear optimization,” Journal of optimization theory and applications
1998
Earlier work this paper cites.
J. Randløv and P. Alstrøm, “Learning to drive a bicycle using reinforcement learning and shaping.,” in ICML
1998
Earlier work this paper cites.
CRC Press, 1999
E. Altman, Constrained Markov decision processes · 1999
Earlier work this paper cites.
R. M. Ryan and E. L. Deci, “Intrinsic and extrinsic motivations: Classic definitions and new directions,” Contemporary educational psychology
2000
Earlier work this paper cites.
F. A. Potra and S. J. Wright, “Interior-point methods,” Journal of Computational and Applied Mathematics
2000
Earlier work this paper cites.
S. J. Wright, “On the convergence of the newton/log-barrier method,” Mathematical Programming
2001
Earlier work this paper cites.
D. M. Endres and J. E. Schindelin, “A new metric for probability distributions,” IEEE Transactions on Information theory
2003
Earlier work this paper cites.
B. Fuglede and F. Topsoe, “Jensen-shannon divergence and hilbert space embedding,” in International Symposium onInformation Theory, 2004. ISIT 2004. Proceedings
2004
Earlier work this paper cites.
University of Illinois at Urbana-Champaign, 2004
A. D. Laud, Theory and application of reward shaping in reinforcement learning · 2004
Earlier work this paper cites.
Springer Science & Business Media, 2006
G. B. Dantzig and M. N. Thapa, Linear programming 2: theory and extensions · 2006
Earlier work this paper cites.
Springer Science & Business Media, 2008
C. Villani, Optimal transport: old and new · 2008
Earlier work this paper cites.
C. P. van Schaik and J. M. Burkart, “Social learning and evolution: the cultural intelligence hypothesis,” Philosophical Transactions of the Royal Society B: Biological Sciences
2011
Earlier work this paper cites.
P. Sequeira, F. S. Melo, R. Prada, and A. Paiva, “Emerging social awareness: Exploring intrinsic motivation in multiagent learning,” in 2011 IEEE International Conference on Development and Learning (ICDL)
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control.,” in IROS
2012
Earlier work this paper cites.
Random House, 2014
Y. N. Harari, Sapiens: A brief history of humankind · 2014
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning
2015
Cited alongside, same era.
2015
Cited alongside, same era.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “Variational information maximizing exploration,” 2016
2016
Cited alongside, same era.
J. K. Pugh, L. B. Soros, and K. O. Stanley, “Quality diversity: A new frontier for evolutionary computation,” Frontiers in Robotics and AI
2016
Cited alongside, same era.
Y. Chow, O. Nachum, E. Duenez-Guzman, and M. Ghavamzadeh, “A lyapunov-based approach to safe reinforcement learning,” in Advances in neural information processing systems
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Thirty-Second AAAI Conference on Artificial Intelligence
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, “Openai gym,” 2016
2016
Cited alongside, same era.
Princeton University Press, 2017
J. Henrich, The secret of our success: How culture is driving human evolution, domesticating our species, and making us smarter · 2017
Cited alongside, same era.
2017
Cited alongside, same era.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops
2017
Cited alongside, same era.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70
2017
Cited alongside, same era.
2017
Cited alongside, same era.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in Proceedings of the 34th International Conference on Machine Learning
2017
Cited alongside, same era.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Jaques, A. Lazaridou, E. Hughes, C. Gulcehre, P. Ortega, D. Strouse, J. Z. Leibo, and N. De Freitas, “Social influence as intrinsic motivation for multi-agent deep reinforcement learning,” in International Conference on Machine Learning
2019
Later among the works it cites.
Y. Zhang, W. Yu, and G. Turk, “Learning novel policies for tasks,” in International Conference on Machine Learning
2019
Later among the works it cites.
Y. Zhang, W. Yu, and G. Turk, “Learning novel policies for tasks,” CoRR
2019
Later among the works it cites.
H. Liu, A. Trott, R. Socher, and C. Xiong, “Competitive experience replay,” CoRR
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Ray, J. Achiam, and D. Amodei, “Benchmarking safe exploration in deep reinforcement learning,” openai
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Ciosek, Q. Vuong, R. Loftin, and K. Hofmann, “Better exploration with optimistic actor critic,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
A. P. Badia, B. Piot, S. Kapturowski, P. Sprechmann, A. Vitvitskyi, Z. D. Guo, and C. Blundell, “Agent57: Outperforming the atari human benchmark,” in International Conference on Machine Learning
2020
Closest in time.
M. Elbarbari, K. Efthymiadis, B. Vanderborght, and A. Nowé, “Ltlf-based reward shaping for reinforcement learning,” in Adaptive and Learning Agents Workshop 2021
2021
Closest in time.