Fetching the paper…
Reading the bibliography…
Reinforcement learning agents are prone to undesired behaviors due to reward mis-specification.
Monte carlo sampling methods using markov chains and their applications
Hastings, W. K · 1970
Earlier work this paper cites.
Subjectivity and correlation in randomized strategies
Aumann, R. J · 1974
Earlier work this paper cites.
Statistical analysis of non-lattice data
Besag, J · 1975
Earlier work this paper cites.
Correlated equilibrium as an expression of bayesian rationality
Aumann, R. J · 1987
Earlier work this paper cites.
Compatible conditional distributions
Arnold, B. C. and Press, S. J · 1989
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Pomerleau, D. A · 1991
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
Quantal response equilibria for normal form games
McKelvey, R. D. and Palfrey, T. R · 1995
Earlier work this paper cites.
Multiagent reinforcement learning: theoretical framework and an algorithm
Hu, J., Wellman, M. P., and Others · 1998
Earlier work this paper cites.
Quantal response equilibria for extensive form games
McKelvey, R. D. and Palfrey, T. R · 1998
Earlier work this paper cites.
Learning agents for uncertain environments
Russell, S · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Hart, S. and Mas-Colell, A · 2000
Earlier work this paper cites.
Dependency networks for inference, collaborative filtering, and data visualization
Heckerman, D., Chickering, D. M., Meek, C., Rounthwaite, R., and Kadie, C · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Theory of point estimation
Lehmann, E. L. and Casella, G · 2006
Earlier work this paper cites.
No-regret learning in convex games
Gordon, G. J., Greenwald, A., and Marks, C · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Multi-agent inverse reinforcement learning
Natarajan, S., Kunapuli, G., Judah, K., Tadepalli, P., Kersting, K., and Shavlik, J · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Cited alongside, same era.
Gibbs ensembles for nearly compatible and incompatible conditional models
Chen, S.-H., Ip, E. H., and Wang, Y. J · 2011
Cited alongside, same era.
Theoretical considerations of potential-based reward shaping for multi-agent systems
Devlin, S. and Kudenko, D · 2011
Cited alongside, same era.
Best-response mechanisms
Nisan, N., Schapira, M., Valiant, G., and Zohar, A · 2011
Cited alongside, same era.
Maximum causal entropy correlated equilibria for markov games
Ziebart, B. D., Bagnell, J. A., and Dey, A. K · 2011
Cited alongside, same era.
The stochastic response dynamic: A new approach to learning and computing equilibrium in continuous games
Gandhi, A · 2012
Model-free imitation learning with policy optimization
Ho, J., Gupta, J., and Ermon, S · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2016
Later among the works it cites.
Making friends on the fly: Cooperating with new teammates
Barrett, S., Rosenfeld, A., Kraus, S., and Stone, P · 2017
Later among the works it cites.
Learning robust rewards with adversarial inverse reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Coordinated Multi-Robot exploration under communication constraints using decentralized markov decision processes
Matignon, L., Jeanpierre, L., Mouaddib, A.-I., and Others · 2012
Cited alongside, same era.
Inverse reinforcement learning for decentralized non-cooperative multiagent systems
Reddy, T. S., Gopikrishna, V., Zaruba, G., and Huber, M · 2012
Cited alongside, same era.
Learning objective functions for manipulation
Kalakrishnan, M., Pastor, P., Righetti, L., and Schaal, S · 2013
Cited alongside, same era.
Computational rationalization: The inverse equilibrium problem
Waugh, K., Ziebart, B. D., and Andrew Bagnell, J · 2013
Cited alongside, same era.
Multi-robot inverse reinforcement learning under occlusion with interactions
Bogert, K. and Doshi, P · 2014
Cited alongside, same era.
Theory and applications of proper scoring rules
Dawid, A. P. and Musio, M · 2014
Cited alongside, same era.
Fu, J., Luo, K., and Levine, S · 2017
Later among the works it cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S., Holly, E., Lillicrap, T., and Levine, S · 2017
Later among the works it cites.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A · 2017
Later among the works it cites.
Coordinated Multi-Agent imitation learning
Le, H. M., Yue, Y., and Carr, P · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Leibo, J. Z., Zambaldi, V., Lanctot, M., Marecki, J., and Graepel, T · 2017
Later among the works it cites.
InfoGAIL: Interpretable imitation learning from visual demonstrations
Li, Y., Song, J., and Ermon, S · 2017
Later among the works it cites.
Multi-Agent Actor-Critic for mixed Cooperative-Competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., and Mordatch, I · 2017
Later among the works it cites.
Multiagent Bidirectionally-Coordinated nets for learning to play StarCraft combat games
Peng, P., Yuan, Q., Wen, Y., Yang, Y., Tang, Z., Long, H., and Wang, J · 2017
Later among the works it cites.
Inverse reinforcement learning in swarm systems
Šošić, A., KhudaBukhsh, W. R., Zoubir, A. M., and Koeppl, H · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Liao, S., Grosse, R., and Ba, J · 2017
Later among the works it cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., and Yu, Y · 2017
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Later among the works it cites.
Multi-agent inverse reinforcement learning for general-sum stochastic games
Lin, X., Adams, S. C., and Beling, P. A · 2018
Later among the works it cites.
Multi-agent generative adversarial imitation learning
Song, J., Ren, H., Sadigh, D., and Ermon, S · 2018
Later among the works it cites.