Fetching the paper…
Reading the bibliography…
Can artificial agents learn to assist others in achieving their goals without knowing what those goals are? Generic reinforcement learning agents could be trained to behave altruistically towards others by rewarding them for altruistic behaviour, i.e., rewarding them for benefiting other agents in a given situation.
1912
Earlier work this paper cites.
Watkins, C. J. C. H. & Dayan, P. (1992), ‘Q-learning’, Machine learning
1992
Earlier work this paper cites.
Tan, M. (1993), Multi-Agent Reinforcement Learning: Independent vs. Cooperative Agents, in
1993
Earlier work this paper cites.
Littman, M. L. (1994), Markov games as a framework for multi-agent reinforcement learning, in
1994
Earlier work this paper cites.
Dowding, K. & Monroe, K. R. (1997), ‘The Heart of Altruism: Perceptions of a Common Humanity’, The British Journal of Sociology
1997
Earlier work this paper cites.
Kaelbling, L. P., Littman, M. L. & Cassandra, A. R. (1998), ‘Planning and acting in partially observable stochastic domains’, Artificial intelligence
1998
Earlier work this paper cites.
Fehr, E. & Fischbacher, U. (2003), ‘The nature of human altruism’
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
Allen, C., Smit, I. & Wallach, W. (2005), ‘Artificial morality: Top-down, bottom-up, and hybrid approaches’, Ethics and Information Technology
2005
Earlier work this paper cites.
Cover, T. M. & Thomas, J. A. (2005), Elements of Information Theory
2005
Earlier work this paper cites.
Klyubin, A. S., Polani, D. & Nehaniv, C. L. (2005), Empowerment: A universal agent-centric measure of control, in
2005
Earlier work this paper cites.
Baker, C. L., Tenenbaum, J. B. & Saxe, R. (2006), ‘Bayesian models of human action understanding’, Advances in neural information processing systems
2006
Earlier work this paper cites.
Klyubin, A. S., Polani, D. & Nehaniv, C. L. (2008), ‘Keep your options open: An information-based driving principle for sensorimotor systems’, PLoS ONE
2008
Earlier work this paper cites.
Omohundro, S. (2008), ‘The basic ai drives. agi08proceedings of the first conference on artificial general intelligence’
2008
Earlier work this paper cites.
Macindoe, O., Kaelbling, L. P. & Lozano-Pérez, T. (2012), Pomcop: Belief space planning for sidekicks in cooperative games, in
2012
Earlier work this paper cites.
Dragan, A. D. & Srinivasa, S. S. (2013), ‘A policy-blending formalism for shared control’, The International Journal of Robotics Research
2013
Earlier work this paper cites.
Galvin, D. (2014), ‘Three tutorial lectures on entropy and counting’, arXiv preprint arXiv:1406.7872
2014
Earlier work this paper cites.
Javdani, S., Srinivasa, S. S. & Bagnell, J. A. (2015), ‘Shared autonomy via hindsight optimization’, Robotics science and systems: online proceedings
2015
Earlier work this paper cites.
Mohamed, S. & Rezende, D. J. (2015), Variational information maximisation for intrinsically motivated reinforcement learning, in
2015
Cited alongside, same era.
Pérez-D’Arpino, C. & Shah, J. A. (2015), Fast target prediction of human reaching motion for cooperative human-robot manipulation tasks using time series classification, in
2015
Cited alongside, same era.
Schulman, J., Levine, S., Moritz, P., Jordan, M. & Abbeel, P. (2015), Trust region policy optimization, in
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Benson-Tilsen, T. & Soares, N. (2016), Formalizing convergent instrumental goals, in
2016
Cited alongside, same era.
Mordatch, I. & Abbeel, P. (2018), Emergence of grounded compositional language in multi-agent populations, in
2018
Later among the works it cites.
Song, J., Ren, H., Ermon, S. & Sadigh, D. (2018), Multi-agent generative adversarial imitation learning, in
2018
Later among the works it cites.
Mao, H., Zhang, Z., Xiao, Z. & Gong, Z. (2019), Modelling the dynamic joint policy of teammates with attention multi-agent ddpg, in
2019
Later among the works it cites.
Russell, S. (2019), Human compatible: Artificial intelligence and the problem of control
2019
Later among the works it cites.
Yu, L., Song, J. & Ermon, S. (2019), Multi-agent adversarial inverse reinforcement learning, in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gregor, K., Rezende, D. J. & Wierstra, D. (2016), Variational intrinsic control, in
2016
Cited alongside, same era.
Guckelsberger, C., Salge, C. & Colton, S. (2016), Intrinsically motivated general companion NPCs via Coupled Empowerment Maximisation, in
2016
Cited alongside, same era.
Hadfield-Menell, D., Russell, S. J., Abbeel, P. & Dragan, A. (2016), Cooperative inverse reinforcement learning, in
2016
Cited alongside, same era.
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D. & Wierstra, D. (2016), Continuous control with deep reinforcement learning, in
2016
Cited alongside, same era.
Pellegrinelli, S., Admoni, H., Javdani, S. & Srinivasa, S. (2016), Human-robot shared workspace collaboration via hindsight optimization, in
2016
Cited alongside, same era.
Bostrom, N. (2017), Superintelligence
2017
Cited alongside, same era.
Fisac, J. F., Gates, M. A., Hamrick, J. B., Liu, C., Hadfield-Menell, D., Palaniappan, M., Malik, D., Sastry, S. S., Griffiths, T. L. & Dragan, A. D. (2017), ‘Pragmatic-pedagogic value alignment’
2017
Cited alongside, same era.
Christianos, F., Schäfer, L. & Albrecht, S. (2020), ‘Shared Experience Actor-Critic for Multi-Agent Reinforcement Learning’, Advances in Neural Information Processing Systems
2020
Later among the works it cites.
Du, Y., Tiomkin, S., Kiciman, E., Polani, D., Abbeel, P. & Dragan, A. (2020), ‘Ave: Assistance via empowerment’, Advances in Neural Information Processing Systems
2020
Later among the works it cites.
Fisac, J. F., Liu, C., Hamrick, J. B., Sastry, S., Hedrick, J. K., Griffiths, T. L. & Dragan, A. D. (2020), Generating plans that predict themselves, in
2020
Later among the works it cites.
Gabriel, I. (2020), ‘Artificial Intelligence, Values, and Alignment’, Minds and Machines
2020
Later among the works it cites.
Jeon, W., Barde, P., Nowrouzezahrai, D. & Pineau, J. (2020), ‘Scalable multi-agent inverse reinforcement learning via actor-attention-critic’
2020
Later among the works it cites.
Mutti, M., Pratissoli, L. & Restelli, M. (2020), ‘A Policy Gradient Method for Task-Agnostic Exploration’
2020
Later among the works it cites.
Papoudakis, G., Christianos, F. & Albrecht, S. V. (2020), ‘Opponent modelling with local information variational autoencoders’
2020
Later among the works it cites.
Terry, J. K., Jayakumar, M., Santos, L., Black, B., Hari, A., Dieffendahl, C., Williams, N. L., Lokesh, Y., Sullivan, R., Horsh, C. & Ravi, P. (2020), ‘Pettingzoo: Gym for multi-agent reinforcement learning’
2020
Later among the works it cites.
Volpi, N. C. & Polani, D. (2020), ‘Goal-Directed Empowerment: Combining Intrinsic Motivation and Task-Oriented Behaviour’, IEEE Transactions on Cognitive and Developmental Systems
2020
Later among the works it cites.
Zhao, R., Lu, K., Abbeel, P. & Tiomkin, S. (2020), Efficient empowerment estimation for unsupervised stabilization, in
2020
Later among the works it cites.
Arora, S. & Doshi, P. (2021), ‘A survey of inverse reinforcement learning: Challenges, methods and progress’, Artificial Intelligence
2021
Closest in time.