Fetching the paper…
Reading the bibliography…
AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A., and Terry, M. E · 1952
Earlier work this paper cites.
On the folly of rewarding A, while hoping for B
Kerr, S · 1975
Earlier work this paper cites.
Markets and hierarchies
Williamson, O. E · 1975
Earlier work this paper cites.
Vertical integration, appropriable rents, and the competitive contracting process
Klein, B., Crawford, R. G., and Alchian, A. A · 1978
Earlier work this paper cites.
Damage measures for breach of contract
Shavell, S · 1980
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
A POMDP formulation of preference elicitation problems
Boutilier, C · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P., and Ng, A. Y · 2004
Earlier work this paper cites.
Preference elicitation and generalized additive utility
Braziunas, D., and Boutilier, C · 2006
Earlier work this paper cites.
The basic AI drives
Omohundro, S. M · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Russell, S. J., and Norvig, P · 2010
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Griffith, S., Subramanian, K., Scholz, J., Isbell, C. L., and Thomaz, A. L · 2013
Cited alongside, same era.
Power to the people: The role of humans in interactive machine learning
Amershi, S., Cakmak, M., Knox, W. B., and Kulesza, T · 2014
Cited alongside, same era.
Superintelligence: Paths, dangers, strategies
Bostrom, N · 2014
Cited alongside, same era.
Faulty reward functions in the wild, 2016
Amodei, D., and Clark, J · 2016
Cited alongside, same era.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A · 2016
Eckersley, P · 2018
Later among the works it cites.
Penalizing side effects using stepwise relative reachability
Krakovna, V., Orseau, L., Kumar, R., Martic, M., and Legg, S · 2018
Later among the works it cites.
Categorizing variants of Goodhart’s law
Manheim, D., and Garrabrant, S · 2018
Later among the works it cites.
Towards a just theory of measurement: A principled social measurement assurance program for machine learning
Andrus, M., and Gilbert, T. K · 2019
Later among the works it cites.
The assistive multi-armed bandit
Chan, L., Hadfield-Menell, D., Srinivasa, S., and Dragan, A. D · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Low impact artificial intelligence
Armstrong, S., and Levinstein, B · 2017
Cited alongside, same era.
Learning robot objectives from physical human interaction
Bajcsy, A., Losey, D. P., O’Malley, M. K., and Dragan, A. D · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences, 2017
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
The off-switch game
Hadfield-Menell, D., Dragan, A., Abbeel, P., and Russell, S · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A · 2017
Cited alongside, same era.
Leike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S · 2017
Cited alongside, same era.
Incomplete contracting and AI alignment
Hadfield-Menell, D., and Hadfield, G. K · 2019
Later among the works it cites.
Jacobs, A. Z., and Wallach, H · 2019
Later among the works it cites.
Human Compatible: Artificial Intelligence and the Problem of Control
Russell, S. J · 2019
Later among the works it cites.
Youtube’s recommendation algorithm has a dark side
Tufekci, Z · 2019
Later among the works it cites.
Inside Twitter’s ambitious plan to change the way we tweet
Wagner, K · 2019
Later among the works it cites.
What are you optimizing for? aligning recommender systems with human values
Stray, J., Adler, S., and Hadfield-Menell, D · 2020
Later among the works it cites.
The problem with metrics is a fundamental problem for AI
Thomas, R., and Uminsky, D · 2020
Later among the works it cites.