Fetching the paper…
Reading the bibliography…
Framed in positive terms, this report examines how technical AI research might be steered in a manner that is more attentive to humanity's long-term prospects for survival as a species.
Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning
Milli, S. and A. D. Dragan (2019) · 1903
Earlier work this paper cites.
Theory of games and economic behavior
Morgenstern, O. and J. Von Neumann (1953) · 1953
Earlier work this paper cites.
Sleeping beauty and self-location: A hybrid model
Bostrom, N. (2007) · 1955
Earlier work this paper cites.
Solution of a problem of leon henkin 1
Löb, M. H. (1955) · 1955
Earlier work this paper cites.
Some moral and technical consequences of automation
Wiener, N. (1960) · 1960
Earlier work this paper cites.
When is a linear control system optimal?
Kalman, R. E. (1964) · 1964
Earlier work this paper cites.
Groupthink
Janis, I. L. (1971) · 1971
Earlier work this paper cites.
Incentives in teams
Groves, T. (1973) · 1973
Earlier work this paper cites.
Agreeing to disagree
Aumann, R. J. (1976) · 1976
Earlier work this paper cites.
Cardinal welfare, individualistic ethics, and interpersonal comparisons of utility
Harsanyi, J. C. (1980) · 1980
Earlier work this paper cites.
Normal accidents: Living with high-risk technologies
Perrow, C. (1984) · 1984
Earlier work this paper cites.
Mind children: The future of robot and human intelligence
Moravec, H. (1988) · 1988
Earlier work this paper cites.
New challenges in organizational research: high reliability organizations
Roberts, K. H. (1989) · 1989
Earlier work this paper cites.
Informal organizational networking as a crisis-avoidance strategy: Us naval flight operations as a case study
Rochlin, G. I. (1989) · 1989
Earlier work this paper cites.
Some characteristics of one type of high reliability organization
Roberts, K. H. (1990) · 1990
Earlier work this paper cites.
Groupthink in government: A study of small groups and policy failure
Hart, P. (1990) · 1991
Earlier work this paper cites.
Human factors in large-scale technological systems’ accidents: Three Mile Island, Bhopal, Chernobyl
Meshkati, N. (1991) · 1991
Earlier work this paper cites.
Do the right thing
Russell, S. and E. Wefald (1991) · 1991
Earlier work this paper cites.
How do we know we have global environmental problems? Science and the globalization of environmental discourse
Taylor, P. J. and F. H. Buttel (1992) · 1992
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and G. E. Hinton (1993) · 1993
Earlier work this paper cites.
Hierarchical learning in stochastic domains: Preliminary results
Kaelbling, L. P. (1993) · 1993
Earlier work this paper cites.
The negotiated order of organizational reliability
Schulman, P. R. (1993) · 1993
Earlier work this paper cites.
Decision dynamics in two high reliability military organizations
Roberts, K. H., S. K. Stout, and J. J. Halpern (1994) · 1994
Earlier work this paper cites.
Organizational culture in high reliability organizations: An extension
Klein, R. L., G. A. Bigley, and K. H. Roberts (1995) · 1995
Earlier work this paper cites.
Regulatory compliance and the ethos of quality enhancement: Surprises in nuclear power plant operations1
LaPorte, T. R. and C. W. Thomas (1995) · 1995
Earlier work this paper cites.
Understanding the intentions of others: Re-enactment of intended acts by 18-month-old children
Meltzoff, A. N. (1995) · 1995
Earlier work this paper cites.
Organizing maintenance work at two american nuclear power plants
Bourrier, M. (1996) · 1996
Earlier work this paper cites.
Historical and technical notes on aqueducts from prehistoric to medieval times
De Feo, G., A. Angelakis, G. Antoniou, F. El-Gohary, B. Haut, C. Passchier, and X. Zheng (2013) · 1996
Earlier work this paper cites.
High reliability organizations: Unlikely, demanding and at risk
LaPorte, T. R. (1996) · 1996
Earlier work this paper cites.
The Challenger launch decision: Risky technology, culture, and deviance at NASA
Vaughan, D. (1996) · 1996
Earlier work this paper cites.
The voluntary provision of a pure public good: The case of reduced CFC emissions and the Montreal Protocol
Murdoch, J. C. and T. Sandler (1997) · 1997
Earlier work this paper cites.
On the interpretation of decision problems with imperfect recall
Piccione, M. and A. Rubinstein (1997) · 1997
Earlier work this paper cites.
Hq-learning
Wiering, M. and J. Schmidhuber (1997) · 1997
Earlier work this paper cites.
How long before superintelligence?
Bostrom, N. (1998) · 1998
Earlier work this paper cites.
Alive and well after 25 years: A review of groupthink research
Esser, J. K. (1998) · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and S. J. Russell (1998) · 1998
Earlier work this paper cites.
Science’s new social contract with society
Gibbons, M. (1999) · 1999
Earlier work this paper cites.
Normal Accidents: Living with High Risk Technologies
Perrow, C. (1999) · 1999
Earlier work this paper cites.
Safe operation as a social construct
Rochlin, G. I. (1999) · 1999
Earlier work this paper cites.
Team errors: definition and taxonomy
Sasou, K. and J. Reason (1999) · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., D. Precup, and S. Singh (1999) · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, T. G. (2000) · 2000
Earlier work this paper cites.
The incident command system: High-reliability organizing for complex and volatile task environments
Bigley, G. A. and K. H. Roberts (2001) · 2001
Earlier work this paper cites.
Secure multi-party computation problems and their applications: a review and open problems
Du, W. and M. J. Atallah (2001) · 2001
Earlier work this paper cites.
Making low probabilities useful
Kunreuther, H., N. Novemsky, and D. Kahneman (2001) · 2001
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Bernstein, D. S., R. Givan, N. Immerman, and S. Zilberstein (2002) · 2002
Earlier work this paper cites.
A POMDP formulation of preference elicitation problems
Boutilier, C. (2002) · 2002
Earlier work this paper cites.
Considered opinions: Deliberative polling in Britain
Luskin, R. C., J. S. Fishkin, and R. Jowell (2002) · 2002
Earlier work this paper cites.
The social contract: And, the first and second discourses
Rousseau, J.-J. and G. May (2002) · 2002
Earlier work this paper cites.
Automatic upgrade of live network devices
San Martin, R., M. Werner, S. Bill, and B. Stacey (2002, July 4) · 2002
Earlier work this paper cites.
Artificial intelligence: a modern approach
Russell, S., P. Norvig, J. F. Canny, J. M. Malik, and D. D. Edwards (2003) · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and A. Y. Ng (2004) · 2004
Earlier work this paper cites.
Social contract theory
Friend, C. (2004) · 2004
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Hutter, M. (2004) · 2004
Earlier work this paper cites.
A critical look at risk assessments for global catastrophes
Kent, A. (2004) · 2004
Earlier work this paper cites.
High reliability and the management of critical infrastructures
Schulman, P., E. Roe, M. v. Eeten, and M. d. Bruijne (2004) · 2004
Earlier work this paper cites.
Program equilibrium
Tennenholtz, M. (2004) · 2004
Earlier work this paper cites.
Intelligent Machinery, A Heretical Theory (c.1951). Reprinted in The Essential Turing , by B. Jack Copeland., 2004
Turing, A. (1951) · 2004
Earlier work this paper cites.
An explainable artificial intelligence system for small-unit tactical behavior
Van Lent, M., W. Fisher, and M. Mancuso (2004) · 2004
Earlier work this paper cites.
The complexity of agreement
Aaronson, S. (2005) · 2005
Earlier work this paper cites.
Node shutdown in clustered computer system
Block, T. R., R. Miller, and K. Thayib (2005, July 12) · 2005
Earlier work this paper cites.
Toward a strategic human resource management model of high reliability organization performance
Ericksen, J. and L. Dyer (2005) · 2005
Earlier work this paper cites.
Experimenting with a democratic ideal: Deliberative polling and public opinion
Fishkin, J. S. and R. C. Luskin (2005) · 2005
Earlier work this paper cites.
A Bayesian view of language evolution by iterated learning
Griffiths, T. L. and M. L. Kalish (2005) · 2005
Earlier work this paper cites.
Human-level artificial intelligence? Be serious!
Nilsson, N. J. (2005) · 2005
Earlier work this paper cites.
Bhopal: Anatomy of a crisis
Shrivastava, P. (1992) · 2005
Earlier work this paper cites.
Computational approaches to preference elicitation
Braziunas, D. (2006) · 2006
Earlier work this paper cites.
Method and system for providing emergency shutdown of a malfunctioning device
Litwin Jr, L. R. and K. Ramaswamy (2006, July 4) · 2006
Earlier work this paper cites.
Probabilistic inference in human semantic memory
Steyvers, M., T. L. Griffiths, and S. Dennis (2006) · 2006
Earlier work this paper cites.
Motivated skepticism in the evaluation of political beliefs
Taber, C. S. and M. Lodge (2006) · 2006
Earlier work this paper cites.
Interpretable convolutional neural networks
Zhang, Q., Y. Nian Wu, and S.-C. Zhu (2018) · 2006
Earlier work this paper cites.
Program obfuscation: a quantitative approach
Anckaert, B., M. Madou, B. De Sutter, B. De Bus, K. De Bosschere, and B. Preneel (2007) · 2007
Earlier work this paper cites.
Artificial general intelligence
Goertzel, B. and C. Pennachin (2007) · 2007
Earlier work this paper cites.
Simplicity and probability in causal explanation
Lombrozo, T. (2007) · 2007
Earlier work this paper cites.
Reducing the risk of human extinction
Matheny, J. G. (2007) · 2007
Earlier work this paper cites.
Theory of games and economic behavior (commemorative edition)
Von Neumann, J. and O. Morgenstern (2007) · 2007
Earlier work this paper cites.
Application of machine learning techniques for supply chain demand forecasting
Carbonneau, R., K. Laframboise, and R. Vahidov (2008) · 2008
Earlier work this paper cites.
Theoretical and empirical evidence for the impact of inductive biases on cultural evolution
Griffiths, T. L., M. L. Kalish, and S. Lewandowsky (2008) · 2008
Earlier work this paper cites.
Avoiding another AI winter
Hendler, J. (2008) · 2008
Earlier work this paper cites.
Technique for graceful shutdown of a routing protocol in a network
Scudder, J. G., M. Sivabalan, and D. D. Ward (2008, April 8) · 2008
Earlier work this paper cites.
Academic search engine optimization (ASEO) optimizing scholarly literature for Google Scholar & co
Beel, J., B. Gipp, and E. Wilde (2009) · 2009
Earlier work this paper cites.
The wisdom of individuals: Exploring people’s knowledge about everyday events using iterated learning
Lewandowsky, S., T. L. Griffiths, and M. L. Kalish (2009) · 2009
Earlier work this paper cites.
Electronic media use and sleep in school-aged children and adolescents: A review
Cain, N. and M. Gradisar (2010) · 2010
Earlier work this paper cites.
A computational decision theory for interactive assistants
Fern, A. and P. Tadepalli (2010) · 2010
Earlier work this paper cites.
Almost common priors
Hellman, Z. (2013) · 2010
Earlier work this paper cites.
Unmanned aircraft systems
Hobbs, A. (2010) · 2010
Earlier work this paper cites.
Causal–explanatory pluralism: How intentions, functions, and mechanisms influence causal ascriptions
Lombrozo, T. (2010) · 2010
Earlier work this paper cites.
Omohundro’s “basic AI drives” and catastrophic risks
Shulman, C. (2010) · 2010
Earlier work this paper cites.
Health effects of media on children and adolescents
Strasburger, V. C., A. B. Jordan, and E. Donnerstein (2010) · 2010
Cited alongside, same era.
Bayesian theory of mind: Modeling joint belief-desire attribution
Baker, C., R. Saxe, and J. Tenenbaum (2011) · 2011
Cited alongside, same era.
Program obfuscation with leaky hardware
Bitansky, N., R. Canetti, S. Goldwasser, S. Halevi, Y. Kalai, and G. Rothblum (2011) · 2011
Cited alongside, same era.
Information hazards: a typology of potential harms from knowledge
Bostrom, N. et al. (2011) · 2011
Cited alongside, same era.
Differential privacy
Dwork, C. (2011) · 2011
Cited alongside, same era.
The microstructure of the flash crash: flow toxicity, liquidity crashes, and the probability of informed trading
Easley, D., M. M. L. De Prado, and M. O’Hara (2011) · 2011
Cited alongside, same era.
Federated learning: Strategies for improving communication efficiency
Konečnỳ, J., H. B. McMahan, F. X. Yu, P. Richtárik, A. T. Suresh, and D. Bacon (2016) · 2016
Later among the works it cites.
Active reinforcement learning: Observing rewards at a cost
Krueger, D., J. Leike, O. Evans, and J. Salvatier (2016) · 2016
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D., K. Narasimhan, A. Saeedi, and J. Tenenbaum (2016) · 2016
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., E. Moore, D. Ramage, S. Hampson, et al. (2016) · 2016
Later among the works it cites.
Formalizing human-robot mutual adaptation: A bounded memory model
Nikolaidis, S., A. Kuznetsov, D. Hsu, and S. Srinivasa (2016) · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I don’t want to think about it now: Decision theory with costly computation
Halpern, J. Y. and R. Pass (2011) · 2011
Cited alongside, same era.
Why the future doesn’t need us
Joy, B. (2011) · 2011
Cited alongside, same era.
The politicization of climate change and polarization in the american public’s views of global warming, 2001–2010
McCright, A. M. and R. E. Dunlap (2011) · 2011
Cited alongside, same era.
Protecting the ozone layer: the United Nations history
Andersen, S. O. and K. M. Sarma (2012) · 2012
Cited alongside, same era.
Groupthink: Collective delusions in organizations and markets
Bénabou, R. (2012) · 2012
Cited alongside, same era.
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Bostrom, N. (2012) · 2012
Cited alongside, same era.
A concise introduction to decentralized POMDPs
Oliehoek, F. A., C. Amato, et al. (2016) · 2016
Later among the works it cites.
“Why Should I Trust You?”: Explaining the Predictions of Any Classifier
Ribeiro, M. T., S. Singh, and C. Guestrin (2016) · 2016
Later among the works it cites.
Motor learning affects car-to-driver handover in automated vehicles
Russell, H. E., L. K. Harbott, I. Nisky, S. Pan, A. M. Okamura, and J. C. Gerdes (2016) · 2016
Later among the works it cites.
Learning multiagent communication with backpropagation
Sukhbaatar, S., R. Fergus, et al. (2016) · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
Szegedy, C., V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna (2016) · 2016
Later among the works it cites.
Alignment for advanced machine learning systems
Taylor, J., E. Yudkowsky, P. LaVictoire, and A. Critch (2016) · 2016
Later among the works it cites.
Strategic attentive writer for learning macro-actions
Vezhnevets, A., V. Mnih, S. Osindero, A. Graves, O. Vinyals, J. Agapiou, et al. (2016) · 2016
Later among the works it cites.
Agent-agnostic human-in-the-loop reinforcement learning
Abel, D., J. Salvatier, A. Stuhlmüller, and O. Evans (2017) · 2017
Later among the works it cites.
Indifference methods for managing agent rewards
Armstrong, S. and X. O’Rourke (2017) · 2017
Later among the works it cites.
Synthesizing robust adversarial examples
Athalye, A., L. Engstrom, A. Ilyas, and K. Kwok (2017) · 2017
Later among the works it cites.
The option-critic architecture
Bacon, P.-L., J. Harb, and D. Precup (2017) · 2017
Later among the works it cites.
Emergent complexity via multi-agent competition
Bansal, T., J. Pachocki, S. Sidor, I. Sutskever, and I. Mordatch (2017) · 2017
Later among the works it cites.
The trouble with autopilots: assisted and autonomous driving on the social road
Brown, B. and E. Laurier (2017) · 2017
Later among the works it cites.
Incorrigibility in the CIRL framework
Carey, R. (2017) · 2017
Later among the works it cites.
Humans consulting HCH (HCH)
Christiano, P. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning from human preferences
Christiano, P. F., J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei (2017) · 2017
Later among the works it cites.
Offer: Off-environment reinforcement learning
Ciosek, K. A. and S. Whiteson (2017) · 2017
Later among the works it cites.
The ethical knob: ethically-customisable automated vehicles and the law
Contissa, G., F. Lagioia, and G. Sartor (2017) · 2017
Later among the works it cites.
A parametric, resource-bounded generalization of Löb’s theorem, and a robust cooperation criterion for open-source game theory
Critch, A. (2019) · 2017
Later among the works it cites.
Modeling Agents with Probabilistic Programs
Evans, O., A. Stuhlmüller, J. Salvatier, and D. Filan (2017) · 2017
Later among the works it cites.
Pragmatic-pedagogic value alignment
Fisac, J. F., M. A. Gates, J. B. Hamrick, C. Liu, D. Hadfield-Menell, M. Palaniappan, D. Malik, S. S. Sastry, T. L. Griffiths, and A. D. Dragan (2017) · 2017
Later among the works it cites.
Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (woebot): A randomized controlled trial
Fitzpatrick, K. K., A. Darcy, and M. Vierhile (2017) · 2017
Later among the works it cites.
Stabilising experience replay for deep multi-agent reinforcement learning
Foerster, J., N. Nardelli, G. Farquhar, T. Afouras, P. H. Torr, P. Kohli, and S. Whiteson (2017) · 2017
Later among the works it cites.
The future of employment: how susceptible are jobs to computerisation?
Frey, C. B. and M. A. Osborne (2017) · 2017
Later among the works it cites.
On calibration of modern neural networks
Guo, C., G. Pleiss, Y. Sun, and K. Q. Weinberger (2017) · 2017
Later among the works it cites.
Hadfield-Menell, D., A. Dragan, P. Abbeel, and S. Russell (2016b) · 2017
Later among the works it cites.
Reluplex: An efficient smt solver for verifying deep neural networks
Katz, G., C. Barrett, D. L. Dill, K. Julian, and M. J. Kochenderfer (2017) · 2017
Later among the works it cites.
What uncertainties do we need in Bayesian deep learning for computer vision?
Kendall, A. and Y. Gal (2017) · 2017
Later among the works it cites.
Avoiding discrimination through causal reasoning
Kilbertus, N., M. R. Carulla, G. Parascandolo, M. Hardt, D. Janzing, and B. Schölkopf (2017) · 2017
Later among the works it cites.
The flash crash: High-frequency trading in an electronic market
Kirilenko, A., A. S. Kyle, M. Samadi, and T. Tuzun (2017) · 2017
Later among the works it cites.
Federated control with hierarchical multi-agent deep reinforcement learning
Kumar, S., P. Shah, D. Hakkani-Tur, and L. Heck (2017) · 2017
Later among the works it cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., A. Pritzel, and C. Blundell (2017) · 2017
Later among the works it cites.
Dart: Noise injection for robust imitation learning
Laskey, M., J. Lee, R. Fox, A. Dragan, and K. Goldberg (2017) · 2017
Later among the works it cites.
A rawlsian algorithm for autonomous vehicles
Leben, D. (2017) · 2017
Later among the works it cites.
Training confidence-calibrated classifiers for detecting out-of-distribution samples
Lee, K., H. Lee, K. Lee, and J. Shin (2017) · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Leibo, J. Z., V. Zambaldi, M. Lanctot, J. Marecki, and T. Graepel (2017) · 2017
Later among the works it cites.
Leike, J., M. Martic, V. Krakovna, P. A. Ortega, T. Everitt, A. Lefrancq, L. Orseau, and S. Legg (2017) · 2017
Later among the works it cites.
Deal or no deal? End-to-end learning for negotiation dialogues
Lewis, M., D. Yarats, Y. N. Dauphin, D. Parikh, and D. Batra (2017) · 2017
Later among the works it cites.
Enhancing the reliability of out-of-distribution image detection in neural networks
Liang, S., Y. Li, and R. Srikant (2017) · 2017
Later among the works it cites.
An approach to reachability analysis for feed-forward relu neural networks
Lomuscio, A. and L. Maganti (2017) · 2017
Later among the works it cites.
Milli, S., D. Hadfield-Menell, A. Dragan, and S. Russell (2017) · 2017
Later among the works it cites.
Feature visualization
Olah, C., A. Mordvintsev, and L. Schubert (2017) · 2017
Later among the works it cites.
Automatic virtual machine termination in a cloud
Rigolet, J.-Y. B. (2017, January 31) · 2017
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W., G. Sastry, A. Stuhlmueller, and O. Evans (2017) · 2017
Later among the works it cites.
Developing bug-free machine learning systems with formal mathematics
Selsam, D., P. Liang, and D. L. Dill (2017) · 2017
Later among the works it cites.
Federated multi-task learning
Smith, V., C.-K. Chiang, M. Sanjabi, and A. S. Talwalkar (2017) · 2017
Later among the works it cites.
One pixel attack for fooling deep neural networks
Su, J., D. V. Vargas, and S. Kouichi (2017) · 2017
Later among the works it cites.
Multiagent cooperation and competition with deep reinforcement learning
Tampuu, A., T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente (2017) · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor (2017) · 2017
Later among the works it cites.
Reachability analysis for neural agent-environment systems
Akintunde, M., A. Lomuscio, L. Maganti, and E. Pirovano (2018) · 2018
Later among the works it cites.
Fairness, Accountability, and Transparency in Machine Learning (workshop)
Barocas, S. and M. Hardt (2014) · 2018
Later among the works it cites.
The vulnerable world hypothesis
Bostrom, N. (2018) · 2018
Later among the works it cites.
Servant of many masters: Shifting priorities in Pareto-optimal sequential decision-making
Critch, A. and S. Russell (2017) · 2018
Later among the works it cites.
Learning confidence for out-of-distribution detection in neural networks
DeVries, T. and G. W. Taylor (2018) · 2018
Later among the works it cites.
A dual approach to scalable verification of deep networks
Dvijotham, K., R. Stanforth, S. Gowal, T. Mann, and P. Kohli (2018) · 2018
Later among the works it cites.
Universal artificial intelligence
Everitt, T. and M. Hutter (2018) · 2018
Later among the works it cites.
Learning with opponent-learning awareness
Foerster, J., R. Y. Chen, M. Al-Shedivat, S. Whiteson, P. Abbeel, and I. Mordatch (2018) · 2018
Later among the works it cites.
When will AI exceed human performance? Evidence from AI Experts
Grace, K., J. Salvatier, A. Dafoe, B. Zhang, and O. Evans (2018) · 2018
Later among the works it cites.
Reliable uncertainty estimates in deep neural networks using noise contrastive priors
Hafner, D., D. Tran, A. Irpan, T. Lillicrap, and J. Davidson (2018) · 2018
Later among the works it cites.
It’s time to do something: Mitigating the negative impacts of computing through a change to the peer review process
Hecht, B., L. Wilcox, J. Bigham, J. Schöning, E. Hoque, J. Ernst, Y. Bisk, L. De Russis, L. Yarosh, B. Anjum, D. Contractor, and C. Wu (2018) · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei (2018) · 2018
Later among the works it cites.
Penalizing side effects using stepwise relative reachability
Krakovna, V., L. Orseau, R. Kumar, M. Martic, and S. Legg (2018) · 2018
Later among the works it cites.
Accurate uncertainties for deep learning using calibrated regression
Kuleshov, V., N. Fenner, and S. Ermon (2018) · 2018
Later among the works it cites.
Software verification with itps should use binary code extraction to reduce the tcb
Kumar, R., E. Mullen, Z. Tatlock, and M. O. Myreen (2018) · 2018
Later among the works it cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg (2018) · 2018
Later among the works it cites.
Functional explanation and the function of explanation
Lombrozo, T. and S. Carey (2006) · 2018
Later among the works it cites.
Categorizing variants of goodhart’s law
Manheim, D. and S. Garrabrant (2018) · 2018
Later among the works it cites.
Scaling shared model governance via model splitting
Martic, M., J. Leike, A. Trask, M. Hessel, S. Legg, and P. Kohli (2018) · 2018
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Mordatch, I. and P. Abbeel (2018) · 2018
Later among the works it cites.
Œuf: minimizing the Coq extraction TCB
Mullen, E., S. Pernsteiner, J. R. Wilcox, Z. Tatlock, and D. Grossman (2018) · 2018
Later among the works it cites.
Society-in-the-loop: programming the algorithmic social contract
Rahwan, I. (2018) · 2018
Later among the works it cites.
Where do you think you’re going?: Inferring beliefs about dynamics from behavior
Reddy, S., A. Dragan, and S. Levine (2018) · 2018
Later among the works it cites.
Bridging near-and long-term concerns about ai
Cave, S. and S. S. ÓhÉigeartaigh (2019) · 2019
Later among the works it cites.
Robust artificial intelligence and robust human organizations
Dietterich, T. G. (2019) · 2019
Later among the works it cites.
Hierarchical game-theoretic planning for autonomous vehicles
Fisac, J. F., E. Bronstein, E. Stefansson, D. Sadigh, S. S. Sastry, and A. D. Dragan (2019) · 2019
Later among the works it cites.
Some Background on Our Views Regarding Advanced Artificial Intelligence
Karnovsky, H. (2016) · 2019
Later among the works it cites.
Decomposing deliberation
Ought.org (2017a) · 2019
Later among the works it cites.
Predicting slow judgements
Ought.org (2017b) · 2019
Later among the works it cites.
In memoryless Cartesian environments, every UDT policy is a CDT+SIA policy
Taylor, J. (2016a) · 2019
Later among the works it cites.
Thoughts on human models
Taylor, J. (2016c) · 2019
Later among the works it cites.
Laws and regulations that apply to your agricultural operation by statute
United States Environmental Protection Agency (EPA) (2019) · 2019
Later among the works it cites.
Superintelligence skepticism as a political tool
Baum, S. (2018) · 2078
Closest in time.