Fetching the paper…
Reading the bibliography…
For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want.
Understanding agent incentives using causal influence diagrams. part i: Single action settings
T. Everitt, P. A. Ortega, E. Barnes, and S. Legg · 1902
Earlier work this paper cites.
Modeling AGI safety frameworks with causal influence diagrams
T. Everitt, R. Kumar, V. Krakovna, and S. Legg · 1906
Earlier work this paper cites.
Eliza—a computer program for the study of natural language communication between man and machine
J. Weizenbaum · 1966
Earlier work this paper cites.
On the folly of rewarding A, while hoping for B
S. Kerr · 1975
Earlier work this paper cites.
Animal signals: information or manipulation
R. Dawkins and J. R. Krebs · 1978
Earlier work this paper cites.
Manipulation
J. Rudinow · 1978
Earlier work this paper cites.
Problems of monetary management: the UK experience
C. A. Goodhart · 1984
Earlier work this paper cites.
Animal signals: mind-reading and manipulation
J. R. Krebs · 1984
Earlier work this paper cites.
The morality of freedom
J. Raz · 1986
Earlier work this paper cites.
The intentional stance
D. C. Dennett · 1989
Earlier work this paper cites.
Dishonesty and the handicap principle
R. A. Johnstone and A. Grafen · 1993
Earlier work this paper cites.
The evolution of communication
M. D. Hauser · 1996
Earlier work this paper cites.
Manipulative actions: a conceptual and moral analysis
R. Noggle · 1996
Earlier work this paper cites.
Evolutionarily stable strategies of age-dependent sexual advertisement
H. Kokko · 1997
Earlier work this paper cites.
‘improving ratings’: audit in the british university system
M. Strathern · 1997
Earlier work this paper cites.
Theory of mind and the evolution of language
R. Dunbar et al · 1998
Earlier work this paper cites.
The evolution of animal communication: reliability and deception in signaling systems
W. A. Searcy and S. Nowicki · 2005
Earlier work this paper cites.
Why talk? speaking as selfish behaviour
T. Scott-Phillips · 2006
Earlier work this paper cites.
Anthropomorphism and AI: Turing’s much misunderstood imitation game
D. Proudfoot · 2011
Earlier work this paper cites.
Thinking inside the box: Controlling and using an oracle AI
S. Armstrong, A. Sandberg, and N. Bostrom · 2012
Earlier work this paper cites.
Between reason and coercion: ethically permissible influence in health care and health policy contexts
J. S. Blumenthal-Barby · 2012
Earlier work this paper cites.
Automatic deception detection in italian court cases
T. Fornaciari and M. Poesio · 2013
Earlier work this paper cites.
What is manipulation
A. Barnhill · 2014
Earlier work this paper cites.
The mens rea and moral status of manipulation
M. Baron · 2014
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
N. Bostrom · 2014
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
K. Narasimhan, T. Kulkarni, and R. Barzilay · 2015
Earlier work this paper cites.
Deception detection using real-life trial data
V. Pérez-Rosas, M. Abouelenien, R. Mihalcea, and M. Burzo · 2015
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, A. Dragan, P. Abbeel, and S. Russell · 2016
Earlier work this paper cites.
The Definition of Lying and Deception
J. E. Mahon · 2016
Earlier work this paper cites.
Deception as a derived function of language
N. Oesch · 2016
Cited alongside, same era.
Safely interruptible agents
L. Orseau and S. Armstrong · 2016
Cited alongside, same era.
Good and safe uses of AI oracles
S. Armstrong and X. O’Rorke · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
The future of employment: How susceptible are jobs to computerisation?
C. B. Frey and M. A. Osborne · 2017
Cited alongside, same era.
J. Leike, M. Martic, V. Krakovna, P. A. Ortega, T. Everitt, A. Lefrancq, L. Orseau, and S. Legg · 2017
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Later among the works it cites.
Clarifying AI alignment, AI alignment forum
P. Christiano · 2020
Later among the works it cites.
Faulty reward functions in the wild
J. Clark and D. Amodei · 2020
Later among the works it cites.
Perspective api
Conversation-AI · 2020
Later among the works it cites.
Underspecification presents challenges for credibility in modern machine learning
A. D’Amour, K. Heller, D. Moldovan, B. Adlam, B. Alipanahi, A. Beutel, C. Chen, J. Deaton, J. Eisenstein, M. D. Hoffman, et al · 2020
Later among the works it cites.
A survey of the state of explainable AI for natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deal or no deal? end-to-end learning for negotiation dialogues
M. Lewis, D. Yarats, Y. N. Dauphin, D. Parikh, and D. Batra · 2017
Cited alongside, same era.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
J. Zhao, T. Wang, M. Yatskar, V. Ordonez, and K.-W. Chang · 2017
Cited alongside, same era.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
M. Brundage, S. Avin, J. Clark, H. Toner, P. Eckersley, B. Garfinkel, A. Dafoe, P. Scharre, T. Zeitzoff, B. Filar, et al · 2018
Cited alongside, same era.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Cited alongside, same era.
T. Everitt, G. Lea, and M. Hutter · 2018
Cited alongside, same era.
Ethical challenges in data-driven dialogue systems
P. Henderson, K. Sinha, N. Angelard-Gontier, N. R. Ke, G. Fried, R. Lowe, and J. Pineau · 2018
Cited alongside, same era.
M. Danilevsky, K. Qian, R. Aharonov, Y. Katsis, B. Kawas, and P. Sen · 2020
Later among the works it cites.
Artificial intelligence, values and alignment
I. Gabriel · 2020
Later among the works it cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Later among the works it cites.
The state and fate of linguistic diversity and inclusion in the NLP world
P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury · 2020
Later among the works it cites.
Specification gaming: the flip side of AI ingenuity
V. Krakovna, J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg · 2020
Later among the works it cites.
The surprising creativity of digital evolution: A collection of anecdotes from the evolutionary computation and artificial life research communities
J. Lehman, J. Clune, D. Misevic, C. Adami, L. Altenberg, J. Beaulieu, P. J. Bentley, S. Bernard, G. Beslon, D. M. Bryson, et al · 2020
Later among the works it cites.
Gender bias in neural natural language processing
K. Lu, P. Mardziel, F. Wu, P. Amancharla, and A. Datta · 2020
Later among the works it cites.
Stereoset: Measuring stereotypical bias in pretrained language models
M. Nadeem, A. Bethke, and S. Reddy · 2020
Later among the works it cites.
The Ethics of Manipulation
R. Noggle · 2020
Later among the works it cites.
Building safe artificial intelligence: specification, robustness and assurance
P. Ortega and V. Maini · 2020
Later among the works it cites.
AI deception: When your artificial intelligence learns to lie
H. Roff · 2020
Later among the works it cites.
Comment on clarifying AI alignment, AI alignment forum
R. Shah · 2020
Later among the works it cites.
Interpretability in ML: A broad overview
O. Shen · 2020
Later among the works it cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Later among the works it cites.
The AI-box experiment
E. Yudkowsky · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
E. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Closest in time.
Agent incentives: A causal perspective
T. Everitt, R. Carey, E. Langlois, P. A. Ortega, and S. Legg · 2021
Closest in time.
Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
T. Everitt, M. Hutter, R. Kumar, and V. Krakovna · 2021
Closest in time.
Relaxed adversarial training for inner alignment
E. Hubinger · 2021
Closest in time.
Giving GPT-3 a turing test
K. Lacker · 2021
Closest in time.
How RL agents behave when their actions are modified
E. D. Langlois and T. Everitt · 2021
Closest in time.
Doctor GPT-3: hype or reality?
A.-L. Rousseau, C. Baudelaire, and K. Riera · 2021
Closest in time.
Teaching GPT-3 to identify nonsense
A. Sabeti · 2021
Closest in time.
Understanding the capabilities, limitations, and societal impact of large language models
A. Tamkin, M. Brundage, J. Clark, and D. Ganguli · 2021
Closest in time.