Fetching the paper…
Reading the bibliography…
The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wants.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Statistical Learning Theory
V. N. Vapnik · 1998
Earlier work this paper cites.
Creating friendly ai 1.0: The analysis and design of benevolent goal architectures
E. Yudkowsky · 2001
Earlier work this paper cites.
The basic ai drives
S. M. Omohundro · 2008
Earlier work this paper cites.
Thinking inside the box: Controlling and using an oracle ai
S. Armstrong, A. Sandberg, and N. Bostrom · 2012
Earlier work this paper cites.
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
N. Bostrom · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Cited alongside, same era.
Aligning superintelligence with human interests: A technical research agenda
N. Soares and B. Fallenstein · 2014
Cited alongside, same era.
Taking superintelligence seriously: Superintelligence: Paths, dangers, strategies by nick bostrom (oxford university press, 2014)
M. Brundage · 2015
Cited alongside, same era.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Cited alongside, same era.
Formalizing convergent instrumental goals
T. Benson-Tilsen and N. Soares · 2016
Cited alongside, same era.
Safely interruptible agents
L. Orseau and M. Armstrong · 2016
Alignment for advanced machine learning systems
J. Taylor, E. Yudkowsky, P. LaVictoire, and A. Critch · 2016
Later among the works it cites.
Ai alignment: why its hard and where to start
E. Yudkowsky · 2016
Later among the works it cites.
Failures of gradient-based deep learning
S. Shalev-Shwartz, O. Shamir, and S. Shammah · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Later among the works it cites.
Why do deep convolutional networks generalize so poorly to small image transformations?
A. Azulay and Y. Weiss · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Should we fear supersmart robots?
S. Russell · 2016
Cited alongside, same era.
On the sample complexity of end-to-end training vs. semantic abstraction training
S. Shalev-Shwartz and A. Shashua · 2016
Cited alongside, same era.
Ai alignment: why its hard and where to start
Y. N. Harari and F. fei Li
Cited in the paper.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Later among the works it cites.
Deep learning: A critical appraisal
G. Marcus · 2018
Later among the works it cites.
Towards a human-like open-domain chatbot
D. Adiwardana, M.-T. Luong, D. R. So, J. Hall, N. Fiedel, R. Thoppilan, Z. Yang, A. Kulshreshtha, G. Nemade, Y. Lu, et al · 2020
Closest in time.