Fetching the paper…
Reading the bibliography…
While certain industrial sectors (e.g., aviation) have a long history of mandatory incident reporting complete with analytical findings, the practice of artificial intelligence (AI) safety benefits from no such mandate and thus analyses must be performed on publicly known ``open source'' AI incidents.
Artificial intelligence: An overview,
V. Honavar, · 2006
Earlier work this paper cites.
Artificial intelligence safety and cybersecurity: A timeline of AI failures,
R. V. Yampolskiy, M. Spellchecker, · 2016
Earlier work this paper cites.
The landscape of AI safety and beneficence research. input for brainstorming at beneficial AI 2017,
R. Mallah, · 2017
Earlier work this paper cites.
Predicting future AI failures from historic examples,
R. V. Yampolskiy, · 2018
Earlier work this paper cites.
Categorizing variants of goodhart’s law,
D. Manheim, S. Garrabrant, · 2018
Earlier work this paper cites.
Cybersecurity research datasets: taxonomy and empirical analysis,
M. Zheng, H. Robbins, Z. Chai, P. Thapa, T. Moore, · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning,
K. Cobbe, O. Klimov, C. Hesse, T. Kim, J. Schulman, · 2019
Earlier work this paper cites.
Reframing superintelligence,
K. E. Drexler, · 2019
Cited alongside, same era.
AI failures: A review of underlying issues,
D. N. Banerjee, S. S. Chanda, · 2020
Cited alongside, same era.
AI watch. defining artificial intelligence. Towards an operational definition and taxonomy of artificial intelligence (2020)
S. Samoili, M. L. Cobo, E. Gomez, G. De Prato, F. Martinez-Plumed, B. Delipetrev, · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits,
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, C. Olah, · 2021
Cited alongside, same era.
Agi control theory,
R. V. Yampolskiy, · 2021
Cited alongside, same era.
E. Yudkowski, Agi ruin: A list of lethalities, https://intelligence.org/2022/06/10/agi-ruin/ , 2022. Accessed October 28th, 2022
2022
Closest in time.
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al., · 2022
Closest in time.
J. Porway, A taxonomy for AI / data for good, https://data.org/news/a-taxonomy-for-ai-data-for-good/ , 2022. Accessed October 28th, 2022
2022
Closest in time.
S. McGregor, The first taxonomy of AI incidents, https://incidentdatabase.ai/blog/the-first-taxonomy-of-ai-incidents , 2021. Accessed October 28th, 2022
2022
Closest in time.
Taxonomy of risks posed by language models,
L. Weidinger, J. Uesato, M. Rauh, C. Griffin, P.-S. Huang, J. Mellor, A. Glaese, M. Cheng, B. Balle, A. Kasirzadeh, et al., · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preventing repeated real world AI failures by cataloging incidents: The AI incident database,
S. McGregor, · 2021
Cited alongside, same era.
Closest in time.
J. Wentworth, You are not measuring what you think you are measuring, https://www.lesswrong.com/posts/9kNxhKWvixtKW5anS/you-are-not-measuring-what-you-think-you-are-measuring , 2022. Accessed October 28th, 2022
2022
Closest in time.