Fetching the paper…
Reading the bibliography…
Existing work on the alignment problem has focused mainly on (1) qualitative descriptions of the alignment problem; (2) attempting to align AI actions with human interests by focusing on value specification and learning; and/or (3) focusing on a single agent or on humanity as a monolith.
AI Research Considerations for Human Existential Safety (ARCHES)
Critch, A.; and Krueger, D. 2020 · 2006
Earlier work this paper cites.
CARLA: An open urban driving simulator
Dosovitskiy, A.; Ros, G.; Codevilla, F.; Lopez, A.; and Koltun, V. 2017 · 2017
Earlier work this paper cites.
Modeling controversy within populations
Jang, M.; Dori-Hacohen, S.; and Allan, J. 2017 · 2017
Earlier work this paper cites.
Classifying global catastrophic risks
Avin, S.; Wintle, B. C.; Weitzdörfer, J.; Ó hÉigeartaigh, S. S.; Sutherland, W. J.; and Rees, M. J. 2018 · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J.; Krueger, D.; Everitt, T.; Martic, M.; Maini, V.; and Legg, S. 2018 · 2018
Earlier work this paper cites.
Long-term trajectories of human civilization
Baum, S. D.; Armstrong, S.; Ekenstedt, T.; Häggström, O.; Hanson, R.; Kuhlemann, K.; Maas, M. M.; Miller, J. D.; Salmela, M.; Sandberg, A.; Sotala, K.; Torres, P.; Turchin, A.; and Yampolskiy, R. V. 2019 · 2019
Earlier work this paper cites.
The Impact of Artificial Intelligence on Strategic Stability and Nuclear Risk, Volume I, Euro-Atlantic Perspectives
Boulanin, V.; Avin, S.; Sauer, F.; Borrie, J.; Scheftelowitsch, D.; Bronk, J.; Stoutland, P. O.; Hagström, M.; Topychkanov, P.; Horowitz, M. C.; Kaspersen, A.; King, C.; Amadae, S.; and Rickli, J.-M. 2019 · 2019
Earlier work this paper cites.
Artificial intelligence & future warfare: implications for international security
Johnson, J. 2019 · 2019
Earlier work this paper cites.
The existential threat from cyber-enabled information warfare
Lin, H. 2019 · 2019
Earlier work this paper cites.
Human Compatible: Artificial Intelligence and the Problem of Control
Russell, S. 2019 · 2019
Earlier work this paper cites.
The Alignment Problem: Machine Learning and Human Values
Christian, B. 2020 · 2020
Cited alongside, same era.
Artificial Intelligence, Values, and Alignment
Gabriel, I. 2020 · 2020
Cited alongside, same era.
Military Artificial Intelligence as Contributor to Global Catastrophic Risk
Maas, M. M.; Matteuci, K.; and Cooke, D. 2022 · 2020
Cited alongside, same era.
The Precipice: Existential Risk and the Future of Humanity
Ord, T. 2020 · 2020
Cited alongside, same era.
Tackling threats to informed decision-making in democratic societies: Promoting epistemic security in a technologicall-advanced world
Seger, E.; Avin, S.; Pearson, G.; Briers, M.; Ó hÉigeartaigh, S.; and Bacon, H. 2020 · 2020
Cited alongside, same era.
Stewardship of global collective behavior
Bak-Coleman, J. B.; Alfano, M.; Barfuss, W.; Bergstrom, C. T.; Centeno, M. A.; Couzin, I. D.; Donges, J. F.; Galesic, M.; Gersick, A. S.; Jacquet, J.; Kao, A. B.; Moran, R. E.; Romanczuk, P.; Rubenstein, D. I.; Tombak, K. J.; Van Bavel, J. J.; and Weber, E. U. 2021 · 2021
Current and Near-Term AI as a Potential Existential Risk Factor
Bucknall, B. S.; and Dori-Hacohen, S. 2022 · 2022
Later among the works it cites.
The Ghost in the Machine has an American accent: value conflict in GPT-3
Johnson, R. L.; Pistilli, G.; Menédez-González, N.; Duran, L. D. D.; Panai, E.; Kalpokiene, J.; and Bertulfo, D. J. 2022 · 2022
Later among the works it cites.
Harms from increasingly agentic algorithmic systems
Chan, A.; Salganik, R.; Markelius, A.; Pang, C.; Rajkumar, N.; Krasheninnikov, D.; Langosco, L.; He, Z.; Duan, Y.; Carroll, M.; et al. 2023 · 2023
Later among the works it cites.
Defining and Mitigating Collusion in Multi-Agent Systems
Foxabbott, J.; Deverett, S.; Senft, K.; Dower, S.; and Hammond, L. 2023 · 2023
Later among the works it cites.
AI safety on whose terms?
Lazar, S.; and Nelson, A. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Value Alignment Verification
Brown, D. S.; Schneider, J.; Dragan, A.; and Niekum, S. 2021 · 2021
Cited alongside, same era.
Restoring Healthy Online Discourse by Detecting and Reducing Controversy, Misinformation, and Toxicity Online , 2627–2628
Dori-Hacohen, S.; Sung, K.; Chou, J.; and Lustig-Gonzalez, J. 2021 · 2021
Cited alongside, same era.
Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot
Leibo, J. Z.; Dueñez-Guzman, E. A.; Vezhnevets, A.; Agapiou, J. P.; Sunehag, P.; Koster, R.; Matyas, J.; Beattie, C.; Mordatch, I.; and Graepel, T. 2021 · 2021
Cited alongside, same era.
Value alignment: a formal approach
Sierra, C.; Osman, N.; Noriega, P.; Sabater-Mir, J.; and Perelló, A. 2021 · 2021
Cited alongside, same era.
Vezhnevets, A. S.; Agapiou, J. P.; Aharon, A.; Ziv, R.; Matyas, J.; Duéñez-Guzmán, E. A.; Cunningham, W. A.; Osindero, S.; Karmon, D.; and Leibo, J. Z. 2023 · 2023
Later among the works it cites.
The Ethics of Advanced AI Assistants
Gabriel, I.; Manzini, A.; Keeling, G.; Hendricks, L. A.; Rieser, V.; Iqbal, H.; Tomašev, N.; Ktena, I.; Kenton, Z.; Rodriguez, M.; El-Sayed, S.; Brown, S.; Akbulut, C.; Trask, A.; Hughes, E.; Bergman, A. S.; Shelby, R.; Marchal, N.; Griffin, C.; Mateos-Garcia, J.; Weidinger, L.; Street, W.; Lange, B.; Ingerman, A.; Lentz, A.; Enger, R.; Barakat, A.; Krakovna, V.; Siy, J. O.; Kurth-Nelson, Z.; McCroskery, A.; Bolina, V.; Law, H.; Shanahan, M.; Alberts, L.; Balle, B.; de Haas, S.; Ibitoye, Y.; Dafoe, A.; Goldberg, B.; Krier, S.; Reese, A.; Witherspoon, S.; Hawkins, W.; Rauh, M.; Wallace, D.; Franklin, M.; Goldstein, J. A.; Lehman, J.; Klenk, M.; Vallor, S.; Biles, C.; Morris, M. R.; King, H.; y Arcas, B. A.; Isaac, W.; and Manyika, J. 2024 · 2024
Closest in time.
AI Alignment: A Comprehensive Survey
Ji, J.; Qiu, T.; Chen, B.; Zhang, B.; Lou, H.; Wang, K.; Duan, Y.; He, Z.; Zhou, J.; Zhang, Z.; Zeng, F.; Ng, K. Y.; Dai, J.; Pan, X.; O’Gara, A.; Lei, Y.; Xu, H.; Tse, B.; Fu, J.; McAleer, S.; Yang, Y.; Wang, Y.; Zhu, S.-C.; Guo, Y.; and Gao, W. 2024 · 2024
Closest in time.
Kapoor, S.; Stroebl, B.; Siegel, Z. S.; Nadgir, N.; and Narayanan, A. 2024 · 2024
Closest in time.
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
Piatti, G.; Jin, Z.; Kleiman-Weiner, M.; Schölkopf, B.; Sachan, M.; and Mihalcea, R. 2024 · 2024
Closest in time.