Fetching the paper…
Reading the bibliography…
Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities.
The Off-Switch game
D. Hadfield-Menell, A. Dragan, P. Abbeel, and S. Russell · 2016
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
M. Brundage, S. Avin, J. Clark, H. Toner, P. Eckersley, B. Garfinkel, A. Dafoe, P. Scharre, T. Zeitzoff, B. Filar, H. Anderson, H. Roff, G. C. Allen, J. Steinhardt, C. Flynn, S. Ó. hÉigeartaigh, S. Beard, H. Belfield, S. Farquhar, C. Lyle, R. Crootof, O. Evans, M. Page, J. Bryson, R. Yampolskiy, and D. Amodei · 2018
Earlier work this paper cites.
Model cards for model reporting
M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru · 2018
Earlier work this paper cites.
Toward trustworthy AI development: Mechanisms for supporting verifiable claims
M. Brundage, S. Avin, J. Wang, H. Belfield, G. Krueger, G. Hadfield, H. Khlaaf, J. Yang, H. Toner, R. Fong, T. Maharaj, P. W. Koh, S. Hooker, J. Leung, A. Trask, E. Bluemke, J. Lebensold, C. O’Keefe, M. Koren, T. Ryffel, J. B. Rubinovitz, T. Besiroglu, F. Carugati, J. Clark, P. Eckersley, S. de Haas, M. Johnson, B. Laurie, A. Ingerman, I. Krawczuk, A. Askell, R. Cammarota, A. Lohn, D. Krueger, C. Stix, P. Henderson, L. Graham, C. Prunkl, B. Martin, E. Seger, N. Zilberman, S. Ó. hÉigeartaigh, F. Kroeger, G. Sastry, R. Kagan, A. Weller, B. Tse, E. Barnes, A. Dafoe, P. Scharre, A. Herbert-Voss, M. Rasser, S. Sodhani, C. Flynn, T. K. Gilbert, L. Dyer, S. Khan, Y. Bengio, and M. Anderljung · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Earlier work this paper cites.
Closing the AI accountability gap: Defining an End-to-End framework for internal algorithmic auditing
I. D. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes · 2020
Earlier work this paper cites.
Agent incentives: A causal perspective
T. Everitt, R. Carey, E. Langlois, P. A. Ortega, and S. Legg · 2021
Earlier work this paper cites.
Optimal policies tend to seek power
A. Turner, L. Smith, R. Shah, A. Critch, and P. Tadepalli · 2021
Earlier work this paper cites.
Why and how governments should monitor AI development
J. Whittlestone and J. Clark · 2021
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
C. Burns, H. Ye, D. Klein, and J. Steinhardt · 2022
Earlier work this paper cites.
Is Power-Seeking AI an existential risk?
J. Carlsmith · 2022
Earlier work this paper cites.
Predictability and surprise in large generative models
D. Ganguli, D. Hernandez, L. Lovitt, N. DasSarma, T. Henighan, A. Jones, N. Joseph, J. Kernion, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, N. Elhage, S. El Showk, S. Fort, Z. Hatfield-Dodds, S. Johnston, S. Kravec, N. Nanda, K. Ndousse, C. Olsson, D. Amodei, D. Amodei, T. Brown, J. Kaplan, S. McCandlish, C. Olah, and J. Clark · 2022
Earlier work this paper cites.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
A. Glaese, N. McAleese, M. Trębacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, L. Campbell-Gillingham, J. Uesato, P.-S. Huang, R. Comanescu, F. Yang, A. See, S. Dathathri, R. Greig, C. Chen, D. Fritz, J. S. Elias, R. Green, S. Mokrá, N. Fernando, B. Wu, R. Foley, S. Young, I. Gabriel, W. Isaac, J. Mellor, D. Hassabis, K. Kavukcuoglu, L. A. Hendricks, and G. Irving · 2022
Cited alongside, same era.
Discovering agents
Z. Kenton, R. Kumar, S. Farquhar, J. Richens, M. MacDermott, and T. Everitt · 2022
Cited alongside, same era.
A hazard analysis framework for code synthesis large language models
H. Khlaaf, P. Mishkin, J. Achiam, G. Krueger, and M. Brundage · 2022
Cited alongside, same era.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Ré, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. Wang, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang · 2023
Closest in time.
Harms from increasingly agentic algorithmic systems
A. Chan, R. Salganik, A. Markelius, C. Pang, N. Rajkumar, D. Krasheninnikov, L. Langosco, Z. He, Y. Duan, M. Carroll, M. Lin, A. Mayhew, K. Collins, M. Molamohammadi, J. Burden, W. Zhao, S. Rismani, K. Voudouris, U. Bhatt, A. Weller, D. Krueger, and T. Maharaj · 2023
Closest in time.
Uncertainty, information, and risk in international technology races
N. Emery-Xu, A. Park, and R. Trager · 2023
Closest in time.
Automatically auditing large language models via discrete optimization
E. Jones, A. Dragan, A. Raghunathan, and J. Steinhardt · 2023
Closest in time.
Power-seeking can be probable and predictive for trained agents
V. Krakovna and J. Kramar · 2023
Closest in time.
Inverse scaling prize
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Michael, A. Holtzman, A. Parrish, A. Mueller, A. Wang, A. Chen, D. Madaan, N. Nangia, R. Y. Pang, J. Phang, and S. R. Bowman · 2022
Cited alongside, same era.
The alignment problem from a deep learning perspective
R. Ngo, L. Chan, and S. Mindermann · 2022
Cited alongside, same era.
Lessons from the development of the atomic bomb
T. Ord · 2022
Cited alongside, same era.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
R. Shah, V. Varma, R. Kumar, M. Phuong, V. Krakovna, J. Uesato, and Z. Kenton · 2022
Cited alongside, same era.
Adversarial training for High-Stakes reliability
D. M. Ziegler, S. Nix, L. Chan, T. Bauman, P. Schmidt-Nielsen, T. Lin, A. Scherlis, N. Nabeshima, B. Weinstein-Raun, D. de Haas, B. Shlegeris, and N. Thomas · 2022
Cited alongside, same era.
Strengthening U.S. AI innovation through an ambitious investment in NIST
Anthropic · 2023
Cited alongside, same era.
Update on ARC’s recent eval efforts
ARC Evals · 2023
Cited alongside, same era.
Exploring the relevance of data Privacy-Enhancing technologies for AI governance use cases
E. Bluemke, T. Collins, B. Garfinkel, and A. Trask · 2023
Cited alongside, same era.
I. McKenzie, A. Lyzhov, A. Parrish, A. Prabhu, A. Mueller, N. Kim, S. Bowman, and E. Perez · 2023
Closest in time.
Auditing large language models: a three-layered approach
J. Mökander, J. Schuett, H. R. Kirk, and L. Floridi · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Lieberum, J. Smith, and J. Steinhardt · 2023
Closest in time.
Safely interruptible agents
L. Orseau and S. Armstrong · 2023
Closest in time.
Do the rewards justify the means? measuring Trade-Offs between rewards and ethical behavior in the MACHIAVELLI benchmark
A. Pan, C. J. Shern, A. Zou, N. Li, S. Basart, T. Woodside, J. Ng, H. Zhang, S. Emmons, and D. Hendrycks · 2023
Closest in time.
From plane crashes to algorithmic harm: Applicability of safety engineering frameworks for responsible ML
S. Rismani, R. Shelby, A. Smart, E. Jatho, J. Kroll, A. Moon, and N. Rostamzadeh · 2023
Closest in time.
Sharing powerful AI models
T. Shevlane · 2023
Closest in time.
Thinking about risks from AI: Accidents, misuse and structure
R. Zwetsloot and A. Dafoe · 2023
Closest in time.