Fetching the paper…
Reading the bibliography…
As artificial intelligence (AI) models are scaled up, new capabilities can emerge unintentionally and unpredictably, some of which might be dangerous.
‘Improving ratings’: Audit in the British university system
M. Strathern · 1997
Earlier work this paper cites.
Existential risks: Analyzing human extinction scenarios and related hazards
N. Bostrom · 2002
Earlier work this paper cites.
Intelligence explosion microeconomics
E. Yudkowsky · 2013
Earlier work this paper cites.
Racing to the precipice: A model of artificial intelligence development
S. Armstrong, N. Bostrom, and C. Shulman · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Guidelines for artificial intelligence containment
J. Babcock, J. Krámar, and R. V. Yampolskiy · 2019
Earlier work this paper cites.
Release strategies and the social impacts of language models
I. Solaiman, M. Brundage, J. Clark, A. Askell, A. Herbert-Voss, J. Wu, A. Radford, G. Krueger, J. W. Kim, S. Kreps, M. McCain, A. Newhouse, J. Blazakis, K. McGuffie, and J. Wang · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
The scaling hypothesis
Gwern · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
The role of section 708 of the Defense Production Act in the Federal Government’s response to COVID-19: Antitrust considerations
J. Lobert · 2020
Earlier work this paper cites.
The race for an artificial general intelligence: implications for public policy
W. Naudé and N. Dimitri · 2020
Earlier work this paper cites.
The precipice: Existential risk and the future of humanity
T. Ord · 2020
Earlier work this paper cites.
Executive Order 13911
The White House · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Earlier work this paper cites.
Corporate governance of artificial intelligence in the public interest
P. Cihon, J. Schuett, and S. D. Baum · 2021
Earlier work this paper cites.
AI & antitrust: Reconciling tensions between competition law and cooperative AI development
S.-S. Hua and H. Belfield · 2021
Earlier work this paper cites.
Preventing repeated real world AI failures by cataloging incidents: The AI Incident Database
S. McGregor · 2021
Earlier work this paper cites.
Process for adapting language models to society (PALMS) with values-targeted datasets
I. Solaiman and C. Dennison · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan · 2022
Earlier work this paper cites.
E. Caballero, K. Gupta, I. Rish, and D. Krueger · 2022
Earlier work this paper cites.
Is power-seeking AI an existential risk?
J. Carlsmith · 2022
Earlier work this paper cites.
Best practices for deploying language models
Cohere, OpenAI, and AI21 · 2022
Earlier work this paper cites.
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover
A. Cotra · 2022
Earlier work this paper cites.
Predictability and surprise in large generative models
D. Ganguli, D. Hernandez, L. Lovitt, A. Askell, Y. Bai, A. Chen, T. Conerly, N. Dassarma, D. Drain, N. Elhage, S. El Showk, S. Fort, Z. Hatfield-Dodds, T. Henighan, S. Johnston, A. Jones, N. Joseph, J. Kernian, S. Kravec, B. Mann, N. Nanda, K. Ndousse, C. Olsson, D. Amodei, T. Brown, J. Kaplan, S. McCandlish, C. Olah, D. Amodei, and J. Clark · 2022
Earlier work this paper cites.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
S. Gehrmann, E. Clark, and T. Sellam · 2022
Earlier work this paper cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Earlier work this paper cites.
Will AI make cyber swords or shields?
A. John and K. Jackson · 2022
Earlier work this paper cites.
Information security considerations for AI and the long term future
J. Ladish and L. Heim · 2022
Earlier work this paper cites.
Illustrating reinforcement learning from human feedback (RLHF)
N. Lampert, L. Castricato, L. V. Werra, and A. Havrilla · 2022
Earlier work this paper cites.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Ré, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. Wang, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda · 2022
Earlier work this paper cites.
AI and compute: How much longer can computing power drive artificial intelligence progress?
A. Lohn and M. Musser · 2022
Earlier work this paper cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
Meta Fundamental AI Research Diplomacy Team (FAIR), A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu, A. P. Jacob, M. Komeili, K. Konath, M. Kwon, A. Lerer, M. Lewis, A. H. Miller, S. Mitts, A. Renduchintala, S. Roller, D. Rowe, W. Shi, J. Spisak, A. Wei, D. Wu, H. Zhang, and M. Zijlstra · 2022
Earlier work this paper cites.
200 concrete open problems in mechanistic interpretability: Introduction
N. Nanda · 2022
Earlier work this paper cites.
The alignment problem from a deep learning perspective
R. Ngo, L. Chan, and S. Mindermann · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
E. Perez, S. Ringer, K. Lukošiūtė, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, A. Jones, A. Chen, B. Mann, B. Israel, B. Seethor, C. McKinnon, C. Olah, D. Yan, D. Amodei, D. Amodei, D. Drain, D. Li, E. Tran-Johnson, G. Khundadze, J. Kernion, J. Landis, J. Kerr, J. Mueller, J. Hyun, J. Landau, K. Ndousse, L. Goldberg, L. Lovitt, M. Lucas, M. Sellitto, M. Zhang, N. Kingsland, N. Elhage, N. Joseph, N. Mercado, N. DasSarma, O. Rausch, R. Larson, S. McCandlish, S. Johnston, S. Kravec, S. E. Showk, T. Lanham, T. Telleen-Lawton, T. Brown, T. Henighan, T. Hume, Y. Bai, Z. Hatfield-Dodds, J. Clark, S. R. Bowman, A. Askell, R. Grosse, D. Hernandez, D. Ganguli, E. Hubinger, N. Schiefer, and J. Kaplan · 2022
Cited alongside, same era.
Differential technology development: A responsible innovation principle for navigating technology risks
J. Sandbrink, H. Hobbs, J. Swett, A. Dafoe, and A. Sandberg · 2022
Cited alongside, same era.
Structured access: An emerging paradigm for safe AI deployment
T. Shevlane · 2022
Cited alongside, same era.
Towards understanding-based safety evaluations
E. Hubinger · 2023
Closest in time.
Don’t pause giant AI for the wrong reasons
M. Ienca · 2023
Closest in time.
Evaluating language-model agents on realistic autonomous tasks
M. Kinniment, L. Jun, K. Sato, H. Du, B. Goodrich, M. Hasin, L. Chan, L. H. Miles, T. R. Lin, H. Wijk, J. Burget, A. Ho, E. Barnes, and P. Christiano · 2023
Closest in time.
L. Koessler and J. Schuett · 2023
Closest in time.
Power-seeking can be probable and predictive for trained agents
V. Krakovna and J. Kramar · 2023
Closest in time.
Some high-level thoughts on the DeepMind alignment team’s strategy
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A seismic shift: The new U.S. semiconductor export controls and the implications for U.S. firms, allies, and the innovation ecosystem
S. Shivakumar, C. Wessner, and T. Howell · 2022
Cited alongside, same era.
Beyond neural scaling laws: Beating power law scaling via data pruning
B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. S. Morcos · 2022
Cited alongside, same era.
Transcending scaling laws with 0.1% extra compute
Y. Tay, J. Wei, H. W. Chung, V. Q. Tran, D. R. So, S. Shakeri, X. Garcia, H. S. Zheng, J. Rao, A. Chowdhery, D. Zhou, D. Metzler, S. Petrov, N. Houlsby, Q. V. Le, and M. Dehghani · 2022
Cited alongside, same era.
Parametrically retargetable decision-makers tend to seek power
A. M. Turner and P. Tadepalli · 2022
Cited alongside, same era.
Dual use of artificial-intelligence-powered drug discovery
F. Urbina, F. Lentzos, C. Invernizzi, and S. Ekins · 2022
Cited alongside, same era.
Commerce implements new export controls on advanced computing and semiconductor manufacturing items to the People’s Republic of China (PRC)
US Bureau of Industry and Security · 2022
Cited alongside, same era.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus · 2022
Cited alongside, same era.
Whitepaper
AI Objectives Institute · 2023
Cited alongside, same era.
V. Krakovna and R. Shah · 2023
Closest in time.
12 tentative ideas for US AI policy
L. Muehlhauser · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Lieberum, J. Smith, and J. Steinhardt · 2023
Closest in time.
Deployment corrections: An incident response fromework for frontier AI models
J. O’Brien, Z. Williams, and S. Ee · 2023
Closest in time.
Interpretability dreams
C. Olah · 2023
Closest in time.
Frontier model forum
OpenAI · 2023
Closest in time.
OpenAI · 2023
Closest in time.
March 20 ChatGPT outage: Here’s what happened
OpenAI · 2023
Closest in time.
AI deception: A survey of examples, risks, and potential solutions
P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks · 2023
Closest in time.
Letter signed by Elon Musk demanding AI research pause sparks controversy
K. Paul · 2023
Closest in time.
J. Sandbrink · 2023
Closest in time.
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo · 2023
Closest in time.
AGI labs need an internal audit function
J. Schuett · 2023
Closest in time.
Towards best practices in AGI safety and governance: A survey of expert opinion
J. Schuett, N. Dreksler, M. Anderljung, D. McCaffary, L. Heim, E. Bluemke, and B. Garfinkel · 2023
Closest in time.
How to design an AI ethics board
J. Schuett, A. Reuel, and A. Carlier · 2023
Closest in time.
Open-sourcing highly capable foundation models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives
E. Seger, N. Dreksler, R. Moulange, E. Dardaman, J. Schuett, K. Wei, C. Winter, M. Arnold, S. ÓhÉigeartaigh, A. Korinek, M. Anderljung, B. Bucknall, A. Chan, E. Stafford, L. Koessler, A. Ovadya, B. Garfinkel, E. Bluemke, M. Aird, P. Levermore, J. Hazell, and A. Gupta · 2023
Closest in time.
Democratising AI: Multiple meanings, goals, and methods
E. Seger, A. Ovadya, B. Garfinkel, D. Siddarth, and A. Dafoe · 2023
Closest in time.
An early warning system for novel AI risks
T. Shevlane · 2023
Closest in time.
Model evaluation for extreme risks
T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt, L. Ho, D. Siddarth, S. Avin, W. Hawkins, B. Kim, I. Gabriel, V. Bolina, J. Clark, Y. Bengio, P. Christiano, and A. Dafoe · 2023
Closest in time.
The gradient of generative AI release: Methods and considerations
I. Solaiman · 2023
Closest in time.
Fact sheet: Biden-Harris administration announces new actions to promote responsible AI innovation that protects Americans’ rights and safety
The White House · 2023
Closest in time.
Fact sheet: Biden-Harris administration secures voluntary commitments from leading artificial intelligence companies to manage the risks posed by AI
The White House · 2023
Closest in time.
The illusion of China’s AI prowess
H. Toner, J. Xiao, and J. Ding · 2023
Closest in time.
Optimal policies tend to seek power
A. M. Turner, L. Smith, R. Shah, A. Critch, and P. Tadepalli · 2023
Closest in time.
AI regulation: A pro-innovation approach
UK Department for Science Innovation and Technology and UK Office for Artificial Intelligence · 2023
Closest in time.
Scaling laws literature review
P. Villalobos · 2023
Closest in time.
Pausing AI developments isn’t enough. We need to shut it all down
E. Yudkowsky · 2023
Closest in time.
FTC investigates OpenAI over data leak and ChatGPT’s inaccuracy
C. Zakrzewski · 2023
Closest in time.