Fetching the paper…
Reading the bibliography…
To understand the risks posed by a new AI system, we must understand what it can and cannot do.
Risks from learned optimization in advanced machine learning systems
E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant · 1906
Earlier work this paper cites.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
The Morality of Freedom
J. Raz · 1988
Earlier work this paper cites.
Information theory, inference and learning algorithms
D. J. MacKay · 2003
Earlier work this paper cites.
Expert political judgement: How good is it? How can we know? , volume 321
P. E. Tetlock · 2005
Earlier work this paper cites.
The radicalization risks of GPT-3 and advanced neural language models
K. McGuffie and A. Newhouse · 2009
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Computational Propaganda: Political Parties, Politicians, and Political Manipulation on Social Media
S. C. Woolley and P. N. Howard · 2018
Earlier work this paper cites.
Technology, autonomy, and manipulation
D. Susser, B. Roessler, and H. Nissenbaum · 2019
Earlier work this paper cites.
Detecting "0-day" vulnerability: An empirical study of secret security patch in OSS
X. Wang, K. Sun, A. Batcheller, and S. Jajodia · 2019
Earlier work this paper cites.
Spinning language models: Risks of Propaganda-As-A-Service and countermeasures
E. Bagdasaryan and V. Shmatikov · 2021
Earlier work this paper cites.
A comprehensive survey of AI-enabled phishing attacks detection techniques
A. Basit, M. Zafar, X. Liu, A. R. Javed, Z. Jalil, and K. Kifayat · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Z. Kenton, T. Everitt, L. Weidinger, I. Gabriel, V. Mikulik, and G. Irving · 2021
Earlier work this paper cites.
SPI: Automated identification of security patches via commits
Y. Zhou, J. K. Siow, C. Wang, S. Liu, and Y. Liu · 2021
Earlier work this paper cites.
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover
A. Cotra · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Earlier work this paper cites.
The alignment problem from a deep learning perspective
R. Ngo · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Earlier work this paper cites.
ReAct: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Earlier work this paper cites.
Anthropic’s responsible scaling policy, Sept. 2023
Anthropic · 2023
Earlier work this paper cites.
Artificial intelligence can persuade humans on political issues, Feb. 2023
H. Bai, J. G. Voelkel, J. C. Eichstaedt, and R. Willer · 2023
Cited alongside, same era.
Taken out of context: On measuring situational awareness in LLMs
L. Berglund, A. C. Stickland, M. Balesni, M. Kaufmann, M. Tong, T. Korbak, D. Kokotajlo, and O. Evans · 2023
Cited alongside, same era.
Purple Llama CyberSecEval: A secure coding benchmark for language models
M. Bhatt, S. Chennabasappa, C. Nikolaidis, S. Wan, I. Evtimov, D. Gabi, D. Song, F. Ahmad, C. Aschermann, L. Fontana, S. Frolov, R. P. Giri, D. Kapil, Y. Kozyrakis, D. LeBlanc, J. Milazzo, A. Straumann, G. Synnaeve, V. Vontimitta, S. Whitman, and J. Saxe · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Cited alongside, same era.
AI deception: A survey of examples, risks, and potential solutions
P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo · 2023
Later among the works it cites.
J. Scheurer, M. Balesni, and M. Hobbhahn · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, N. Cheng, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, S. Kravec, T. Maxwell, S. McCandlish, K. Ndousse, O. Rausch, N. Schiefer, D. Yan, M. Zhang, and E. Perez · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Burtell and T. Woodside · 2023
Cited alongside, same era.
Characterizing manipulation from AI systems
M. Carroll, A. Chan, H. Ashton, and D. Krueger · 2023
Cited alongside, same era.
Harms from increasingly agentic algorithmic systems
A. Chan, R. Salganik, A. Markelius, C. Pang, N. Rajkumar, D. Krasheninnikov, L. Langosco, Z. He, Y. Duan, M. Carroll, et al · 2023
Cited alongside, same era.
AI capabilities can be significantly improved without expensive retraining
T. Davidson, J.-S. Denain, P. Villalobos, and G. Bas · 2023
Cited alongside, same era.
Language modeling is compression
G. Delétang, A. Ruoss, P.-A. Duquenne, E. Catt, T. Genewein, C. Mattern, J. Grau-Moya, L. K. Wenliang, M. Aitchison, L. Orseau, et al · 2023
Cited alongside, same era.
C. Gao, H. Jiang, D. Cai, S. Shi, and W. Lam · 2023
Cited alongside, same era.
Can AI write persuasive propaganda?, Apr. 2023
J. A. Goldstein, J. Chao, S. Grossman, A. Stamos, and M. Tomz · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Google · 2023
Cited alongside, same era.
Practices for governing agentic AI systems, 2023
Y. Shavit, S. Agarwal, M. Brundage, S. Adler, C. O’Keefe, R. Campbell, T. Lee, P. Mishkin, T. Eloundou, A. Hickey, et al · 2023
Later among the works it cites.
Model evaluation for extreme risks
T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt, L. Ho, D. Siddarth, S. Avin, W. Hawkins, B. Kim, I. Gabriel, V. Bolina, J. Clark, Y. Bengio, P. Christiano, and A. Dafoe · 2023
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
N. Shinn, B. Labash, and A. Gopinath · 2023
Later among the works it cites.
Fact sheet: President Biden issues Executive Order on safe, secure, and trustworthy artificial intelligence, Oct. 2023a
The White House · 2023
Later among the works it cites.
Ensuring safe, secure, and trustworthy AI, July 2023b
The White House · 2023
Later among the works it cites.
Introducing the AI Safety Institute, Nov. 2023
UK AI Safety Institute · 2023
Later among the works it cites.
A survey on large language model based autonomous agents
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al · 2023
Later among the works it cites.
Honesty is the best policy: Defining and mitigating AI deception
F. R. Ward, F. Belardinelli, F. Toni, and T. Everitt · 2023
Later among the works it cites.
Jailbroken: How does LLM safety training fail?
A. Wei, N. Haghtalab, and J. Steinhardt · 2023
Later among the works it cites.
Sociotechnical safety evaluation of generative AI systems
L. Weidinger, M. Rauh, N. Marchal, A. Manzini, L. A. Hendricks, J. Mateos-Garcia, S. Bergman, J. Kay, C. Griffin, B. Bariach, I. Gabriel, V. Rieser, and W. Isaac · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, et al · 2023
Later among the works it cites.
InterCode: Standardizing and benchmarking interactive coding with execution feedback
J. Yang, A. Prabhakar, K. Narasimhan, and S. Yao · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Later among the works it cites.
Large language models are not robust multiple choice selectors
C. Zheng, H. Zhou, F. Meng, J. Zhou, and M. Huang · 2023
Later among the works it cites.
The Claude 3 model family: Opus, Sonnet, Haiku, Mar. 2024
Anthropic · 2024
Closest in time.
Building an early warning system for llm-aided biological threat creation, Jan. 2024
OpenAI · 2024
Closest in time.
GPT-4V(ision) is a generalist web agent, if grounded
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su · 2024
Closest in time.