Fetching the paper…
Reading the bibliography…
As AI systems appear to exhibit ever-increasing capability and generality, assessing their true potential and safety becomes paramount.
An Introduction to Comparative Psychology
C.L. Morgan · 1903
Earlier work this paper cites.
Risks from Learned Optimization in Advanced Machine Learning Systems, December 2021
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Science and the modern world
Alfred North Whitehead · 1925
Earlier work this paper cites.
The abilities of man: Their nature and measurement
C. Spearman · 1927
Earlier work this paper cites.
On a distinction between hypothetical constructs and intervening variables
Kenneth MacCorquodale and Paul E. Meehl · 1939
Earlier work this paper cites.
The total library
Jorge Luis Borges · 1939
Earlier work this paper cites.
Spearman’s g found in 31 non-Western nations: Strong evidence that g is a universal phenomenon
Russell T. Warne and Cassidy Burningham · 1939
Earlier work this paper cites.
An Experimental Study of Apparent Behavior
Fritz Heider and Marianne Simmel · 1944
Earlier work this paper cites.
Construct validity in psychological tests
L. J. Cronbach and P. E. Meehl · 1955
Earlier work this paper cites.
Experimental and quasi-experimental designs for research
Donald Thomas Campbell, Julian C. Stanley, and Nathaniel Lees Gage · 1963
Earlier work this paper cites.
Theory of fluid and crystallized intelligence: A critical experiment
R. B. Cattell · 1963
Earlier work this paper cites.
The science of animal behaviour
P. L. (Peter Lovell) Broadhurst · 1963
Earlier work this paper cites.
Instructional technology and the measurement of learing outcomes: Some questions
Robert Glaser · 1963
Earlier work this paper cites.
Animal behavior
Keller Breland and Marian Breland · 1966
Earlier work this paper cites.
ELIZA – A Computer Program For the Study of Natural Language Communication Between Man And Machine
J. Weizenbaum · 1966
Earlier work this paper cites.
Information Retrieval
C.J. Van Rijsbergen · 1979
Earlier work this paper cites.
Problems of Monetary Management: The UK Experience
C. A. E. Goodhart · 1984
Earlier work this paper cites.
The Structure and Space of Possible Minds
Aaron Sloman · 1984
Earlier work this paper cites.
The reconstruction of hominid behavioral evolution through strategic modeling
John Tooby and Irven DeVore · 1987
Earlier work this paper cites.
The intentional stance
Daniel Clement Dennett · 1989
Earlier work this paper cites.
What one intelligence test measures: A theoretical account of the processing in the Raven Progressive Matrices Test
P. A. Carpenter, M. A. Just, and P. Shell · 1990
Earlier work this paper cites.
MUC-4 Evaluation Metrics
Nancy Chinchor · 1992
Earlier work this paper cites.
VALIDITY OF PSYCHOLOGICAL ASSESSMENT: VALIDATION OF INFERENCES FROM PERSONS’ RESPONSES AND PERFORMANCES AS SCIENTIFIC INQUIRY INTO SCORE MEANING
Samuel Messick · 1994
Earlier work this paper cites.
Testing heuristics: We have it all wrong
J. N. Hooker · 1995
Earlier work this paper cites.
The demon-haunted world: science as a candle in the dark
Carl Sagan · 1997
Earlier work this paper cites.
PerMIS 2000 White Paper: Measuring Performance and Intelligence of Systems with Autonomy
A. Meystel · 2000
Earlier work this paper cites.
Item response theory for psychologists
Susan E. Embretson and Steven P. Reise · 2000
Earlier work this paper cites.
The Predictive Value of IQ
Robert J. Sternberg, Elena L. Grigorenko, and Donald A. Bundy · 2001
Earlier work this paper cites.
Deep Blue
M. Campbell, A. J. Hoane, and F. Hsu · 2002
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
The Newell Test for a theory of cognition
J.R Anderson and C. Lebiere · 2003
Earlier work this paper cites.
Raven Progressive Matrices
John and Jean Raven · 2003
Earlier work this paper cites.
Agent57: Outperforming the Atari Human Benchmark
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, and Charles Blundell · 2003
Earlier work this paper cites.
Evaluating Models’ Local Decision Boundaries via Contrast Sets, 2020
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hanna Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou · 2004
Earlier work this paper cites.
At the mountains of madness
H. P. Lovecraft · 2005
Earlier work this paper cites.
AI Research Considerations for Human Existential Safety (ARCHES), May 2020
Andrew Critch and David Krueger · 2006
Earlier work this paper cites.
Higher-order truths about chmess
Daniel C. Dennett · 2006
Earlier work this paper cites.
The BICA cognitive decathlon: A test suite for biologically-inspired cognitive agents
S.T. Mueller, M. Jones, B.S. Minnery, and Julia M.H. Hiland · 2007
Earlier work this paper cites.
Challenges in Nonlinear Structural Equation Modeling
Polina Dimitruk, Karin Schermelleh-Engel, Augustin Kelava, and Helfried Moosbrugger · 2007
Earlier work this paper cites.
The Hawthorne Effect: a randomised, controlled trial
Rob McCarney, James Warner, Steve Iliffe, Robbert Van Haselen, Mark Griffin, and Peter Fisher · 2007
Earlier work this paper cites.
Adapting the Turing Test for embodied neurocognitive evaluation of biologically-inspired cognitive agents
S. T. Mueller and B. S. Minnery · 2008
Earlier work this paper cites.
Refining the Cognitive Decathlon
Robert L. Simpson and Charles R. Twardy · 2008
Earlier work this paper cites.
Effective Policing: Understanding How Polygraph Tests Work and Are Used
William G. Iacono · 2008
Earlier work this paper cites.
‘Improving ratings’: audit in the British University system
Marilyn Strathern · 2009
Earlier work this paper cites.
Multidimensional item response theory models
Mark D Reckase · 2009
Earlier work this paper cites.
Demand Characteristics and the Concept of Quasi-Controls
Martin T. Orne · 2009
Earlier work this paper cites.
Dynamic programming
Richard Bellman and Stuart Dreyfus · 2010
Earlier work this paper cites.
Cattell–Horn–Carroll abilities and cognitive tests: What we’ve learned from 20 years of research
Timothy Keith and Matthew R Reynolds · 2010
Earlier work this paper cites.
Intuitive physical reasoning about occluded objects by inexperienced chicks
Cinzia Chiandetti and Giorgio Vallortigara · 2010
Earlier work this paper cites.
Learning Word Vectors for Sentiment Analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Systematic review of the Hawthorne effect: New concepts are needed to study research participation effects
Jim McCambridge, John Witton, and Diana R. Elbourne · 2013
Earlier work this paper cites.
Superintelligence: Paths, dangers, strategies
N. Bostrom · 2014
Earlier work this paper cites.
Teaching Machines to Read and Comprehend
Karl Moritz Hermann, Tomáš Kočiskỳ, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Detecting Animal Deception
Shane D. Courtland · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond
Ramesh Nallapati, Bowen Zhou, Cicero Dos Santos, Caglar Gulcehre, and Bing Xiang · 2016
Cited alongside, same era.
Racing to the precipice: a model of artificial intelligence development
Stuart Armstrong, Nick Bostrom, and Carl Shulman · 2016
Cited alongside, same era.
Computer models solving intelligence test problems: Progress and implications
José Hernández-Orallo, Fernando Martínez-Plumed, Ute Schmid, Michael Siebers, and David L Dowe · 2016
Cited alongside, same era.
Deception in Game Theory: A Survey and Multiobjective Model
Everybody Lies: Deception Levels in Various Domains of Life
Kristina Šekrst · 2022
Later among the works it cites.
Measuring Progress on Scalable Oversight for Large Language Models, November 2022
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Later among the works it cites.
expert survey on progress in AI
Zach Stein-Perlman, Benjamin Weinstein-Raun, and Katja Grace · 2022
Later among the works it cites.
An Overview of Catastrophic AI Risks, September 2023
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Austin Davis · 2016
Cited alongside, same era.
The evolution of general intelligence
Judith M Burkart, Michèle N Schubiger, and Carel P van Schaik · 2017
Cited alongside, same era.
Deep Reinforcement Learning from Human Preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Cited alongside, same era.
Deal or No Deal? End-to-End Learning of Negotiation Dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Specification gaming examples in AI, 2018
Victoria Krakovna · 2018
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
John Burden, Konstantinos Voudouris, Ryan Burnell, Danaja Rutar, Lucy Cheke, and José Hernández-Orallo · 2023
Later among the works it cites.
Evaluating General-Purpose AI with Psychometrics, December 2023
Xiting Wang, Liming Jiang, Jose Hernandez-Orallo, David Stillwell, Luning Sun, Fang Luo, and Xing Xie · 2023
Later among the works it cites.
Sparks of Artificial General Intelligence: Early experiments with GPT-4, April 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Later among the works it cites.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Rahane, Anantharaman S. Iyer, Anders Johan Andreassen, Andrea Madotto, Andrea Santilli, Andreas Stuhlmüller, Andrew M. Dai, Andrew La, Andrew Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karakaş, B. Ryan Roberts, Bao Sheng Loe, Barret Zoph, Bart\lomiej Bojanowski, Batuhan Özyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Howald, Bryan Orinion, Cameron Diao, Cameron Dour, Catherine Stinson, Cedrick Argueta, Cesar Ferri, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Christopher Waites, Christian Voigt, Christopher D. Manning, Christopher Potts, Cindy Ramirez, Clara E. Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Dan Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, C. Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel Moseguí González, Danielle Perszyk, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, David Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Dylan Schrader, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth Donoway, Ellie Pavlick, Emanuele Rodolà, Emma Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A. Chi, Ethan Dyer, Ethan Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fanyue Xia, Fatemeh Siar, Fernando Martínez-Plumed, Francesca Happé, Francois Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, Germán Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Xinyue Wang, Gonzalo Jaimovitch-Lopez, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Hannah Kim, Hannah Rashkin, Hannaneh Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Francis Anthony Shevlin, Hinrich Schuetze, Hiromu Yakura, Hongming Zhang, Hugh Mee Wong, Ian Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, Jackson Kernion, Jacob Hilton, Jaehoon Lee, Jaime Fernández Fisac, James B. Simon, James Koppel, James Zheng, James Zou, Jan Kocon, Jana Thompson, Janelle Wingfield, Jared Kaplan, Jarema Radom, Jascha Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jennifer Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Joan Waweru, John Burden, John Miller, John U. Balis, Jonathan Batchelder, Jonathan Berant, Jörg Frohberg, Jos Rozen, Jose Hernandez-Orallo, Joseph Boudeman, Joseph Guerr, Joseph Jones, Joshua B. Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakrishnan, Katerina Ignatyeva, Katja Markert, Kaustubh Dhole, Kevin Gimpel, Kevin Omondi, Kory Wallace Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras-Ochando, Louis-Philippe Morency, Luca Moschella, Lucas Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros-Colón, Luke Metz, Lütfi Kerem Senel, Maarten Bosma, Maarten Sap, Maartje Ter Hoeve, Maheen Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, Maria Jose Ramirez-Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew L. Leavitt, Matthias Hagen, Mátyás Schubert, Medina Orduna Baitemirova, Melody Arnaud, Melvin McElrath, Michael Andrew Yee, Michael Cohen, Michael Gu, Michael Ivanitskiy, Michael Starritt, Michael Strube, Micha\l Swędrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Mitch Walker, Mo Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, Mukund Varma T, Nanyun Peng, Nathan Andrew Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas Roberts, Nick Doiron, Nicole Martinez, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha S. Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Pu Liang, Paul Vicol, Pegah Alipoormolabashi, Peiyuan Liao, Percy Liang, Peter W. Chang, Peter Eckersley, Phu Mon Htut, Pinyu Hwang, Piotr Mi\lkowski, Piyush Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, Qing Lyu, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ramon Risco, Raphaël Millière, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan Le Bras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Russ Salakhutdinov, Ryan Andrew Chi, Seungjae Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib Singh, Saif M. Mohammad, Sajant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Samuel R. Bowman, Samuel Stern Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Toshniwal, Shyam Upadhyay, Shyamolima Shammie Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo-Hwan Lee, Spencer Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Biderman, Stephanie Lin, Stephen Prasad, Steven Piantadosi, Stuart Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq Ali, Tatsunori Hashimoto, Te-Lin Wu, Théo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, Timofei Kornev, Titus Tunduny, Tobias Gerstenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Venkatesh Ramasesh, vinay uday prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, Wout Vossen, Xiang Ren, Xiaoyu Tong, Xinran Zhao, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yangqiu Song, Yasaman Bahri, Yejin Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yufang Hou, Yuntao Bai, Zachary Seid, Zhuoye Zhao, Zijian Wang, Zijie J. Wang, Zirui Wang, and Ziyi Wu · 2023
Later among the works it cites.
Rethink reporting of evaluation results in AI
Ryan Burnell, Wout Schellaert, John Burden, Tomer D. Ullman, Fernando Martinez-Plumed, Joshua B. Tenenbaum, Danaja Rutar, Lucy G. Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, Douwe Kiela, Murray Shanahan, Ellen M. Voorhees, Anthony G. Cohn, Joel Z. Leibo, and Jose Hernandez-Orallo · 2023
Later among the works it cites.
Apollo Research, 2023
Apollo Research · 2023
Later among the works it cites.
Frontier Threats Red Teaming for AI Safety, 2023
Anthropic · 2023
Later among the works it cites.
Google’s AI Red Team: the ethical hackers making AI safer, July 2023
Google · 2023
Later among the works it cites.
Microsoft AI Red Team building future of safer AI, August 2023
Ram Shankar Siva Kumar · 2023
Later among the works it cites.
NVIDIA AI Red Team: An Introduction, June 2023
Will Pearce and Joseph Lucas · 2023
Later among the works it cites.
Prompt Injection attack against LLM-integrated Applications, June 2023
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models, July 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Evaluating Language-Model Agents on Realistic Autonomous Tasks, July 2023
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, Elizabeth Barnes, and Paul Christiano · 2023
Later among the works it cites.
Model evaluation for extreme risks, 2023
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano, and Allan Dafoe · 2023
Later among the works it cites.
Anthropic’s Responsible Scaling Policy
Anthropic · 2023
Later among the works it cites.
Frontier AI Regulation: Managing Emerging Risks to Public Safety, September 2023
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf · 2023
Later among the works it cites.
Towards best practices in AGI safety and governance: A survey of expert opinion, May 2023
Jonas Schuett, Noemi Dreksler, Markus Anderljung, David McCaffary, Lennart Heim, Emma Bluemke, and Ben Garfinkel · 2023
Later among the works it cites.
Language models show human-like content effects on reasoning tasks, October 2023
Ishita Dasgupta, Andrew K. Lampinen, Stephanie C. Y. Chan, Hannah R. Sheahan, Antonia Creswell, Dharshan Kumaran, James L. McClelland, and Felix Hill · 2023
Later among the works it cites.
Using cognitive psychology to understand GPT-3
Marcel Binz and Eric Schulz · 2023
Later among the works it cites.
Harms from Increasingly Agentic Algorithmic Systems
Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco di Langosco, Zhonghao He, Yawen Duan, Micah Carroll, Michelle Lin, Alex Mayhew, Katherine Collins, Maryam Molamohammadi, John Burden, Wanru Zhao, Shalaleh Rismani, Konstantinos Voudouris, Umang Bhatt, Adrian Weller, David Krueger, and Tegan Maharaj · 2023
Later among the works it cites.
Why an Octopus-like Creature Has Come to Symbolize the State of A.I
Kevin Roose · 2023
Later among the works it cites.
ChatGPT: US lawyer admits using AI for case research
Kathryn Armstrong · 2023
Later among the works it cites.
GPT-4 Beats 90% Of Lawyers Trying To Pass The Bar, 2023
John Koetsier · 2023
Later among the works it cites.
OpenAI announces GPT-4, claims it can beat 90% of humans on the SAT, March 2023
Kif Leswing · 2023
Later among the works it cites.
Talking About Large Language Models, February 2023
Murray Shanahan · 2023
Later among the works it cites.
Defining Deception in Structural Causal Games
Francis Rhys Ward, Francesca Toni, and Francesco Belardinelli · 2023
Later among the works it cites.
Lorenzo Pacchiardi, Alex J. Chan, Sören Mindermann, Ilan Moscovitz, Alexa Y. Pan, Yarin Gal, Owain Evans, and Jan Brauner · 2023
Later among the works it cites.
Samuel Marks and Max Tegmark · 2023
Later among the works it cites.
Evaluating Superhuman Models with Consistency Checks, June 2023
Lukas Fluri, Daniel Paleka, and Florian Tramèr · 2023
Later among the works it cites.
ChemCrow: Augmenting large-language models with chemistry tools, 2023
Andres M. Bran, Sam Cox, Andrew D. White, and Philippe Schwaller · 2023
Later among the works it cites.
Building a Culture of Safety for AI: Perspectives and Challenges, June 2023
David Manheim · 2023
Later among the works it cites.
Sociotechnical Safety Evaluation of Generative AI Systems, October 2023
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability, January 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Towards Automated Circuit Discovery for Mechanistic Interpretability, July 2023
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Later among the works it cites.
Language models can explain neurons in language models, 2023
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Later among the works it cites.
Toward Sociotechnical AI: Mapping Vulnerabilities for Machine Learning in Context
Roel Dobbe and Anouk Wolters · 2024
Closest in time.
What Is It for a Machine Learning Model to Have a Capability?
Jacqueline Harding and Nathaniel Sharadin · 2024
Closest in time.
METR, 2024
METR · 2024
Closest in time.
Evaluating Frontier Models for Dangerous Capabilities, April 2024
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe, and Toby Shevlane · 2024
Closest in time.
Managing extreme AI risks amid rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner, and Sören Mindermann · 2024
Closest in time.
We need a Science of Evals, 2024
Apollo Research · 2024
Closest in time.
CogBench: a large language model walks into a psychology lab, February 2024
Julian Coda-Forno, Marcel Binz, Jane X. Wang, and Eric Schulz · 2024
Closest in time.
Visual cognition in multimodal large language models, January 2024
Luca M. Schulze Buschoff, Elif Akata, Matthias Bethge, and Eric Schulz · 2024
Closest in time.
From “AI”
Nanna Inie, Stefania Druga, Peter Zukerman, and Emily M. Bender · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training, 2024
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
Simple probes can catch sleeper agents, April 2024
Monte MacDiarmid, Timothy Maxwell, Nicholas Schiefer, Jesse Mu, Jared Kaplan, David Duvenaud, Sam Bowman, Alex Tamkin, Ethan Perez, Mrinank Sharma, Carson Denison, and Evan Hubinger · 2024
Closest in time.
Mechanistic Interpretability for AI Safety – A Review, April 2024
Leonard Bereska and Efstratios Gavves · 2024
Closest in time.
Thousands of AI Authors on the Future of AI, April 2024
Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler, Stephen Thomas, Ben Weinstein-Raun, and Jan Brauner · 2024
Closest in time.
General intelligence disentangled via a generality metric for natural and artificial intelligence
José Hernández-Orallo, Bao Sheng Loe, Lucy Cheke, Fernando Martínez-Plumed, and Seán Ó hÉigeartaigh · 2045
Closest in time.