Fetching the paper…
Reading the bibliography…
Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks.
Economic possibilities for our grandchildren - Essays in persuasion
J. Keynes · 1931
Earlier work this paper cites.
Ignition of the atmosphere with nuclear bombs
E. Konopinski, C. Marvin, and E. Teller · 1946
Earlier work this paper cites.
Games against nature
J. W. Milnor · 1954
Earlier work this paper cites.
Rational choice and the structure of the environment
H. A. Simon · 1956
Earlier work this paper cites.
Speculations on perceptrons and other automata
I. J. Good · 1959
Earlier work this paper cites.
The new science of management decision
H. A. Simon · 1960
Earlier work this paper cites.
The Structure of Scientific Revolutions
T. S. Kuhn · 1962
Earlier work this paper cites.
Cat’s Cradle
K. Vonnegut · 1963
Earlier work this paper cites.
Cramming more components onto integrated circuits
G. E. Moore · 1965
Earlier work this paper cites.
Some future social repercussions of computers
I. J. Good · 1970
Earlier work this paper cites.
More is different: Broken symmetry and the nature of the hierarchical structure of science
P. W. Anderson · 1972
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty
A. Tversky and D. Kahneman · 1974
Earlier work this paper cites.
Progress in digital integrated electronics
G. E. Moore et al · 1975
Earlier work this paper cites.
The “false consensus effect”: An egocentric bias in social perception and attribution processes
L. Ross, D. Greene, and P. House · 1977
Earlier work this paper cites.
Issues in assessing the contribution of research and development to productivity growth
Z. Griliches · 1979
Earlier work this paper cites.
The need for biases in learning generalizations
T. M. Mitchell · 1980
Earlier work this paper cites.
The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring
B. S. Bloom · 1984
Earlier work this paper cites.
Fundamentals of expert systems
B. G. Buchanan and R. G. Smith · 1988
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Aviation deregulation and safety: Theory and evidence
L. N. Moses and I. Savage · 1990
Earlier work this paper cites.
The contribution of latent human failures to the breakdown of complex systems
J. Reason · 1990
Earlier work this paper cites.
Response strategies for coping with the cognitive demands of attitude measures in surveys
J. A. Krosnick · 1991
Earlier work this paper cites.
Four types of learning curves
S.-i. Amari, N. Fujita, and S. Shinomoto · 1992
Earlier work this paper cites.
Information-based objective functions for active data selection
D. J. C. MacKay · 1992
Earlier work this paper cites.
Self-enhancement biases and negotiator judgment: Effects of self-esteem and mood
R. M. Kramer, E. Newton, and P. L. Pommerenke · 1993
Earlier work this paper cites.
Active learning with statistical models
D. Cohn, Z. Ghahramani, and M. Jordan · 1994
Earlier work this paper cites.
Trust, self-confidence, and operators’ adaptation to automation
J. D. Lee and N. Moray · 1994
Earlier work this paper cites.
R & D-based models of economic growth
C. I. Jones · 1995
Earlier work this paper cites.
The private and social returns to research and development
B. H. Hall · 1996
Earlier work this paper cites.
How long before superintelligence
N. Bostrom · 1998
Earlier work this paper cites.
When will computer hardware match the human brain?
H. Moravec · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
R. M. French · 1999
Earlier work this paper cites.
Code of the street: Decency, violence, and the moral life of the inner city
E. Anderson · 2000
Earlier work this paper cites.
Intrusion Detection Systems
R. Bace and P. Mell · 2000
Earlier work this paper cites.
Long-term growth as a sequence of exponential modes, 2000
R. Hanson · 2000
Earlier work this paper cites.
Ultimate physical limits to computation
S. Lloyd · 2000
Earlier work this paper cites.
Scaling to very very large corpora for natural language disambiguation
M. Banko and E. Brill · 2001
Earlier work this paper cites.
Economic growth given machine intelligence, 2001
R. Hanson · 2001
Earlier work this paper cites.
Deep Blue
M. Campbell, A. J. Hoane Jr, and F.-h. Hsu · 2002
Earlier work this paper cites.
By 2029 no computer - or "machine intelligence" - will have passed the Turing Test, 2002
R. Kurzweil · 2002
Earlier work this paper cites.
In two minds: dual-process accounts of reasoning
J. S. B. Evans · 2003
Earlier work this paper cites.
Columbia and challenger: organizational failure at nasa
J. L. Hall · 2003
Earlier work this paper cites.
Bank secrecy act, anti-money laundering, and office of foreign assets control
Federal Deposit Insurance Corporation · 2004
Earlier work this paper cites.
Universal limits on computation
L. M. Krauss and G. D. Starkman · 2004
Earlier work this paper cites.
Trust in automation: Designing for appropriate reliance
J. D. Lee and K. A. See · 2004
Earlier work this paper cites.
The singularity is near
R. Kurzweil · 2005
Earlier work this paper cites.
Primacy and recency effects on clicking behavior
J. Murphy, C. Hofacker, and R. Mizerski · 2006
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
S. Legg and M. Hutter · 2007
Earlier work this paper cites.
The basic AI drives
S. M. Omohundro · 2008
Earlier work this paper cites.
Recursive self-improvement, 2008
E. Yudkowsky · 2008
Earlier work this paper cites.
Anomaly detection: A survey
V. Chandola, A. Banerjee, and V. Kumar · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
How to explain individual classification decisions
D. Baehrens, T. Schroeter, S. Harmeling, M. Kawanabe, K. Hansen, and K.-R. Müller · 2010
Earlier work this paper cites.
The singularity: A philosophical analysis
D. Chalmers · 2010
Earlier work this paper cites.
Building Watson: An overview of the DeepQA project
D. Ferrucci, E. Brown, J. Chu-Carroll, J. Fan, D. Gondek, A. A. Kalyanpur, A. Lally, J. W. Murdock, E. Nyberg, J. Prager, et al · 2010
Earlier work this paper cites.
The economic effects of airline deregulation
S. Morrison and C. Winston · 2010
Earlier work this paper cites.
Ontological crises in artificial agents’ value systems
P. De Blanc · 2011
Earlier work this paper cites.
Thinking, fast and slow
D. Kahneman · 2011
Earlier work this paper cites.
The better angels of our nature: The decline of violence in history and its causes
S. Pinker · 2011
Earlier work this paper cites.
Losing humanity: The case against killer robots
B. L. Docherty · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Anniversary - 80 years ago, Leo Szilard envisioned neutron chain reaction, Sep 2013
R. Adams · 2013
Earlier work this paper cites.
Organizational issues in the implementation and adoption of health information technology innovations: an interpretative review
K. Cresswell and A. Sheikh · 2013
Earlier work this paper cites.
Algorithmic progress in six domains
K. Grace · 2013
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
D. Kahneman and A. Tversky · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Deterrence: A review of the evidence by a criminologist for economists
D. S. Nagin · 2013
Earlier work this paper cites.
An overview of models of technological singularity
A. Sandberg · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
K. Simonyan · 2013
Earlier work this paper cites.
Intelligence explosion microeconomics
E. Yudkowsky · 2013
Earlier work this paper cites.
The operational role of security information and event management systems
S. Bhatt, P. K. Manadhata, and L. Zomlot · 2014
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
N. Bostrom · 2014
Earlier work this paper cites.
Approval directed agents, 2014
P. Christiano · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller · 2014
Earlier work this paper cites.
How we’re predicting AI–or failing to
S. Armstrong and K. Sotala · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Earlier work this paper cites.
Corrigibility
N. Soares, B. Fallenstein, S. Armstrong, and E. Yudkowsky · 2015
Earlier work this paper cites.
F. Wang and C. Rudin · 2015
Earlier work this paper cites.
A survey of sparse representation: algorithms and applications
Z. Zhang, Y. Xu, J. Yang, X. Li, and D. Zhang · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain · 2016
Earlier work this paper cites.
DeepDGA: Adversarially-tuned domain generation and detection
H. S. Anderson, J. Woodbridge, and B. Filar · 2016
Earlier work this paper cites.
Racing to the precipice: a model of artificial intelligence development
S. Armstrong, N. Bostrom, and C. Shulman · 2016
Earlier work this paper cites.
Conditional computation in neural networks for faster models
E. Bengio, P.-L. Bacon, J. Pineau, and D. Precup · 2016
Earlier work this paper cites.
Learning the preferences of ignorant, inconsistent agents
O. Evans, A. Stuhlmüller, and N. Goodman · 2016
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Earlier work this paper cites.
The phylogenetic roots of human lethal violence
J. M. Gómez, M. Verdú, A. González-Megías, and M. Méndez · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, S. J. Russell, P. Abbeel, and A. Dragan · 2016
Earlier work this paper cites.
The age of Em: Work, love, and life when robots rule the earth
R. Hanson · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
MAWPS: A math word problem repository
R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther · 2016
Earlier work this paper cites.
Engineering a safer world: Systems thinking applied to safety
N. G. Leveson · 2016
Earlier work this paper cites.
Fast and flexible monotonic functions with ensembles of lattices
M. Milani Fard, K. Canini, A. Cotter, J. Pfeifer, and M. Gupta · 2016
Earlier work this paper cites.
Grad-cam: Why did you say that?
R. R. Selvaraju, A. Das, R. Vedantam, M. Cogswell, D. Parikh, and D. Batra · 2016
Earlier work this paper cites.
Not just a black box: Learning important features through propagating activation differences
A. Shrikumar, P. Greenside, A. Shcherbina, and A. Kundaje · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Quantilizers: A safer alternative to maximizers for limited optimization
J. Taylor · 2016
Earlier work this paper cites.
Alignment for advanced machine learning systems
J. Taylor, E. Yudkowsky, P. LaVictoire, and A. Critch · 2016
Earlier work this paper cites.
T. White · 2016
Earlier work this paper cites.
Corrigibility, 2017
P. Christiano · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. M. A. Patwary, Y. Yang, and Y. Zhou · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
P. W. Koh and P. Liang · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2017
Earlier work this paper cites.
Z. Li and D. Hoiem · 2017
Earlier work this paper cites.
A tutorial on Fisher information
A. Ly, M. Marsman, J. Verhagen, R. P. Grasman, and E.-J. Wagenmakers · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
D. Smilkov, N. Thorat, B. Kim, F. Viégas, and M. Wattenberg · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Earlier work this paper cites.
Correlation and causation between the un human development index and national and personal wealth and resource exploitation
J. Sušnik and P. van der Zaag · 2017
Earlier work this paper cites.
Safety management requirements for defence systems: Part 1: Requirements
UK Ministry of Defence · 2017
Earlier work this paper cites.
A. Vaswani · 2017
Earlier work this paper cites.
Peeking inside the black-box: a survey on explainable artificial intelligence (XAI)
A. Adadi and M. Berrada · 2018
Earlier work this paper cites.
Sanity checks for saliency maps
J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim · 2018
Earlier work this paper cites.
AI and compute
D. Amodei and D. Hernandez · 2018
Earlier work this paper cites.
Learning to evade static PE machine learning malware models via reinforcement learning
H. S. Anderson, A. Kharkar, B. Filar, D. Evans, and P. Roth · 2018
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
S. Arora, Y. Li, Y. Liang, T. Ma, and A. Risteski · 2018
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
A. Athalye, N. Carlini, and D. Wagner · 2018
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
M. Brundage, S. Avin, J. Clark, H. Toner, P. Eckersley, B. Garfinkel, A. Dafoe, P. Scharre, T. Zeitzoff, B. Filar, et al · 2018
Earlier work this paper cites.
The psychology of human-computer interaction
S. K. Card · 2018
Earlier work this paper cites.
An AI race for strategic advantage: rhetoric and risks
S. Cave and S. S. ÓhÉigeartaigh · 2018
Earlier work this paper cites.
Uncertainty in forecasts of long-run economic growth
P. Christensen, K. Gillingham, and W. Nordhaus · 2018
Earlier work this paper cites.
Takeoff speeds, 2018
P. Christiano · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Earlier work this paper cites.
Loss-calibrated approximate inference in bayesian neural networks
A. D. Cobb, S. J. Roberts, and Y. Gal · 2018
Earlier work this paper cites.
Inoculating science against potential pandemics and information hazards
K. M. Esvelt · 2018
Earlier work this paper cites.
T. Everitt, G. Lea, and M. Hutter · 2018
Earlier work this paper cites.
Safety-first AI for autonomous data centre cooling and industrial control
C. Gamble and J. Gao · 2018
Earlier work this paper cites.
J. Gilmer, L. Metz, F. Faghri, S. S. Schoenholz, M. Raghu, M. Wattenberg, and I. Goodfellow · 2018
Earlier work this paper cites.
When will AI exceed human performance? evidence from AI experts
K. Grace, J. Salvatier, A. Dafoe, B. Zhang, and O. Evans · 2018
Earlier work this paper cites.
21 Lessons for the 21st Century
Y. N. Harari · 2018
Earlier work this paper cites.
G. Irving, P. Christiano, and D. Amodei · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
B. Lake and M. Baroni · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Mądry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2018
Earlier work this paper cites.
Understanding learning dynamics of language models with svcca
N. Saphra and A. Lopez · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Earlier work this paper cites.
C4 dataset, 2019
AllenAI · 2019
Earlier work this paper cites.
Understanding multi-head attention in abstractive summarization
J. Baan, M. ter Hoeve, M. van der Wees, A. Schuth, and M. de Rijke · 2019
Earlier work this paper cites.
A. Bajcsy, S. Bansal, E. Bronstein, V. Tolani, and C. J. Tomlin · 2019
Earlier work this paper cites.
The vulnerable world hypothesis
N. Bostrom · 2019
Earlier work this paper cites.
Activation atlas
S. Carter, Z. Armstrong, L. Schubert, I. Johnson, and C. Olah · 2019
Earlier work this paper cites.
Machine learning interpretability: A survey on methods and metrics
D. V. Carvalho, E. M. Pereira, and J. S. Cardoso · 2019
Earlier work this paper cites.
Worst-case guarantees (revisited), 2019
P. Christiano · 2019
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
E. Dinan, S. Humeau, B. Chintagunta, and J. Weston · 2019
Earlier work this paper cites.
Reframing superintelligence: Comprehensive AI services as general intelligence
K. E. Drexler · 2019
Earlier work this paper cites.
Adversarial policies: Attacking deep reinforcement learning
A. Gleave, M. Dennis, C. Wild, N. Kant, S. Levine, and S. Russell · 2019
Earlier work this paper cites.
Gradient hacking
E. Hubinger · 2019
Earlier work this paper cites.
AI safety needs social scientists
G. Irving and A. Askell · 2019
Earlier work this paper cites.
Hidden incentives for auto-induced distributional shift
D. Krueger, T. Maharaj, and J. Leike · 2019
Earlier work this paper cites.
Limitations of the empirical fisher approximation for natural gradient descent
F. Kunstner, P. Hennig, and L. Balles · 2019
Earlier work this paper cites.
V. Lai and C. Tan · 2019
Earlier work this paper cites.
The emergence of number and syntax units in LSTM language models
Y. Lakretz, G. Kruszewski, T. Desbordes, D. Hupkes, S. Dehaene, and M. Baroni · 2019
Earlier work this paper cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
F. Locatello, S. Bauer, M. Lucic, G. Raetsch, S. Gelly, B. Schölkopf, and O. Bachem · 2019
Earlier work this paper cites.
Discovery of natural language concepts in individual units of cnns
S. Na, Y. J. Choe, D.-H. Lee, and G. Kim · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
On the feasibility of learning, rather than assuming, human biases for reward inference
R. Shah, N. Gundotra, P. Abbeel, and A. Dragan · 2019
Earlier work this paper cites.
The bitter lesson
R. Sutton · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
J. Vig and Y. Belinkov · 2019
Earlier work this paper cites.
EDA: Easy data augmentation techniques for boosting performance on text classification tasks
J. Wei and K. Zou · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Thinking about risks from AI: Accidents, misuse and structure
R. Zwetsloot and A. Dafoe · 2019
Earlier work this paper cites.
Towards a human-like open-domain chatbot
D. Adiwardana, M.-T. Luong, D. R. So, J. Hall, N. Fiedel, R. Thoppilan, Z. Yang, A. Kulshreshtha, G. Nemade, Y. Lu, et al · 2020
Earlier work this paper cites.
Neurosymbolic reinforcement learning with formally verified exploration
G. Anderson, A. Verma, I. Dillig, and S. Chaudhuri · 2020
Earlier work this paper cites.
Debate update: Obfuscated arguments problem, 2020
B. Barnes · 2020
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
D. Bau, J.-Y. Zhu, H. Strobelt, A. Lapedriza, B. Zhou, and A. Torralba · 2020
Earlier work this paper cites.
Are ideas getting harder to find?
N. Bloom, C. I. Jones, J. Van Reenen, and M. Webb · 2020
Earlier work this paper cites.
OpenAI API, Jun 2020
G. Brockman, M. Murati, and P. Welinder · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown · 2020
Earlier work this paper cites.
Learning the prior, 2020
P. Christiano · 2020
Earlier work this paper cites.
Forecasting transformative AI with biological anchors, 2020
A. Cotra · 2020
Earlier work this paper cites.
Why we drive: on freedom, risk and taking back control
M. Crawford · 2020
Earlier work this paper cites.
Ai research considerations for human existential safety (ARCHES)
A. Critch and D. Krueger · 2020
Earlier work this paper cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
M. Dennis, N. Jaques, E. Vinitsky, A. Bayen, S. Russell, A. Critch, and S. Levine · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy · 2020
Earlier work this paper cites.
Likelihood of discontinuous progress around the development of AGI, 2020
K. Grace · 2020
Earlier work this paper cites.
Detoxify
L. Hanu and Unitary team · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Measuring the algorithmic efficiency of neural networks
D. Hernandez and T. B. Brown · 2020
Earlier work this paper cites.
Squeeze-and-excitation networks
J. Hu, L. Shen, S. Albanie, G. Sun, and E. Wu · 2020
Earlier work this paper cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
H. J. Jeon, S. Milli, and A. Dragan · 2020
Earlier work this paper cites.
A calculation of the social returns to innovation
B. F. Jones and L. H. Summers · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Specification gaming: the flip side of AI ingenuity
V. Krakovna, J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg · 2020
Earlier work this paper cites.
Explainable AI: A review of machine learning interpretability methods
P. Linardatos, V. Papastefanopoulos, and S. Kotsiantis · 2020
Earlier work this paper cites.
Compositional explanations of neurons
J. Mu and J. Andreas · 2020
Earlier work this paper cites.
Interpreting GPT: the logit lens, 2020
nostalgebraist · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Earlier work this paper cites.
The precipice: Existential risk and the future of humanity
T. Ord · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
G. Pruthi, F. Liu, S. Kale, and M. Sundararajan · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
The primacy order effect in complex decision making
A. Rey, K. Le Goff, M. Abadie, and P. Courrieu · 2020
Earlier work this paper cites.
Improved protein structure prediction using potentials from deep learning
A. W. Senior, R. Evans, J. Jumper, J. Kirkpatrick, L. Sifre, T. Green, C. Qin, A. Žídek, A. W. Nelson, A. Bridgland, et al · 2020
Earlier work this paper cites.
AI alignment 2018-19 review
R. Shah · 2020
Earlier work this paper cites.
Benefits of assistance over reward learning, 2020
R. Shah, P. Freire, N. Alex, R. Freedman, D. Krasheninnikov, L. Chan, M. D. Dennis, P. Abbeel, A. Dragan, and S. Russell · 2020
Earlier work this paper cites.
The offense-defense balance of scientific knowledge: Does publishing AI research reduce misuse?
T. Shevlane and A. Dafoe · 2020
Earlier work this paper cites.
Avoiding tampering incentives in deep RL via decoupled approval
J. Uesato, R. Kumar, V. Krakovna, T. Everitt, R. Ngo, and S. Legg · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer, and S. Shieber · 2020
Cited alongside, same era.
Y. Zhang, Q. V. Liao, and R. K. Bellamy · 2020
Cited alongside, same era.
A review of uncertainty quantification in deep learning: Techniques, applications and challenges
M. Abdar, F. Pourpanah, S. Hussain, D. Rezazadegan, L. Liu, M. Ghavamzadeh, P. Fieguth, X. Cao, A. Khosravi, U. R. Acharya, et al · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, et al · 2021
Cited alongside, same era.
Does the whole exceed its parts? the effect of AI explanations on complementary team performance
Refusal in language models is mediated by a single direction
A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda · 2024
Later among the works it cites.
Training language models to win debates with self-play improves judge accuracy
S. Arnesen, D. Rein, and J. Michael · 2024
Later among the works it cites.
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
L. Bailey, E. Ong, S. Russell, and S. Emmons · 2024
Later among the works it cites.
Towards evaluations-based safety cases for AI scheming
M. Balesni, M. Hobbhahn, D. Lindner, A. Meinke, T. Korbak, J. Clymer, B. Shlegeris, J. Scheurer, C. Stix, R. Shah, et al · 2024
Later among the works it cites.
International scientific report on the safety of advanced AI: Interim report
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar, M. T. Ribeiro, and D. Weld · 2021
Cited alongside, same era.
Imitative generalisation (aka ‘learning the prior’), 2021
B. Barnes · 2021
Cited alongside, same era.
An interpretability illusion for BERT
T. Bolukbasi, A. Pearce, A. Yuan, A. Coenen, E. Reif, F. Viégas, and M. Wattenberg · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Cited alongside, same era.
L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot · 2021
Cited alongside, same era.
Curve circuits
N. Cammarata, G. Goh, S. Carter, C. Voss, L. Schubert, and C. Olah · 2021
Cited alongside, same era.
Mundane solutions to exotic problems, may 2021
P. Christiano · 2021
Cited alongside, same era.
Eliciting latent knowledge: How to tell if your eyes deceive you, Dec 2021
P. Christiano, A. Cotra, and M. Xu · 2021
Cited alongside, same era.
Y. Bengio, B. Fox, et al · 2024
Later among the works it cites.
Sabotage evaluations for frontier models
J. Benton, M. Wagner, E. Christiansen, C. Anil, E. Perez, J. Srivastav, E. Durmus, D. Ganguli, S. Kravec, B. Shlegeris, et al · 2024
Later among the works it cites.
Jailbreaking large language models with symbolic mathematics
E. Bethany, M. Bethany, J. A. N. Flores, S. K. Jha, and P. Najafirad · 2024
Later among the works it cites.
Refining minimax regret for unsupervised environment design
M. Beukman, S. Coward, M. Matthews, M. Fellows, M. Jiang, M. Dennis, and J. Foerster · 2024
Later among the works it cites.
Large language monkeys: Scaling inference compute with repeated sampling
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini · 2024
Later among the works it cites.
M. D. Buhl, G. Sett, L. Koessler, J. Schuett, and M. Anderljung · 2024
Later among the works it cites.
Truth is universal: Robust detection of lies in LLMs
L. Bürger, F. A. Hamprecht, and B. Nadler · 2024
Later among the works it cites.
Stitching SAEs of different sizes, 2024
P. Bussmann, Bart ands Leask, J. Bloom, C. Tigges, and N. Nanda · 2024
Later among the works it cites.
Some lessons from adversarial machine learning
N. Carlini · 2024
Later among the works it cites.
Stealing part of a production language model
N. Carlini, D. Paleka, K. D. Dvijotham, T. Steinke, J. Hayase, A. F. Cooper, K. Lee, M. Jagielski, M. Nasr, A. Conmy, et al · 2024
Later among the works it cites.
Defending against unforeseen failure modes with latent adversarial training
S. Casper, L. Schulze, O. Patel, and D. Hadfield-Menell · 2024
Later among the works it cites.
Ccrl 40/15 rating list — all engines
CCRL · 2024
Later among the works it cites.
Scalable influence and fact tracing for large language model pretraining
T. A. Chang, D. Rajagopal, T. Bolukbasi, L. Dixon, and I. Tenney · 2024
Later among the works it cites.
Designing a dashboard for transparency and control of conversational AI
Y. Chen, A. Wu, T. DePodesta, C. Yeh, K. Li, N. C. Marin, O. Patel, J. Riecke, S. Raval, O. Seow, et al · 2024
Later among the works it cites.
All time top 100 ranklist by highest elo rating, 2016
ChessDB · 2024
Later among the works it cites.
Scaling automatic neuron description, 2024
D. Choi, V. Huang, K. Meng, D. D. Johnson, J. Steinhardt, and S. Schwettmann · 2024
Later among the works it cites.
Arc prize 2024: Technical report
F. Chollet, M. Knoop, G. Kamradt, and B. Landers · 2024
Later among the works it cites.
Gradient routing: Masking gradients to localize computation in neural networks
A. Cloud, J. Goldman-Wetzler, E. Wybitul, J. Miller, and A. M. Turner · 2024
Later among the works it cites.
Safety cases: Justifying the safety of advanced AI systems
J. Clymer, N. Gabrieli, D. Krueger, and T. Larsen · 2024
Later among the works it cites.
A. F. Cooper, C. A. Choquette-Choo, M. Bogen, M. Jagielski, K. Filippova, K. Z. Liu, A. Chouldechova, J. Hayes, Y. Huang, N. Mireshghallah, et al · 2024
Later among the works it cites.
The rising costs of training frontier AI models
B. Cottier, R. Rahman, L. Fattorini, N. Maslej, and D. Owen · 2024
Later among the works it cites.
Towards guaranteed safe AI: A framework for ensuring robust and reliable AI systems
D. Dalrymple, J. Skalse, Y. Bengio, S. Russell, M. Tegmark, S. Seshia, S. Omohundro, C. Szegedy, B. Goldhaber, N. Ammann, et al · 2024
Later among the works it cites.
Do unlearning methods remove information from language model weights?
A. Deeb and F. Roger · 2024
Later among the works it cites.
DeepSeek-R1-Lite-Preview is now live: unleashing supercharged reasoning power!, 2024
DeepSeek · 2024
Later among the works it cites.
MASTERKEY: automated jailbreaking of large language model chatbots
G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu · 2024
Later among the works it cites.
Sycophancy to subterfuge: Investigating reward-tampering in large language models
C. Denison, M. MacDiarmid, F. Barez, D. Duvenaud, S. Kravec, S. Marks, N. Schiefer, R. Soklaski, A. Tamkin, J. Kaplan, et al · 2024
Later among the works it cites.
DiPaCo: Distributed path composition
A. Douillard, Q. Feng, A. A. Rusu, A. Kuncoro, Y. Donchev, R. Chhaparia, I. Gog, M. Ranzato, J. Shen, and A. Szlam · 2024
Later among the works it cites.
h4rm3l: A dynamic benchmark of composable jailbreak attacks for llm safety assessment
M. K. B. Doumbouya, A. Nandi, G. Poesia, D. Ghilardi, A. Goldie, F. Bianchi, D. Jurafsky, and C. D. Manning · 2024
Later among the works it cites.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Later among the works it cites.
Transcoders find interpretable LLM feature circuits
J. Dunefsky, P. Chlenski, and N. Nanda · 2024
Later among the works it cites.
DsDm: Model-aware dataset selection with datamodels
L. Engstrom, A. Feldmann, and A. Madry · 2024
Later among the works it cites.
Data on notable AI models, 2024
Epoch AI · 2024
Later among the works it cites.
Estimating idea production: A methodological survey
E. Erdil, T. Besiroglu, and A. Ho · 2024
Later among the works it cites.
Teams of LLM agents can exploit zero-day vulnerabilities
R. Fang, R. Bindu, A. Gupta, Q. Zhan, and D. Kang · 2024
Later among the works it cites.
Applying sparse autoencoders to unlearn knowledge in language models
E. Farrell, Y.-T. Lau, and A. Conmy · 2024
Later among the works it cites.
Red-teaming for generative AI: Silver bullet or security theater?
M. Feffer, A. Sinha, W. H. Deng, Z. C. Lipton, and H. Heidari · 2024
Later among the works it cites.
On the similarity of circuits across languages: a case study on the subject-verb agreement task
J. Ferrando and M. R. Costa-jussà · 2024
Later among the works it cites.
A primer on the inner workings of transformer-based language models
J. Ferrando, G. Sarti, A. Bisazza, and M. R. Costa-jussà · 2024
Later among the works it cites.
Evaluating superhuman models with consistency checks
L. Fluri, D. Paleka, and F. Tramèr · 2024
Later among the works it cites.
The ethics of advanced AI assistants
I. Gabriel, A. Manzini, G. Keeling, L. A. Hendricks, V. Rieser, H. Iqbal, N. Tomašev, I. Ktena, Z. Kenton, M. Rodriguez, et al · 2024
Later among the works it cites.
Language models scale reliably with over-training and on downstream tasks
S. Y. Gadre, G. Smyrnis, V. Shankar, S. Gururangan, M. Wortsman, R. Shao, J. Mercat, A. Fang, J. Li, S. Keh, et al · 2024
Later among the works it cites.
Bias and fairness in large language models: A survey
I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Dernoncourt, T. Yu, R. Zhang, and N. K. Ahmed · 2024
Later among the works it cites.
Erasing conceptual knowledge from language models
R. Gandikota, S. Feucht, S. Marks, and D. Bau · 2024
Later among the works it cites.
Scaling and evaluating sparse autoencoders
L. Gao, T. D. la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu · 2024
Later among the works it cites.
Algorithmic discrimination in the credit domain: what do we know about it?
A. C. B. Garcia, M. G. P. Garcia, and R. Rigobon · 2024
Later among the works it cites.
Dred: Zero-shot transfer in reinforcement learning via data-regularised environment design
S. Garcin, J. Doran, S. Guo, C. G. Lucas, and S. V. Albrecht · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Later among the works it cites.
Project naptime: Evaluating offensive security capabilities of large language models
S. Glazunov and M. Brand · 2024
Later among the works it cites.
Gemini 2.0 Flash Thinking mode, 2024
Google · 2024
Later among the works it cites.
How google protects its production services, 2024
Google Cloud · 2024
Later among the works it cites.
Thousands of AI authors on the future of AI
K. Grace, H. Stewart, J. F. Sandkühler, S. Thomas, B. Weinstein-Raun, and J. Brauner · 2024
Later among the works it cites.
Seizing the AI for Science opportunity, 2024
C. Griffin, D. Wallace, J. Mateos-Garcia, H. Schieve, and P. Kohli · 2024
Later among the works it cites.
Compact proofs of model performance via mechanistic interpretability
J. Gross, R. Agrawal, T. Kwa, E. Ong, C. H. Yip, A. Gibson, S. Noubir, and L. Chan · 2024
Later among the works it cites.
Three sketches of ASL-4 safety case components
R. Grosse · 2024
Later among the works it cites.
Deliberative alignment: Reasoning enables safer language models
M. Y. Guan, M. Joglekar, E. Wallace, S. Jain, B. Barak, A. Heylar, R. Dias, A. Vallone, H. Ren, J. Wei, et al · 2024
Later among the works it cites.
Mechanistic unlearning: Robust knowledge unlearning and editing via mechanistic localization
P. Guo, A. Syed, A. Sheshadri, A. Ewart, and G. K. Dziugaite · 2024
Later among the works it cites.
Global universal basic skills: Current deficits and implications for world development
S. Gust, E. A. Hanushek, and L. Woessmann · 2024
Later among the works it cites.
Covert malicious finetuning: Challenges in safeguarding llm adaptation
D. Halawi, A. Wei, E. Wallace, T. T. Wang, N. Haghtalab, and J. Steinhardt · 2024
Later among the works it cites.
M. Hanna, O. Liu, and A. Variengien · 2024
Later among the works it cites.
Training large language models to reason in a continuous latent space
S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian · 2024
Later among the works it cites.
Nexus: A brief history of information networks from the Stone Age to AI
Y. N. Harari · 2024
Later among the works it cites.
How to use and interpret activation patching
S. Heimersheim and N. Nanda · 2024
Later among the works it cites.
O. Henkel, L. Hills, A. Boxer, B. Roberts, and Z. Levonian · 2024
Later among the works it cites.
Algorithmic progress in language models
A. Ho, T. Besiroglu, E. Erdil, D. Owen, R. Rahman, Z. C. Guo, D. Atkinson, N. Thompson, and J. Sevilla · 2024
Later among the works it cites.
The developmental landscape of in-context learning
J. Hoogland, G. Wang, M. Farrugia-Roberts, L. Carroll, S. Wei, and D. Murfet · 2024
Later among the works it cites.
Effects of scale on language model robustness
N. Howe, I. McKenzie, O. Hollinsworth, M. Zajac, T. Tseng, A. Tucker, P.-L. Bacon, and A. Gleave · 2024
Later among the works it cites.
Sleeper agents: Training deceptive LLMs that persist through safety training
E. Hubinger, C. Denison, J. Mu, M. Lambert, M. Tong, M. MacDiarmid, T. Lanham, D. M. Ziegler, T. Maxwell, N. Cheng, et al · 2024
Later among the works it cites.
An Introduction to Universal Artificial Intelligence
M. Hutter, D. Quarel, and E. Catt · 2024
Later among the works it cites.
Safety cases at AISI
G. Irving · 2024
Later among the works it cites.
Automation collapse, 2024
G. Irving, T. Korbak, and B. Hilton · 2024
Later among the works it cites.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Later among the works it cites.
Adversaries can misuse combinations of safe models
E. Jones, A. Dragan, and J. Steinhardt · 2024
Later among the works it cites.
Sieve: SAEs beat baselines on a real-world task (a code generation case study)
A. Karvonen, D. Pai, M. Wang, and B. Keigwin · 2024
Later among the works it cites.
The Maes-Garreau point, 2007
K. Kelly · 2024
Later among the works it cites.
On scalable oversight with weak LLMs judging strong LLMs
Z. Kenton, N. Y. Siegel, J. Kramár, J. Brown-Cohen, S. Albanie, J. Bulian, R. Agarwal, D. Lindner, Y. Tang, N. D. Goodman, et al · 2024
Later among the works it cites.
Debating with more persuasive LLMs leads to more truthful answers
A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez · 2024
Later among the works it cites.
The road to artificial superintelligence: A comprehensive survey of superalignment
H. Kim, X. Yi, J. Yao, J. Lian, M. Huang, S. Duan, J. Bak, and X. Xie · 2024
Later among the works it cites.
Economic policy challenges for the age of AI
A. Korinek · 2024
Later among the works it cites.
Semantic entropy probes: Robust and cheap hallucination detection in LLMs
J. Kossen, J. Han, M. Razzak, L. Schut, S. Malik, and Y. Gal · 2024
Later among the works it cites.
AtP*: An efficient and scalable method for localizing LLM behaviour to components
J. Kramár, T. Lieberum, R. Shah, and N. Nanda · 2024
Later among the works it cites.
Me, myself, and AI: The situational awareness dataset (SAD) for LLMs
R. Laine, B. Chughtai, J. Betley, K. Hariharan, M. Balesni, J. Scheurer, M. Hobbhahn, A. Meinke, and O. Evans · 2024
Later among the works it cites.
Cluster-norm for unsupervised probing of knowledge
W. Laurito, S. Maiya, G. Dhimoïla, K. Hänni, et al · 2024
Later among the works it cites.
Programming refusal with conditional activation steering
B. W. Lee, I. Padhi, K. N. Ramamurthy, E. Miehling, P. Dognin, M. Nagireddy, and A. Dhurandhar · 2024
Later among the works it cites.
A comparative study on annotation quality of crowdsourcing and LLM via label aggregation
J. Li · 2024
Later among the works it cites.
A hitchhiker’s guide to jailbreaking chatgpt via prompt engineering
Y. Liu, G. Deng, Z. Xu, Y. Li, Y. Zheng, Y. Zhang, L. Zhao, T. Zhang, and K. Wang · 2024
Later among the works it cites.
Improve mathematical reasoning in language models by automated process supervision
L. Luo, Y. Liu, R. Liu, S. Phatale, M. Guo, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, et al · 2024
Later among the works it cites.
Switching between tasks can cause AI to lose the ability to learn
C. Lyle and R. Pascanu · 2024
Later among the works it cites.
Eight methods to evaluate robust unlearning in LLMs
A. Lynch, P. Guo, A. Ewart, S. Casper, and D. Hadfield-Menell · 2024
Later among the works it cites.
Evolving diverse red-team language models in multi-round multi-agent games
C. Ma, Z. Yang, H. Ci, J. Gao, M. Gao, X. Pan, and Y. Yang · 2024
Later among the works it cites.
Subversion strategy eval: Evaluating ai’s stateless strategic capabilities against control protocols
A. Mallen, C. Griffin, A. Abate, and B. Shlegeris · 2024
Later among the works it cites.
Should users trust advanced AI assistants? Justified trust as a function of competence and alignment
A. Manzini, G. Keeling, N. Marchal, K. R. McKee, V. Rieser, and I. Gabriel · 2024
Later among the works it cites.
Loneliness and suicide mitigation for students using GPT3-enabled chatbots
B. Maples, M. Cerit, A. Vishwanath, and R. Pea · 2024
Later among the works it cites.
Generative AI misuse: A taxonomy of tactics and insights from real-world data
N. Marchal, R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac · 2024
Later among the works it cites.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
S. Marks, C. Rager, E. J. Michaud, Y. Belinkov, D. Bau, and A. Mueller · 2024
Later among the works it cites.
Hidden in plain text: Emergence & mitigation of steganographic collusion in LLMs
Y. Mathew, O. Matthews, R. McCarthy, J. Velja, C. S. de Witt, D. Cope, and N. Schoots · 2024
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, et al · 2024
Later among the works it cites.
LLM critics help catch LLM bugs
N. McAleese, R. M. Pokorny, J. F. C. Uribe, E. Nitishinskaya, M. Trebacz, and J. Leike · 2024
Later among the works it cites.
Understanding and steering Llama 3 with sparse autoencoders, 2024
T. McGrath, D. Balsam, M. Deng, and E. Ho · 2024
Later among the works it cites.
Frontier models are capable of in-context scheming
A. Meinke, B. Schoen, J. Scheurer, M. Balesni, R. Shah, and M. Hobbhahn · 2024
Later among the works it cites.
Monitor: An AI-driven observability interface, 2024
K. Meng, V. Huang, N. Chowdhury, D. Choi, J. Steinhardt, and S. Schwettmann · 2024
Later among the works it cites.
Before 2030, will an AI complete the Turing Test in the Kurzweil/Kapor Longbet?
Metaculus · 2024
Later among the works it cites.
Guidelines for capability elicitation
METR · 2024
Later among the works it cites.
Staying ahead of threat actors in the age of AI
Microsoft Threat Intelligence · 2024
Later among the works it cites.
Transformer circuit faithfulness metrics are not robust
J. Miller, B. Chughtai, and W. Saunders · 2024
Later among the works it cites.
How can AI accelerate science, and how can our government help?, July 2024
T. M. Mitchell · 2024
Later among the works it cites.
Secret collusion among generative AI agents
S. R. Motwani, M. Baranchuk, M. Strohmeier, V. Bolina, P. H. Torr, L. Hammond, and C. S. de Witt · 2024
Later among the works it cites.
The operational risks of AI in large-scale biological attacks, 2024
C. A. Mouton, C. Lucas, and E. Guest · 2024
Later among the works it cites.
A. Mueller, J. Brinkmann, M. Li, S. Marks, K. Pal, N. Prakash, C. Rager, A. Sankaranarayanan, A. S. Sharma, J. Sun, et al · 2024
Later among the works it cites.
Securing AI model weights: Preventing theft and misuse of frontier models
S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H.-A. Bradley, and J. Alstott · 2024
Later among the works it cites.
The alignment problem from a deep learning perspective
R. Ngo, L. Chan, and S. Mindermann · 2024
Later among the works it cites.
The case for CoT unfaithfulness is overstated, 2024
nostalgebraist · 2024
Later among the works it cites.
Response to NIST Executive Order on AI, Feb 2024
OpenAI · 2024
Later among the works it cites.
Disrupting malicious uses of AI by state-affiliated threat actors, 2024
OpenAI · 2024
Later among the works it cites.
Learning to reason with LLMs, 2024
OpenAI · 2024
Later among the works it cites.
How predictable is language model benchmark performance?
D. Owen · 2024
Later among the works it cites.
Consistency checks for language model forecasters
D. Paleka, A. P. Sudhir, A. Alvarez, V. Bhat, A. Shen, E. Wang, and F. Tramèr · 2024
Later among the works it cites.
BOLT: Privacy-preserving, accurate and efficient inference for transformers
Q. Pang, J. Zhu, H. Möllering, W. Zheng, and T. Schneider · 2024
Later among the works it cites.
The geometry of categorical and hierarchical concepts in large language models
K. Park, Y. J. Choe, Y. Jiang, and V. Veitch · 2024
Later among the works it cites.
Preventing memorized completions via white-box filtering
O. Patel and R. Wang · 2024
Later among the works it cites.
GoEX: perspectives and designs towards a runtime for autonomous LLM applications
S. G. Patil, T. Zhang, V. Fang, R. Huang, A. Hao, M. Casado, J. E. Gonzalez, R. A. Popa, I. Stoica, et al · 2024
Later among the works it cites.
Automatically interpreting millions of features in large language models
G. Paulo, A. Mallen, C. Juang, and N. Belrose · 2024
Later among the works it cites.
A. Peppin, A. Reuel, S. Casper, E. Jones, A. Strait, U. Anwar, A. Agrawal, S. Kapoor, S. Koyejo, M. Pellat, et al · 2024
Later among the works it cites.
Let’s think dot by dot: Hidden computation in transformer language models
J. Pfau, W. Merrill, and S. R. Bowman · 2024
Later among the works it cites.
Evaluating frontier models for dangerous capabilities
M. Phuong, M. Aitchison, E. Catt, S. Cogan, A. Kaskasoli, V. Krakovna, D. Lindner, M. Rahtz, Y. Assael, S. Hodkinson, et al · 2024
Later among the works it cites.
Reducing misuse to an access problem
O. Popeti · 2024
Later among the works it cites.
QwQ: Reflect deeply on the boundaries of the unknown, 2024
Qwen Team · 2024
Later among the works it cites.
Improving sparse decomposition of language model activations with gated sparse autoencoders
S. Rajamanoharan, A. Conmy, L. Smith, T. Lieberum, V. Varma, J. Kramár, R. Shah, and N. Nanda · 2024
Later among the works it cites.
Mixture-of-depths: Dynamically allocating compute in transformer-based language models
D. Raposo, S. Ritter, B. Richards, T. Lillicrap, P. C. Humphreys, and A. Santoro · 2024
Later among the works it cites.
AI overviews: About last week, May 2024
L. Reid · 2024
Later among the works it cites.
NYU code debates update/postmortem, 2024
D. Rein · 2024
Later among the works it cites.
A small-molecule TNIK inhibitor targets fibrosis in preclinical and clinical models
F. Ren, A. Aliper, J. Chen, H. Zhao, S. Rao, C. Kuppe, I. V. Ozerov, M. Zhang, K. Witte, C. Kruse, et al · 2024
Later among the works it cites.
Mathematical discoveries from program search with large language models
B. Romera-Paredes, M. Barekatain, A. Novikov, M. Balog, M. P. Kumar, E. Dupont, F. J. Ruiz, J. S. Ellenberg, P. Wang, O. Fawzi, et al · 2024
Later among the works it cites.
Observational scaling laws and the predictability of language model performance
Y. Ruan, C. J. Maddison, and T. Hashimoto · 2024
Later among the works it cites.
Procedural knowledge in pretraining drives reasoning in large language models
L. Ruis, M. Mozes, J. Bae, S. R. Kamalakara, D. Talupuru, A. Locatelli, R. Kirk, T. Rocktäschel, E. Grefenstette, and M. Bartolo · 2024
Later among the works it cites.
Explorations of self-repair in language models
C. Rushing and N. Nanda · 2024
Later among the works it cites.
Rainbow teaming: Open-ended generation of diverse adversarial prompts, 2024
M. Samvelyan, S. C. Raparthy, A. Lupu, E. Hambro, A. H. Markosyan, M. Bhatt, Y. Mao, M. Jiang, J. Parker-Holder, J. Foerster, T. Rocktäschel, and R. Raileanu · 2024
Later among the works it cites.
Why has predicting downstream capabilities of frontier AI models with scale remained elusive?
R. Schaeffer, H. Schoelkopf, B. Miranda, G. Mukobi, V. Madan, A. Ibrahim, H. Bradley, S. Biderman, and S. Koyejo · 2024
Later among the works it cites.
AI-augmented predictions: LLM assistants improve human forecasting accuracy
P. Schoenegger, P. S. Park, E. Karger, S. Trott, and P. E. Tetlock · 2024
Later among the works it cites.
Training compute of frontier AI models grows by 4-5x per year, 2024
J. Sevilla and E. Roldán · 2024
Later among the works it cites.
Can AI scaling continue through 2030?, 2024
J. Sevilla, T. Besiroglu, B. Cottier, J. You, E. Roldán, P. Villalobos, and E. Erdil · 2024
Later among the works it cites.
A multimodal automated interpretability agent
T. R. Shaham, S. Schwettmann, F. Wang, A. Rajaram, E. Hernandez, J. Andreas, and A. Torralba · 2024
Later among the works it cites.
O. Shorinwa, Z. Mei, J. Lidard, A. Z. Ren, and A. Majumdar · 2024
Later among the works it cites.
N. Y. Siegel, O.-M. Camburu, N. Heess, and M. Perez-Ortiz · 2024
Later among the works it cites.
The strong feature hypothesis could be wrong, 2024
L. Smith · 2024
Later among the works it cites.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Later among the works it cites.
A strongREJECT for empty jailbreaks
A. Souly, Q. Lu, D. Bowen, T. Trinh, E. Hsieh, S. Pandey, P. Abbeel, J. Svegliato, S. Emmons, O. Watkins, et al · 2024
Later among the works it cites.
Steering without side effects: Improving post-deployment control of language models
A. C. Stickland, A. Lyzhov, J. Pfau, S. Mahdi, and S. R. Bowman · 2024
Later among the works it cites.
How will advanced AI systems impact democracy?
C. Summerfield, L. Argyle, M. Bakker, T. Collins, E. Durmus, T. Eloundou, I. Gabriel, D. Ganguli, K. Hackenburg, G. Hadfield, et al · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, et al · 2024
Later among the works it cites.
Regulatory compliance on adverse media screening, feb 2024
Thomson Reuters Legal · 2024
Later among the works it cites.
LLM circuit analyses are consistent across training and scale
C. Tigges, M. Hanna, Q. Yu, and S. Biderman · 2024
Later among the works it cites.
Connecting the dots: LLMs can infer and verbalize latent structure from disparate training data
J. Treutlein, D. Choi, J. Betley, S. Marks, C. Anil, R. B. Grosse, and O. Evans · 2024
Later among the works it cites.
Measuring feature sensitivity using dataset filtering, 2024
N. L. Turner, A. Jermyn, and J. Batson · 2024
Later among the works it cites.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
M. Turpin, J. Michael, E. Perez, and S. Bowman · 2024
Later among the works it cites.
When combinations of humans and AI are useful: A systematic review and meta-analysis
M. Vaccaro, A. Almaatouq, and T. Malone · 2024
Later among the works it cites.
Trading off compute in training and inference, 2023
P. Villalobos and D. Atkinson · 2024
Later among the works it cites.
The instruction hierarchy: Training LLMs to prioritize privileged instructions
E. Wallace, K. Xiao, R. Leike, L. Weng, J. Heidecke, and A. Beutel · 2024
Later among the works it cites.
S. Wan, C. Nikolaidis, D. Song, D. Molnar, J. Crnkovich, J. Grace, M. Bhatt, S. Chennabasappa, S. Whitman, S. Ding, et al · 2024
Later among the works it cites.
Jailbroken: How does LLM safety training fail?
A. Wei, N. Haghtalab, and J. Steinhardt · 2024
Later among the works it cites.
Adaptive deployment of untrusted llms reduces distributed threats
J. Wen, V. Hebbar, C. Larson, A. Bhatt, A. Radhakrishnan, M. Sharma, H. Sleight, S. Feng, H. He, E. Perez, et al · 2024
Later among the works it cites.
Readout of President Joe Biden’s meeting with president Xi Jinping of the People’s Republic of China
White House · 2024
Later among the works it cites.
Nuclear astrophysicists at war
M. Wiescher and K. Langanke · 2024
Later among the works it cites.
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
H. Wijk, T. Lin, J. Becker, S. Jawhar, N. Parikh, T. Broadley, L. Chan, M. Chen, J. Clymer, J. Dhyani, et al · 2024
Later among the works it cites.
Less: Selecting influential data for targeted instruction tuning
M. Xia, S. Malladi, S. Gururangan, S. Arora, and D. Chen · 2024
Later among the works it cites.
SafeDecoding: Defending against jailbreak attacks via safety-aware decoding
Z. Xu, F. Jiang, L. Niu, J. Jia, B. Y. Lin, and R. Poovendran · 2024
Later among the works it cites.
Robust LLM safeguarding via refusal feature adversarial training
L. Yu, V. Do, K. Hambardzumyan, and N. Cancedda · 2024
Later among the works it cites.
ShieldGemma: Generative AI content moderation based on gemma
W. Zeng, Y. Liu, R. Mullins, L. Peran, J. Fernandez, H. Harkous, K. Narasimhan, D. Proud, P. Kumar, B. Radharapu, et al · 2024
Later among the works it cites.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, J. Z. Kolter, M. Fredrikson, and D. Hendrycks · 2024
Later among the works it cites.
Trust and safety warnings and appeals, 2025
Anthropic · 2025
Closest in time.
Claude sonnet 3.7 (often) knows when it’s in alignment evaluations, 2025
Apollo · 2025
Closest in time.
Chain-of-thought reasoning in the wild is not always faithful
I. Arcuschin, J. Janiak, R. Krzyzanowski, S. Rajamanoharan, N. Nanda, and A. Conmy · 2025
Closest in time.
Monitoring reasoning models for misbehavior and the risks of promoting obfuscation
B. Baker, J. Huizinga, L. Gao, Z. Dou, M. Y. Guan, A. Madry, W. Zaremba, J. Pachocki, and D. Farhi · 2025
Closest in time.
Open problems in machine unlearning for AI safety
F. Barez, T. Fu, A. Prabhu, S. Casper, A. Sanyal, A. Bibi, A. O’Gara, R. Kirk, B. Bucknall, T. Fist, et al · 2025
Closest in time.
International AI safety report
Y. Bengio, S. Mindermann, D. Privitera, T. Besiroglu, R. Bommasani, S. Casper, Y. Choi, P. Fox, B. Garfinkel, D. Goldfarb, et al · 2025
Closest in time.
Ctrl-Z: Controlling AI agents via resampling, 2025
A. Bhatt, C. Rushing, A. Kaufman, T. Tracy, V. Georgiv, D. Matolcsi, A. Khan, and B. Shlegeris · 2025
Closest in time.
Know your client (KYC), 2025
CFI Team · 2025
Closest in time.
Are DeepSeek R1 and other reasoning models more faithful?
J. Chua and O. Evans · 2025
Closest in time.
Guide to comply with canada’s anti-money laundering and anti-terrorist financing legislation, 2022
CPA Canada · 2025
Closest in time.
Navigating OFAC compliance: The crucial role of IP address geolocation in sanctions screening, 2025
Descartes Systems Group · 2025
Closest in time.
Gate: An integrated assessment model for ai automation, 2025
E. Erdil, A. Potlogea, T. Besiroglu, E. Roldan, A. Ho, J. Sevilla, M. Barnett, M. Vrzla, and R. Sandler · 2025
Closest in time.
MONA: myopic optimization with non-myopic approval can mitigate multi-step reward hacking
S. Farquhar, V. Varma, D. Lindner, D. Elson, C. Biddulph, I. Goodfellow, and R. Shah · 2025
Closest in time.
31 c.f.r. § 1020.315 transactions of exempt persons
Financial Crimes Enforcement Network · 2025
Closest in time.
Detecting strategic deception using linear probes, 2025
N. Goldowsky-Dill, B. Chughtai, S. Heimersheim, and M. Hobbhahn · 2025
Closest in time.
Updating the Frontier Safety Framework
Google DeepMind · 2025
Closest in time.
Safe LoRA: the silver lining of reducing safety risks when fine-tuning large language models
C.-Y. Hsu, Y.-L. Tsai, C.-H. Lin, P.-Y. Chen, C.-M. Yu, and C.-Y. Huang · 2025
Closest in time.
Training on documents about reward hacking induces reward hacking
N. Hu, B. Wright, C. Denison, S. Marks, J. Treutlein, J. Uesato, and E. Hubinger · 2025
Closest in time.
Scaling sparse feature circuits for studying in-context learning, 2025
D. Kharlapenko, S. Shabalin, F. Barez, N. Nanda, and A. Conmy · 2025
Closest in time.
A sketch of an AI control safety case
T. Korbak, J. Clymer, B. Hilton, B. Shlegeris, and G. Irving · 2025
Closest in time.
Gradual disempowerment: Systemic existential risks from incremental AI development
J. Kulveit, R. Douglas, N. Ammann, D. Turan, D. Krueger, and D. Duvenaud · 2025
Closest in time.
Auditing language models for hidden objectives
S. Marks, J. Treutlein, T. Bricken, J. Lindsey, J. Marcus, S. Mishra-Sharma, D. Ziegler, et al · 2025
Closest in time.
Preparing for the intelligence explosion
F. Moorhouse and W. MacAskill · 2025
Closest in time.
Risk management: Threat modelling, 2023
National Cyber Security Centre (NCSC) · 2025
Closest in time.
Disrupting deceptive uses of AI by covert influence operations, may 2024
OpenAI · 2025
Closest in time.
Building an early warning system for LLM-aided biological threat creation, Jan 2024a
OpenAI · 2025
Closest in time.
M. Sharma, M. Tong, J. Mu, J. Wei, J. Kruthoff, S. Goodfriend, E. Ong, A. Peng, R. Agarwal, C. Anil, et al · 2025
Closest in time.
Toward expert-level medical question answering with large language models
K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, K. Clark, S. R. Pfohl, H. Cole-Lewis, et al · 2025
Closest in time.
Self-fulfilling misalignment data might be poisoning our AI models
A. M. Turner · 2025
Closest in time.
AI sandbagging: Language models can strategically underperform on evaluations
T. van der Weij, F. Hofstätter, O. Jaffe, S. F. Brown, and F. R. Ward · 2025
Closest in time.
The rise and potential of large language model based agents: A survey
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, et al · 2025
Closest in time.