Fetching the paper…
Reading the bibliography…
In this position paper, we argue that understanding the relation between structure in the data distribution and structure in trained models is central to AI alignment.
Risks from learned optimization in advanced machine learning systems, June 2019
Hubinger, E., Merwijk, C. v., Mikulik, V., Skalse, J., and Garrabrant, S · 1906
Earlier work this paper cites.
Emergent properties of the local geometry of neural loss landscapes, October 2019
Fort, S. and Ganguli, S · 1910
Earlier work this paper cites.
Timaeus. Critias. Cleitophon. Menexenus. Epistles , chapter 1, pp. 1–254
Bury, R. G · 1929
Earlier work this paper cites.
Timaeus
Plato · 1929
Earlier work this paper cites.
The Essential Turing: Seminal Writings in Computing, Logic, Philosophy, Artificial Intelligence, and Artificial Life: Plus The Secrets of Enigma , chapter 10, pp. 404–432
Copeland, B. J. (ed.) · 1948
Earlier work this paper cites.
Some moral and technical consequences of automation
Wiener, N · 1960
Earlier work this paper cites.
A formal theory of inductive inference. Part I
Solomonoff, R. J · 1964
Earlier work this paper cites.
Speculations concerning the first ultraintelligent machine
Good, I. J · 1966
Earlier work this paper cites.
From Watt to Clausius: The Rise of Thermodynamics in the Early Industrial Age
Cardwell, D. S. L · 1971
Earlier work this paper cites.
The Internal Constitution of the Stars
Eddington, A. S · 1988
Earlier work this paper cites.
Stochastic Complexity in Statistical Inquiry , volume 15
Rissanen, J · 1989
Earlier work this paper cites.
The Lever of Riches: Technological Creativity and Economic Progress
Mokyr, J · 1992
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Schmidhuber, J · 1992
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Statistical modeling: The two cultures (with comments and a rejoinder by the author)
Breiman, L · 2001
Earlier work this paper cites.
Scaling laws for neural language models, January 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Semantic Cognition
Rogers, T. T. and McClelland, J. L · 2004
Earlier work this paper cites.
Intelligent machinery
Turing, A · 2004
Earlier work this paper cites.
Coherent extrapolated volition
Yudkowsky, E · 2004
Earlier work this paper cites.
Principles of Statistical Inference
Cox, D. R · 2006
Earlier work this paper cites.
The hidden complexity of wishes
Yudkowsky, E · 2007
Earlier work this paper cites.
Algebraic Geometry and Statistical Learning Theory
Watanabe, S · 2009
Earlier work this paper cites.
Free energy and dendritic self-organization
Kiebel, S. J. and Friston, K. J · 2011
Earlier work this paper cites.
Between order and chaos
Crutchfield, J. P · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N · 2014
Earlier work this paper cites.
Developmental Cognitive Neuroscience: An Introduction
Johnson, M. H. and de Haan, M. D · 2015
Earlier work this paper cites.
Visualizing representations: Deep learning and human beings, January 2015
Olah, C · 2015
Earlier work this paper cites.
Thermodynamics in Nuclear Power Plant Systems
Zohuri, B. and McDaniel, P · 2015
Earlier work this paper cites.
Concrete problems in AI safety, June 2016
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
What does the universal prior actually look like?
Christiano, P · 2016
Earlier work this paper cites.
Model compression and acceleration for deep neural networks: The principles, progress, and challenges
Cheng, Y., Wang, D., Zhou, P., and Zhang, T · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically, December 2017
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information, March 2017
Shwartz-Ziv, R. and Tishby, N · 2017
Earlier work this paper cites.
Critical learning periods in deep networks
Achille, A., Rovere, M., and Soatto, S · 2018
Earlier work this paper cites.
AI and compute
Amodei, D. and Hernandez, D · 2018
Earlier work this paper cites.
Clarifying AI alignment
Christiano, P · 2018
Earlier work this paper cites.
AGI safety literature review
Everitt, T., Lea, G., and Hutter, M · 2018
Earlier work this paper cites.
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
The learning theoretic alignment agenda
Kosoy, V · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review, May 2018
Levine, S · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Cited alongside, same era.
The basic AI drives
Omohundro, S. M · 2018
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2018
Managing extreme AI risks amid rapid progress, October 2023
Bengio, Y., Hinton, G., Yao, A., Song, D., Abbeel, P., Darrell, T., Harari, Y. N., Zhang, Y.-Q., Xue, L., Shalev-Shwartz, S., Hadfield, G., Clune, J., Maharaj, T., Hutter, F., Baydin, A. G., McIlraith, S., Gao, Q., Acharya, A., Krueger, D., Dragan, A., Torr, P., Russell, S., Kahneman, D., Brauner, J., and Mindermann, S · 2023
Later among the works it cites.
The scaling hypothesis, 2020
Branwen, G · 2023
Later among the works it cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al · 2023
Later among the works it cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T. T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P. J., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Biyik, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mathematical Theory of Bayesian Statistics
Watanabe, S · 2018
Cited alongside, same era.
What failure looks like
Christiano, P · 2019
Cited alongside, same era.
Are minimal circuits deceptive?
Hubinger, E · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Cited alongside, same era.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Papyan, V · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Cited alongside, same era.
Natural abstractions: Key claims, theorems, and critiques
Chan, L., Lang, L., and Jenner, E · 2023
Later among the works it cites.
A toy model of universality: Reverse engineering how networks learn group operations
Chughtai, B., Chan, L., and Nanda, N · 2023
Later among the works it cites.
Language modeling is compression
Delétang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., Hutter, M., and Veness, J · 2023
Later among the works it cites.
BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B, October 2023
Gade, P., Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Later among the works it cites.
Lieberum, T., Rahtz, M., Kramár, J., Nanda, N., Irving, G., Shah, R., and Mikulik, V · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Later among the works it cites.
Deep deceptiveness
Soares, N · 2023
Later among the works it cites.
Meta-posterior consistency for the bayesian inference of metastable system, August 2024
Adams, Z. P. and Mukherjee, S · 2024
Later among the works it cites.
Foundational challenges in assuring alignment and safety of large language models
Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E. S., Jenner, E., Casper, S., Sourbut, O., Edelman, B. L., Zhang, Z., Günther, M., Korinek, A., Hernandez-Orallo, J., Hammond, L., Bigelow, E. J., Pan, A., Langosco, L., Korbak, T., Zhang, H. C., Zhong, R., hÉigeartaigh, S. O., Recchia, G., Corsi, G., Chan, A., Anderljung, M., Edwards, L., Petrov, A., de Witt, C. S., Motwani, S. R., Bengio, Y., Chen, D., Torr, P., Albanie, S., Maharaj, T., Foerster, J. N., Tramèr, F., He, H., Kasirzadeh, A., Choi, Y., and Krueger, D · 2024
Later among the works it cites.
RL, but don’t do anything I wouldn’t do, 2024
Cohen, M. K., Hutter, M., Bengio, Y., and Russell, S · 2024
Later among the works it cites.
AI control: Improving safety despite intentional subversion
Greenblatt, R., Shlegeris, B., Sachan, K., and Roger, F · 2024
Later among the works it cites.
Deliberative alignment: Reasoning enables safer language models, December 2024
Guan, M. Y., Joglekar, M., Wallace, E., Jain, S., Barak, B., Helyar, A., Dias, R., Vallone, A., Ren, H., Wei, J., Chung, H. W., Toyer, S., Heidecke, J., Beutel, A., and Glaese, A · 2024
Later among the works it cites.
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
He, T., Doshi, D., Das, A., and Gromov, A · 2024
Later among the works it cites.
The developmental landscape of in-context learning, February 2024
Hoogland, J., Wang, G., Farrugia-Roberts, M., Carroll, L., Wei, S., and Murfet, D · 2024
Later among the works it cites.
Sleeper agents: Training deceptive LLMs that persist through safety training, January 2024
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., Ganguli, D., Barez, F., Clark, J., Ndousse, K., Sachan, K., Sellitto, M., Sharma, M., DasSarma, N., Grosse, R., Kravec, S., Bai, Y., Witten, Z., Favaro, M., Brauner, J., Karnofsky, H., Christiano, P., Bowman, S. R., Graham, L., Kaplan, J., Mindermann, S., Greenblatt, R., Shlegeris, B., Schiefer, N., and Perez, E · 2024
Later among the works it cites.
The platonic representation hypothesis
Huh, M., Cheung, B., Wang, T., and Isola, P · 2024
Later among the works it cites.
The queen’s dilemma: A paradox of control
Murfet, D · 2024
Later among the works it cites.
OpenAI o1 system card, December 2024
OpenAI, :, Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., Iftimie, A., Karpenko, A., Passos, A. T., Neitz, A., Prokofiev, A., Wei, A., Tam, A., Bennett, A., Kumar, A., Saraiva, A., Vallone, A., Duberstein, A., Kondrich, A., Mishchenko, A., Applebaum, A., Jiang, A., Nair, A., Zoph, B., Ghorbani, B., Rossen, B., Sokolowsky, B., Barak, B., McGrew, B., Minaiev, B., Hao, B., Baker, B., Houghton, B., McKinzie, B., Eastman, B., Lugaresi, C., Bassin, C., Hudson, C., Li, C. M., Bourcy, C. d., Voss, C., Shen, C., Zhang, C., Koch, C., Orsinger, C., Hesse, C., Fischer, C., Chan, C., Roberts, D., Kappler, D., Levy, D., Selsam, D., Dohan, D., Farhi, D., Mely, D., Robinson, D., Tsipras, D., Li, D., Oprica, D., Freeman, E., Zhang, E., Wong, E., Proehl, E., Cheung, E., Mitchell, E., Wallace, E., Ritter, E., Mays, E., Wang, F., Such, F. P., Raso, F., Leoni, F., Tsimpourlas, F., Song, F., Lohmann, F. v., Sulit, F., Salmon, G., Parascandolo, G., Chabot, G., Zhao, G., Brockman, G., Leclerc, G., Salman, H., Bao, H., Sheng, H., Andrin, H., Bagherinezhad, H., Ren, H., Lightman, H., Chung, H. W., Kivlichan, I., O’Connell, I., Osband, I., Gilaberte, I. C., Akkaya, I., Kostrikov, I., Sutskever, I., Kofman, I., Pachocki, J., Lennon, J., Wei, J., Harb, J., Twore, J., Feng, J., Yu, J., Weng, J., Tang, J., Yu, J., Candela, J. Q., Palermo, J., Parish, J., Heidecke, J., Hallman, J., Rizzo, J., Gordon, J., Uesato, J., Ward, J., Huizinga, J., Wang, J., Chen, K., Xiao, K., Singhal, K., Nguyen, K., Cobbe, K., Shi, K., Wood, K., Rimbach, K., Gu-Lemberg, K., Liu, K., Lu, K., Stone, K., Yu, K., Ahmad, L., Yang, L., Liu, L., Maksin, L., Ho, L., Fedus, L., Weng, L., Li, L., McCallum, L., Held, L., Kuhn, L., Kondraciuk, L., Kaiser, L., Metz, L., Boyd, M., Trebacz, M., Joglekar, M., Chen, M., Tintor, M., Meyer, M., Jones, M., Kaufer, M., Schwarzer, M., Shah, M., Yatbaz, M., Guan, M. Y., Xu, M., Yan, M., Glaese, M., Chen, M., Lampe, M., Malek, M., Wang, M., Fradin, M., McClay, M., Pavlov, M., Wang, M., Wang, M., Murati, M., Bavarian, M., Rohaninejad, M., McAleese, N., Chowdhury, N., Chowdhury, N., Ryder, N., Tezak, N., Brown, N., Nachum, O., Boiko, O., Murk, O., Watkins, O., Chao, P., Ashbourne, P., Izmailov, P., Zhokhov, P., Dias, R., Arora, R., Lin, R., Lopes, R. G., Gaon, R., Miyara, R., Leike, R., Hwang, R., Garg, R., Brown, R., James, R., Shu, R., Cheu, R., Greene, R., Jain, S., Altman, S., Toizer, S., Toyer, S., Miserendino, S., Agarwal, S., Hernandez, S., Baker, S., McKinney, S., Yan, S., Zhao, S., Hu, S., Santurkar, S., Chaudhuri, S. R., Zhang, S., Fu, S., Papay, S., Lin, S., Balaji, S., Sanjeev, S., Sidor, S., Broda, T., Clark, A., Wang, T., Gordon, T., Sanders, T., Patwardhan, T., Sottiaux, T., Degry, T., Dimson, T., Zheng, T., Garipov, T., Stasi, T., Bansal, T., Creech, T., Peterson, T., Eloundou, T., Qi, V., Kosaraju, V., Monaco, V., Pong, V., Fomenko, V., Zheng, W., Zhou, W., McCabe, W., Zaremba, W., Dubois, Y., Lu, Y., Chen, Y., Cha, Y., Bai, Y., He, Y., Zhang, Y., Wang, Y., Shao, Z., and Li, Z · 2024
Later among the works it cites.
Position: Understanding LLMs requires more than statistical generalization
Reizinger, P., Ujváry, S., Mészáros, A., Kerekes, A., Brendel, W., and Huszár, F · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
LLM circuit analyses are consistent across training and scale
Tigges, C., Hanna, M., Yu, Q., and Biderman, S · 2024
Later among the works it cites.
Structure Development in List Sorting Transformers
Urdshals, E. and Nasufi, J · 2024
Later among the works it cites.
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
Wang, G., Hoogland, J., Van Wingerden, S., Furman, Z., and Murfet, D · 2024
Later among the works it cites.
Watanabe, S · 2024
Later among the works it cites.
Open problems in machine unlearning for AI safety, January 2025
Barez, F., Fu, T., Prabhu, A., Casper, S., Sanyal, A., Bibi, A., O’Gara, A., Kirk, R., Bucknall, B., Fist, T., Ong, L., Torr, P., Lam, K.-Y., Trager, R., Krueger, D., Mindermann, S., Hernandez-Orallo, J., Geva, M., and Gal, Y · 2025
Closest in time.
International AI safety report: The international scientific report on the safety of advanced AI
Bengio, Y., Mindermann, S., Privitera, D., Besiroglu, T., Bommasani, R., Casper, S., Choi, Y., Fox, P., Garfinkel, B., Goldfarb, D., Heidari, H., Ho, A., Kapoor, S., Khalatbari, L., Longpre, S., Manning, S., Mavroudis, V., Mazeika, M., Michael, J., Newman, J., Ng, K. Y., Okolo, C. T., Raji, D., Sastry, G., Seger, E., Skeadas, T., South, T., Strubell, E., Tramèr, F., Velasco, L., Wheeler, N., Acemoglu, D., Adekanmbi, O., Dalrymple, D., Dietterich, T. G., Felten, E. W., Fung, P., Gourinchas, P.-O., Heintz, F., Hinton, G., Jennings, N., Krause, A., Leavy, S., Liang, P., Ludermir, T., Marda, V., Margetts, H., McDermid, J., Munga, J., Narayanan, A., Nelson, A., Neppel, C., Oh, A., Ramchurn, G., Russell, S., Schaake, M., Schölkopf, B., Song, D., Soto, A., Tiedrich, L., Varoquaux, G., Yao, A., Zhang, Y.-Q., Ajala, O., Albalawi, F., Alserkal, M., Avrin, G., Busch, C., de Carvalho, A. C. P. d. L. F., Fox, B., Gill, A. S., Hatip, A. H., Heikkilä, J., Johnson, C., Jolly, G., Katzir, Z., Khan, S. M., Kitano, H., Krüger, A., Lee, K. M., Ligot, D. V., López Portillo, J. R., Molchanovskyi, O., Monti, A., Mwamanzi, N., Nemer, M., Oliver, N., Pezoa Rivera, R., Ravindran, B., Riza, H., Rugege, C., Seoighe, C., Sheehan, J., Sheikh, H., Wong, D., and Zeng, Y · 2025
Closest in time.
Dynamics of transient structure in in-context linear regression transformers, January 2025
Carroll, L., Hoogland, J., Farrugia-Roberts, M., and Murfet, D · 2025
Closest in time.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning, January 2025
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Bao, H., Xu, H., Wang, H., Ding, H., Xin, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Wang, J., Chen, J., Yuan, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Ye, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Zhao, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Xu, Y., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z · 2025
Closest in time.
Introduction to AI Safety, Ethics, and Society
Hendrycks, D · 2025
Closest in time.
A sketch of an AI control safety case, January 2025
Korbak, T., Clymer, J., Hilton, B., Shlegeris, B., and Irving, G · 2025
Closest in time.
The local learning coefficient: A singularity-aware complexity measure
Lau, E., Furman, Z., Wang, G., Murfet, D., and Wei, S · 2025
Closest in time.
Open problems in mechanistic interpretability, January 2025
Sharkey, L., Chughtai, B., Batson, J., Lindsey, J., Wu, J., Bushnaq, L., Goldowsky-Dill, N., Heimersheim, S., Ortega, A., Bloom, J., Biderman, S., Garriga-Alonso, A., Conmy, A., Nanda, N., Rumbelow, J., Wattenberg, M., Schoots, N., Miller, J., Michaud, E. J., Casper, S., Tegmark, M., Saunders, W., Bau, D., Todd, E., Geiger, A., Geva, M., Hoogland, J., Murfet, D., and McGrath, T · 2025
Closest in time.