Fetching the paper…
Reading the bibliography…
Mathematical reasoning is an increasingly important indicator of large language model (LLM) capabilities, yet we lack understanding of how LLMs process even simple mathematical tasks.
Liii. on lines and planes of closest fit to systems of points in space
F.R.S., K. P · 1901
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Pay attention to MLPs
Liu, H., Dai, Z., So, D., and Le, Q. V · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Earlier work this paper cites.
Toy models of superposition
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C · 2022
Earlier work this paper cites.
Towards understanding grokking: An effective theory of representation learning
Liu, Z., Kitouni, O., Nolte, N., Michaud, E. J., Tegmark, M., and Williams, M · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A. J., and Belinkov, Y · 2022
Earlier work this paper cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Localizing model behavior with path patching, 2023
Goldowsky-Dill, N., MacLeod, C., Sato, L., and Arora, A · 2023
Earlier work this paper cites.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Hanna, M., Liu, O., and Variengien, A · 2023
Earlier work this paper cites.
The linear representation hypothesis and the geometry of large language models
Park, K., Choe, Y. J., and Veitch, V · 2023
Cited alongside, same era.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
Stolfo, A., Belinkov, Y., and Sachan, M · 2023
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2023
Cited alongside, same era.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2023
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges, 2024
Ahn, J., Verma, R., Lou, R., Liu, D., Zhang, R., and Yin, W · 2024
Cited alongside, same era.
Towards principled evaluations of sparse autoencoders for interpretability and control
Makelov, A., Lange, G., and Nanda, N · 2024
Later among the works it cites.
Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A · 2024
Later among the works it cites.
Arithmetic without algorithms: Language models solve math with a bag of heuristics
Nikankin, Y., Reusch, A., Mueller, A., and Belinkov, Y · 2024
Later among the works it cites.
What is a linear representation? what is a multidimensional feature?, 2024
Olah, C. and Jermyn, A · 2024
Later among the works it cites.
Improving dictionary learning with gated sparse autoencoders, 2024
Rajamanoharan, S., Conmy, A., Smith, L., Lieberum, T., Varma, V., Kramár, J., Shah, R., and Nanda, N · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M · 2024
Cited alongside, same era.
Nnsight and ndif: Democratizing access to foundation model internals
Fiotto-Kaufman, J., Loftus, A. R., Todd, E., Brinkmann, J., Juang, C., Pal, K., Rager, C., Mueller, A., Marks, S., Sharma, A. S., Lucchetti, F., Ripa, M., Belfki, A., Prakash, N., Multani, S., Brodley, C., Guha, A., Bell, J., Wallace, B., and Bau, D · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders, 2024
Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J · 2024
Cited alongside, same era.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024
Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., Olsson, C. F., Denain, J.-S., Ho, A., de Oliveira Santos, E., Järviniemi, O., Barnett, M., Sandler, R., Vrzala, M., Sevilla, J., Ren, Q., Pratt, E., Levine, L., Barkley, G., Stewart, N., Grechuk, B., Grechuk, T., Enugandla, S. V., and Wildon, M · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., et al · 2024
Cited alongside, same era.
How to use and interpret activation patching, 2024
Heimersheim, S. and Nanda, N · 2024
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Huben, R., Cunningham, H., Smith, L. R., Ewart, A., and Sharkey, L · 2024
Cited alongside, same era.
Later among the works it cites.
Gemma 2: Improving open language models at a practical size, 2024
Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., et al · 2024
Later among the works it cites.
Can llms master math? investigating large language models on math stack exchange
Satpute, A., Gießing, N., Greiner-Petter, A., Schubotz, M., Teschke, O., Aizawa, A., and Gipp, B · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
Modular addition without black-boxes: Compressing explanations of mlps that compute numerical integration, 2024
Yip, C. H., Agrawal, R., Chan, L., and Gross, J · 2024
Later among the works it cites.
Pre-trained large language models use fourier features to compute addition
Zhou, T., Fu, D., Sharan, V., and Jia, R · 2024
Later among the works it cites.
Fault-tolerant neural networks from biological error correction codes
Zlokapa, A., Tan, A. K., Martyn, J. M., Fiete, I. R., Tegmark, M., and Chuang, I. L · 2024
Later among the works it cites.
interpreting GPT: the logit lens — LessWrong — lesswrong.com
nostalgebraist · 2025
Closest in time.
Language models encode the value of numbers linearly
Zhu, F., Dai, D., and Sui, Z · 2025
Closest in time.