Fetching the paper…
Reading the bibliography…
In this paper, we investigate how model distillation impacts the development of reasoning features in large language models (LLMs).
Distilling the knowledge in a neural network
Geoffrey Hinton · 2015
Earlier work this paper cites.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Model compression via distillation and quantization
Antonio Polino, Razvan Pascanu, and Dan Alistarh · 2018
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Implicit representations of meaning in neural language models
Belinda Z Li, Maxwell Nye, and Jacob Andreas · 2021
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark · 2022
Earlier work this paper cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2023
Cited alongside, same era.
Improving activation steering in language models with mean-centring
Ole Jorgensen, Dylan Cope, Nandi Schoots, and Murray Shanahan · 2023
Cited alongside, same era.
Samuel Marks and Max Tegmark · 2023
Cited alongside, same era.
Sparse crosscoders for cross-layer features and model diffing
Jack Lindsey, Adly Templeton, Jonathan Marcus, Thomas Conerly, Joshua Batson, and Christopher Olah · 2024
Later among the works it cites.
Opening the ai black box: program synthesis via mechanistic interpretability
Eric J Michaud, Isaac Liao, Vedang Lad, Ziming Liu, Anish Mudide, Chloe Loughridge, Zifan Carl Guo, Tara Rezaei Kheirkhah, Mateja Vukelić, and Max Tegmark · 2024
Later among the works it cites.
The geometry of categorical and hierarchical concepts in large language models
Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch · 2024
Later among the works it cites.
Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders
Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mansi Sakarvadia, Aswathy Ajith, Arham Khan, Daniel Grzenda, Nathaniel Hudson, André Bauer, Kyle Chard, and Ian Foster · 2023
Cited alongside, same era.
Modeldiff: A framework for comparing learning algorithms
Harshay Shah, Sung Min Park, Andrew Ilyas, and Aleksander Madry · 2023
Cited alongside, same era.
Function vectors in large language models
Eric Todd, Millicent L Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau · 2023
Cited alongside, same era.
Steering language models with activation engineering
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid · 2023
Cited alongside, same era.
Generalization from starvation: Hints of universality in llm knowledge graph learning
David D Baek, Yuxiao Li, and Max Tegmark · 2024
Cited alongside, same era.
Not all language model features are linear
Joshua Engels, Eric J Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Cited alongside, same era.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Cited alongside, same era.
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, et al · 2024
Later among the works it cites.
Extending activation steering to broad skills and multiple behaviours
Teun van der Weij, Massimo Poesio, and Nandi Schoots · 2024
Later among the works it cites.
A survey on knowledge distillation of large language models
Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou · 2024
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2024
Later among the works it cites.
Harmonic loss trains interpretable ai models
David D Baek, Ziming Liu, Riya Tyagi, and Max Tegmark · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Language models use trigonometry to do addition
Subhash Kantamneni and Max Tegmark · 2025
Closest in time.
Transcoders beat sparse autoencoders for interpretability
Gonçalo Paulo, Stepan Shabalin, and Nora Belrose · 2025
Closest in time.