Fetching the paper…
Reading the bibliography…
Current frontier large-language models rely on reasoning to achieve state-of-the-art performance.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Causal abstractions of neural networks, 2021
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models, 2021
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Earlier work this paper cites.
Prompt programming for large language models: Beyond the few-shot paradigm, 2021
Laria Reynolds and Kyle McDonell · 2021
Earlier work this paper cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Chris Olah · 2022
Earlier work this paper cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Shapley value attribution in chain of thought, 2023
Leo Gao · 2023
Earlier work this paper cites.
Roscoe: A suite of metrics for scoring step-by-step reasoning, 2023
Olga Golovneva, Moya Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz · 2023
Cited alongside, same era.
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Cited alongside, same era.
A circuit for python docstrings in a 4-layer attention-only transformer, 2023
Stefan Heimersheim and Jett Janiak · 2023
Cited alongside, same era.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek · 2025
Closest in time.
Causal abstraction: A theoretical foundation for mechanistic interpretability, 2025
Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, and Thomas Icard · 2025
Closest in time.
Under the hood of a reasoning model
Goodfire · 2025
Closest in time.
Chain of thought monitorability: A new and fragile opportunity for ai safety
Tomek Korbak, Mikita Balesni, Elizabeth Barnes, Yoshua Bengio, Joe Benton, Joseph Bloom, Mark Chen, Alan Cooney, Allan Dafoe, Anca Dragan, et al · 2025
Closest in time.
On the biology of a large language model
Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivoire, Thomas Conerly, Chris Olah, and Joshua Batson · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Cited alongside, same era.
Skyler Wu, Eric Meng Shen, Charumathi Badrinath, Jiaqi Ma, and Himabindu Lakkaraju · 2023
Cited alongside, same era.
Forking paths in neural text generation, 2024
Eric Bigelow, Ari Holtzman, Hidenori Tanaka, and Tomer Ullman · 2024
Cited alongside, same era.
Chain-of-thought reasoning in the wild is not always faithful, 2025
Iván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan, Neel Nanda, and Arthur Conmy · 2025
Cited alongside, same era.
Monitoring reasoning models for misbehavior and the risks of promoting obfuscation, 2025
Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y. Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi · 2025
Cited alongside, same era.
Reasoning models don’t always say what they think, 2025
Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Samuel R. Bowman, Jan Leike, Jared Kaplan, and Ethan Perez · 2025
Cited alongside, same era.
Measuring faithfulness in chain-of-thought reasoning, 2023a
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez
Cited in the paper.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al
Cited in the paper.
Closest in time.
s1: Simple test-time scaling, 2025
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
o1: Introducing our first reasoning model
OpenAI · 2025
Closest in time.
An approach to technical agi safety and security, 2025
Rohin Shah, Alex Irpan, Alexander Matt Turner, Anna Wang, Arthur Conmy, David Lindner, Jonah Brown-Cohen, Lewis Ho, Neel Nanda, Raluca Ada Popa, Rishub Jain, Rory Greig, Samuel Albanie, Scott Emmons, Sebastian Farquhar, Sébastien Krier, Senthooran Rajamanoharan, Sophie Bridgers, Tobi Ijitoye, Tom Everitt, Victoria Krakovna, Vikrant Varma, Vladimir Mikulik, Zachary Kenton, Dave Orr, Shane Legg, Noah Goodman, Allan Dafoe, Four Flynn, and Anca Dragan · 2025
Closest in time.
Understanding reasoning in thinking language models via steering vectors
Constantin Venhoff, Iván Arcuschin, Philip Torr, Arthur Conmy, and Neel Nanda · 2025
Closest in time.