Fetching the paper…
Reading the bibliography…
Large Reasoning Models (LRMs) have shown impressive capabilities in multi-step reasoning tasks.
The cambridge dictionary of statistics
B. Everitt · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
The conjunction fallacy: A misunderstanding about conjunction?
K. Tentori, N. Bonini, and D. Osherson · 2004
Earlier work this paper cites.
Thinking, fast and slow
D. Kahneman · 2011
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
X. Ho, A.-K. D. Nguyen, S. Sugawara, and A. Aizawa · 2020
Earlier work this paper cites.
Uncertainty estimation in autoregressive structured prediction
A. Malinin and M. Gales · 2020
Earlier work this paper cites.
Interpreting GPT: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Tversky and kahneman’s cognitive illusions: who can solve them, and why?
G. Bruckmaier, S. Krauss, K. Binder, S. Hilbert, and M. Brunner · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
M. Geva, R. Schuster, J. Berant, and O. Levy · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
Language models (mostly) know what they know
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al · 2022
Earlier work this paper cites.
Measuring and narrowing the compositionality gap in language models
O. Press, M. Zhang, S. Min, L. Schmidt, N. A. Smith, and M. Lewis · 2022
Earlier work this paper cites.
Out-of-distribution detection and selective generation for conditional language models
J. Ren, J. Luo, Y. Zhao, K. Krishna, M. Saleh, B. Lakshminarayanan, and P. J. Liu · 2022
Earlier work this paper cites.
Musique: Multihop questions via single-hop question composition
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Lm vs lm: Detecting factual errors via cross examination
R. Cohen, M. Hamri, M. Geva, and A. Globerson · 2023
Cited alongside, same era.
Chainpoll: A high efficacy method for llm hallucination detection
R. Friel and A. Sanyal · 2023
Cited alongside, same era.
Let’s verify step by step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
P. Manakul, A. Liusie, and M. Gales · 2023
Cited alongside, same era.
Receval: Evaluating reasoning chains via correctness and informativeness
A. Prasad, S. Saha, X. Zhou, and M. Bansal · 2023
Cited alongside, same era.
Investigating truthfulness in a pre-release o3 model, 2024
Transluce Research · 2024
Later among the works it cites.
Llms still can’t plan; can lrms? a preliminary evaluation of openai’s o1 on planbench
K. Valmeekam, K. Stechly, and S. Kambhampati · 2024
Later among the works it cites.
Switching between different cognitive strategies induces switch costs as evidenced by switches between manual and mental object rotation
P. P. Weis and W. Kunde · 2024
Later among the works it cites.
Retrieval head mechanistically explains long-context factuality
W. Wu, Y. Wang, G. Xiao, H. Peng, and Y. Fu · 2024
Later among the works it cites.
Can we verify step by step for incorrect answer detection?
X. Xu, S. Diao, C. Yang, and Y. Wang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
A. Stolfo, Y. Belinkov, and M. Sachan · 2023
Cited alongside, same era.
Characterizing mechanisms for factual recall in language models
Q. Yu, J. Merullo, and E. Pavlick · 2023
Cited alongside, same era.
Inside: Llms’ internal states retain the power of hallucination detection
C. Chen, K. Liu, Z. Chen, Y. Gu, Y. Wu, M. Tao, Z. Fu, and J. Ye · 2024
Cited alongside, same era.
Information flow routes: Automatically interpreting language models at scale
J. Ferrando and E. Voita · 2024
Cited alongside, same era.
A primer on the inner workings of transformer-based language models
J. Ferrando, G. Sarti, A. Bisazza, and M. R. Costa-jussà · 2024
Cited alongside, same era.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
M. Hanna, O. Liu, and A. Variengien · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space
S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian · 2024
Cited alongside, same era.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al · 2024
Later among the works it cites.
Unibias: Unveiling and mitigating llm bias through internal attention and ffn manipulation
H. Zhou, Z. Feng, Z. Zhu, J. Qian, and K. Mao · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D.-A. et al · 2025
Closest in time.
Can large language models detect errors in long chain-of-thought reasoning?, 2025
Y. He, S. Li, J. Liu, W. Wang, X. Bu, G. Zhang, Z. Peng, Z. Zhang, Z. Zheng, W. Su, and B. Zheng · 2025
Closest in time.
Math-verify: A rule-based mathematical answer verification library, 2025
Hugging Face · 2025
Closest in time.
Arithmetic without algorithms: Language models solve math with a bag of heuristics
Y. Nikankin, A. Reusch, A. Mueller, and Y. Belinkov · 2025
Closest in time.
Openai o3 and o4-mini system card, April 2025
OpenAI · 2025
Closest in time.
Deepseek-r1 hallucinates more than deepseek-v3, 2025
Vectara Research · 2025
Closest in time.
Do larger language models imply better reasoning? a pretraining scaling law for reasoning
X. Wang, S. Tan, M. Jin, W. Y. Wang, R. Panda, and Y. Shen · 2025
Closest in time.
K. Yan, Y. Xu, Z. Du, X. Yao, Z. Wang, X. Guo, and J. Chen · 2025
Closest in time.
Z. Zeng, Q. Cheng, Z. Yin, Y. Zhou, and X. Qiu · 2025
Closest in time.
The lessons of developing process reward models in mathematical reasoning
Z. Zhang, C. Zheng, Y. Wu, B. Zhang, R. Lin, B. Yu, D. Liu, J. Zhou, and J. Lin · 2025
Closest in time.