Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are typically trained by reinforcement learning (RL) with verifiable rewards (RLVR) and supervised fine-tuning (SFT) on reasoning traces to improve their reasoning abilities.
A Statistical Method for Evaluating Systematic Relationships
R.R. Sokal, C.D. Michener, and University of Kansas · 1958
Earlier work this paper cites.
Hierarchical grouping to optimize an objective function
Jr. Ward, Joe H · 1963
Earlier work this paper cites.
Algebraic connectivity of graphs
Miroslav Fiedler · 1973
Earlier work this paper cites.
A set of measures of centrality based on betweenness
Linton C Freeman · 1977
Earlier work this paper cites.
Centrality in social networks conceptual clarification
Linton C. Freeman · 1978
Earlier work this paper cites.
Collective dynamics of ‘small-world’ networks
Duncan J. Watts and Steven H. Strogatz · 1998
Earlier work this paper cites.
Efficient behavior of small-world networks
Vito Latora and Massimo Marchiori · 2001
Earlier work this paper cites.
Community structure in social and biological networks
M. Girvan and M. E. J. Newman · 2002
Earlier work this paper cites.
Assortative mixing in networks
M. E. J. Newman · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
The structure and function of complex networks
M. E. J. Newman · 2003
Earlier work this paper cites.
Superfamilies of evolved and designed networks
Ron Milo, Shalev Itzkovitz, Nadav Kashtan, Reuven Levitt, Shai Shen-Orr, Inbal Ayzenshtat, Michal Sheffer, and Uri Alon · 2004
Earlier work this paper cites.
Modeling interactome: scale-free or geometric?
N. Pržulj, D. G. Corneil, and I. Jurisica · 2004
Earlier work this paper cites.
Efficient estimation of graphlet frequency distributions in protein–protein interaction networks
N. Pržulj, D. G. Corneil, and I. Jurisica · 2006
Earlier work this paper cites.
Biological network comparison using graphlet degree distribution
Nataša Pržulj · 2007
Earlier work this paper cites.
Model Selection for Social Networks Using Graphlets
Jeannette Janssen, Matt Hurshman, and Nauzer Kalyaniwalla · 2012
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović · 2015
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, and Greg Brockman et al · 2021
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Large language models can self-improve
Jiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han · 2023
Earlier work this paper cites.
Why think step by step? reasoning emerges from the locality of experience
Ben Prystawski, Michael Y. Li, and Noah Goodman · 2023
Earlier work this paper cites.
Stream of search (SoS): Learning to search in language
Kanishk Gandhi, Denise H J Lee, Gabriel Grand, Muxin Liu, Winson Cheng, Archit Sharma, and Noah Goodman · 2024
Earlier work this paper cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, and Alex Vaughan et al · 2024
Earlier work this paper cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, and Alex Carney et al · 2024
Earlier work this paper cites.
Language model self-improvement by reinforcement learning contemplation
Jing-Cheng Pang, Pengyuan Wang, Kaiyuan Li, Xiong-Hui Chen, Jiacheng Xu, Zongzhang Zhang, and Yang Yu · 2024
Cited alongside, same era.
DeepSeekMath: Pushing the limits of mathematical reasoning in open language models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Cited alongside, same era.
Understanding reasoning ability of language models from the perspective of reasoning paths aggregation
Xinyi Wang, Alfonso Amayuelas, Kexun Zhang, Liangming Pan, Wenhu Chen, and William Yang Wang · 2024
Cited alongside, same era.
C-pack: Packed resources for general chinese embeddings
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie · 2024
Cited alongside, same era.
Qwen2.5-Math technical report: Toward mathematical expert model via self-improvement
Topology of reasoning: Understanding large reasoning models through reasoning graph properties
Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa, and Yutaka Matsuo · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
Heuristics considered harmful: Rl with random rewards should not make LLMs reason, 2025
Owen Oertell, Wenhao Zhan, Gokul Swamy, Zhiwei Steven Wu, Kiante Brantley, Jason Lee, and Wen Sun · 2025
Closest in time.
Maximizing confidence alone improves reasoning
Mihir Prabhudesai, Lili Chen, Alex Ippoliti, Katerina Fragkiadaki, Hao Liu, and Deepak Pathak · 2025
Closest in time.
Decomposing elements of problem solving: What ”math” does rl teach?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Keming Lu, Mingfeng Xue, Runji Lin, Tianyu Liu, Xingzhang Ren, and Zhenru Zhang · 2024
Cited alongside, same era.
mGTE: Generalized long-context text representation and reranking models for multilingual text retrieval
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang · 2024
Cited alongside, same era.
Thought Anchors: Which LLM reasoning steps matter?
Paul C. Bogdan, Uzay Macar, Neel Nanda, and Arthur Conmy · 2025
Cited alongside, same era.
Incorrect baseline evaluations call into question recent LLM-RL claims, 2025
Nikhil Chandak, Shashwat Goel, and Ameya Prabhu · 2025
Cited alongside, same era.
Reasoning with exploration: An entropy perspective on reinforcement learning for LLMs
Daixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai, Wayne Xin Zhao, Zhenliang Zhang, and Furu Wei · 2025
Cited alongside, same era.
SFT memorizes, RL generalizes: A comparative study of foundation model post-training
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V Le, Sergey Levine, and Yi Ma · 2025
Cited alongside, same era.
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, and Evan Rosen et al · 2025
Cited alongside, same era.
Weight ensembling improves reasoning in language models
Xingyu Dang, Christina Baek, Kaiyue Wen, Zico Kolter, and Aditi Raghunathan · 2025
Cited alongside, same era.
Tian Qin, Core Francisco Park, Mujin Kwun, Aaron Walsman, Eran Malach, Nikhil Anand, Hidenori Tanaka, and David Alvarez-Melis · 2025
Closest in time.
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu · 2025
Closest in time.
e3: Learning to explore enables extrapolation of test-Time compute for LLMs
Amrith Setlur, Matthew Y. R. Yang, Charlie Victor Snell, Jeremiah Greer, Ian Wu, Virginia Smith, Max Simchowitz, and Aviral Kumar · 2025
Closest in time.
Rethinking reflection in pre-Training
Darsh J Shah, Peter Rushton, Somanshu Singla, Mohit Parmar, Kurt Smith, Yash Vanjani, Ashish Vaswani, Adarsh Chaluvaraju, Andrew Hojel, Andrew Ma, Anil Thomas, Anthony Polloreno, Ashish Tanwer, Burhan Drak Sibai, Divya S Mansingka, Divya Shivaprasad, Ishaan Shah, Karl Stratos, Khoi Nguyen, Michael Callahan, Michael Pust, Mrinal Iyer, Philip Monk, Platon Mazarakis, Ritvik Kapila, Saurabh Srivastava, and Tim Romanski · 2025
Closest in time.
Spurious rewards: Rethinking training signals in rlvr
Rulin Shao, Shuyue Stella Li, Rui Xin, Scott Geng, Yiping Wang, Sewoong Oh, Simon Shaolei Du, Nathan Lambert, Sewon Min, Ranjay Krishna, Yulia Tsvetkov, Hannaneh Hajishirzi, Pang Wei Koh, and Luke Zettlemoyer · 2025
Closest in time.
RL’s razor: Why online reinforcement learning forgets less
Idan Shenfeld, Jyothish Pari, and Pulkit Agrawal · 2025
Closest in time.
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Na Zou, Hanjie Chen, and Xia Hu · 2025
Closest in time.
Understanding reasoning in thinking language models via steering vectors
Constantin Venhoff, Iván Arcuschin, Philip Torr, Arthur Conmy, and Neel Nanda · 2025
Closest in time.
Xumeng Wen, Zihan Liu, Shun Zheng, Zhijian Xu, Shengyu Ye, Zhirong Wu, Xiao Liang, Yang Wang, Junjie Li, Ziming Miao, Jiang Bian, and Mao Yang · 2025
Closest in time.
The invisible leash: Why rlvr may not escape its origin
Fang Wu, Weihao Xuan, Ximing Lu, Zaid Harchaoui, and Yejin Choi · 2025
Closest in time.
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, and Chenxu Lv et al · 2025
Closest in time.
LIMO: Less is more for reasoning
Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu · 2025
Closest in time.
Hiroshi Yoshihara, Taiki Yamaguchi, and Yuichi Inoue · 2025
Closest in time.
Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model?
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, and Gao Huang · 2025
Closest in time.
SimpleRL-Zoo: Investigating and taming zero reinforcement learning for open base models in the wild
Weihao Zeng, Yuzhen Huang, Qian Liu, Wei Liu, Keqing He, Zejun Ma, and Junxian He · 2025
Closest in time.
First return, entropy-eliciting explore
Tianyu Zheng, Tianshun Xing, Qingshui Gu, Taoran Liang, Xingwei Qu, Xin Zhou, Yizhi Li, Zhoufutu Wen, Chenghua Lin, Wenhao Huang, Qian Liu, Ge Zhang, and Zejun Ma · 2025
Closest in time.
Reinforcing general reasoning without verifiers
Xiangxin Zhou, Zichen Liu, Anya Sims, Haonan Wang, Tianyu Pang, Chongxuan Li, Liang Wang, Min Lin, and Chao Du · 2025
Closest in time.
Graphlet-based characterization of directed networks
Anida Sarajlić, Noël Malod-Dognin, Ömer Nebil Yaveroğlu, and Nataša Pržulj · 2045
Closest in time.