Fetching the paper…
Reading the bibliography…
A key paradigm to improve the reasoning capabilities of large language models (LLMs) is to allocate more inference-time compute to search against a verifier or reward model.
The dondition of a finite Markov chain and perturbation bounds for the limiting probabilities
Carl Dean Meyer · 1980
Earlier work this paper cites.
Stochastic complementation, uncoupling Markov chains, and the theory of nearly reducible systems
Carl D. Meyer · 1989
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
On using extended statistical queries to avoid membership queries
Nader Bshouty and Vitaly Feldman · 2001
Earlier work this paper cites.
Markov chain decomposition for convergence rate analysis
Neal Madras and Dana Randall · 2001
Earlier work this paper cites.
Metastability and low lying spectra in reversible Markov chains
Anton Bovier, Michael Eckhoff, Véronique Gayrard, and Markus Klein · 2002
Earlier work this paper cites.
An algorithm for computing stochastically stable distributions with applications to mltiagent learning in repeated games
John Wicks and Amy Greenwald · 2005
Earlier work this paper cites.
An SVD approach to identifying metastable states of Markov chains
David Fritzsche, Volker Mehrmann, Daniel Szyld, and Elena Virnik · 2008
Earlier work this paper cites.
Markov Chains and Mixing Times
David Asher Levin, Yuval Peres, and Elizabeth Lee Wilmer · 2009
Earlier work this paper cites.
A robust spectral method for finding lumpings and meta-stable states of non-reversible Markov chains
Martin Nilsson Jacobi · 2010
Earlier work this paper cites.
Metastability of reversible finite state Markov processes
J. Beltrán and C. Landim · 2011
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
On an SVD-based algorithm for identifying meta-stable states of Markov chains
Ryan Tifenbach · 2011
Earlier work this paper cites.
Metastability for a non-reversible dynamics: The evolution of the condensate in totally asymmetric zero range processes
C. Landim · 2012
Earlier work this paper cites.
Metastability for general dynamics with rare transitions: escape time and critical configurations
Emilio Cirillo, Francesca Nardi, and Julien Sohier · 2014
Earlier work this paper cites.
Asymptotically exponential hitting times and metastability: A pathwise approach without reversibility
Roberto Fernandez, Francesco Manzo, Francesca Nardi, and Elisabetta Scoppola · 2014
Earlier work this paper cites.
Metastability of finite state Markov chains: A recursive procedure to identify slow variables for model reduction
Claudio Landim and Tiecheng Xu · 2015
Earlier work this paper cites.
Concentration inequalities for Markov chains by Marton couplings and spectral methods
Daniel Paulin · 2015
Earlier work this paper cites.
Multi-scale metastable dynamics and the asymptotic stationary distribution of perturbed Markov chains
Volker Betz and Stéphane Le Roux · 2016
Earlier work this paper cites.
Metastable states, quasi-stationary distributions and soft measures
Alessandra Bianchi and Alexandre Gaudillière · 2016
Earlier work this paper cites.
Conditioned, quasi-stationary, restricted measures and escape from metastable states
R. Fernandez, F. Manzo, F. R. Nardi, E. Scoppola, and J. Sohier · 2016
Earlier work this paper cites.
A general characterization of the statistical query complexity
Vitaly Feldman · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Earlier work this paper cites.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A. Efros · 2018
Earlier work this paper cites.
Spectral clustering for non-reversible Markov chains
Konstantin Fackeldey, Alexander Sikorski, and M. Weber · 2018
Cited alongside, same era.
C. Landim · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Distribution-specific hardness of learning neural networks
Ohad Shamir · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Exploration by random network distillation
Graph of thoughts: solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al · 2024
Later among the works it cites.
Understanding in-context learning in transformers and LLMs by learning to learn discrete functions
Satwik Bhattamishra, Arkil Patel, Phil Blunsom, and Varun Kanade · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
The evolution of statistical induction heads: in-context learning Markov chains
Ezra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, eran malach, and Surbhi Goel · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Cited alongside, same era.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
What can neural networks reason about?
Keyulu Xu, Jingling Li, Mozhi Zhang, Simon S Du, Ken-ichi Kawarabayashi, and Stefanie Jegelka · 2019
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Scaling scaling laws with board games
Andy L Jones · 2021
Cited alongside, same era.
Show your work: scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Cited alongside, same era.
Constitutional AI: harmlessness from AI feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Kanishk Gandhi, Denise Lee, Gabriel Grand, Muxin Liu, Winson Cheng, Archit Sharma, and Noah D Goodman · 2024
Later among the works it cites.
Unveiling the statistical foundations of chain-of-thought prompting methods
Xinyang Hu, Fengzhuo Zhang, Siyu Chen, and Zhuoran Yang · 2024
Later among the works it cites.
From self-attention to Markov models: unveiling the dynamics of generative transformers
Muhammed Emrullah Ildiz, Yixiao Huang, Yingcong Li, Ankit Singh Rawat, and Samet Oymak · 2024
Later among the works it cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Later among the works it cites.
Transformers provably solve parity efficiently with chain of thought
Juno Kim and Taiji Suzuki · 2024
Later among the works it cites.
Training language models to self-correct via reinforcement learning
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al · 2024
Later among the works it cites.
How do nonlinear transformers acquire generalization-guaranteed CoT ability?
Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui, and Pin-Yu Chen · 2024
Later among the works it cites.
Attention with Markov: A framework for principled analysis of transformers via Markov chains
Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Martin Jaggi, Hyeji Kim, and Michael Gastpar · 2024
Later among the works it cites.
How transformers learn causal structure with gradient descent
Eshaan Nichani, Alex Damian, and Jason D. Lee · 2024
Later among the works it cites.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Solving olympiad geometry without human demonstrations
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong · 2024
Later among the works it cites.
Kaiyue Wen, Huaqing Zhang, Hongzhou Lin, and Jingzhao Zhang · 2024
Later among the works it cites.
Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang · 2024
Later among the works it cites.
Monte Carlo tree search boosts reasoning via iterative preference learning
Yuxi Xie, Anirudh Goyal, Wenyue Zheng, Min-Yen Kan, Timothy P. Lillicrap, Kenji Kawaguchi, and Michael Shieh · 2024
Later among the works it cites.
Tree of thoughts: deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Large language models as Markov chains
Oussama Zekri, Ambroise Odonnat, Abdelhakim Benechehab, Linus Bleistein, Nicolas Boullé, and Ievgen Redko · 2024
Later among the works it cites.
DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Kimi k1.5: scaling reinforcement learning with LLMs
Team Kimi, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Spinning up: proximal policy optimization (PPO), 2018
OpenAI · 2025
Closest in time.
Towards System 2 reasoning in LLMs: learning how to think with meta chain-of-thought
Violet Xiang, Charlie Snell, Kanishk Gandhi, Alon Albalak, Anikait Singh, Chase Blagden, Duy Phung, Rafael Rafailov, Nathan Lile, Dakota Mahan, et al · 2025
Closest in time.