Fetching the paper…
Reading the bibliography…
We propose a novel framework, Meta Chain-of-Thought (Meta-CoT), which extends traditional Chain-of-Thought (CoT) by explicitly modeling the underlying reasoning required to arrive at a particular CoT.
Efficient off-policy meta-reinforcement learning via probabilistic context variables, 2019
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 1903
Earlier work this paper cites.
Meta reinforcement learning as task inference, 2019
Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A. Ortega, Yee Whye Teh, and Nicolas Heess · 1905
Earlier work this paper cites.
A formal theory of inductive inference. part i
Ray J Solomonoff · 1964
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Serial order: A parallel distributed processing approach
Michael I Jordan · 1997
Earlier work this paper cites.
Bandit based Monte-Carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
On effective parallelization of monte carlo tree search, 2020
Anji Liu, Yitao Liang, Ji Liu, Guy Van den Broeck, and Jianshu Chen · 2006
Earlier work this paper cites.
Multi-armed bandits with episode context
Christopher D. Rosin · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning, 2016
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
One-shot learning with memory-augmented neural networks, 2016
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. Maddison, et al · 2016
Earlier work this paper cites.
B-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Learning word vectors for 157 languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy P. Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Some considerations on learning to explore via meta-reinforcement learning, 2019
Bradly C. Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever · 2019
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Earlier work this paper cites.
The lean mathematical library
The mathlib Community · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Why generalization in rl is difficult: Epistemic pomdps and implicit partial observability, 2021
Dibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang, Ryan P. Adams, and Sergey Levine · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Scaling scaling laws with board games, 2021
Andy L. Jones · 2021
Earlier work this paper cites.
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen · 2022
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Learning to summarize from human feedback, 2022
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2022
Earlier work this paper cites.
Chain of thought imitation with procedure cloning, 2022
Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, and Ofir Nachum · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Earlier work this paper cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication, 2023
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S. Morcos · 2023
Earlier work this paper cites.
GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch, 2023
Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Jason Phang, Shivanshu Purohit, Hailey Schoelkopf, Dashiell Stander, Tri Songz, Curt Tigges, Benjamin Thérien, Phil Wang, and Samuel Weinbach · 2023
Cited alongside, same era.
Strategic reasoning with language models, 2023
Kanishk Gandhi, Dorsa Sadigh, and Noah D. Goodman · 2023
Cited alongside, same era.
The false promise of imitating proprietary llms, 2023
Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model, 2023
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Cited alongside, same era.
Beyond a*: Better planning with transformers via search dynamics bootstrapping, 2024
Lucas Lehnert, Sainbayar Sukhbaatar, DiJia Su, Qinqing Zheng, Paul Mcvay, Michael Rabbat, and Yuandong Tian · 2024
Later among the works it cites.
Numinamath
Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Costa Huang, Kashif Rasul, Longhui Yu, Albert Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu · 2024
Later among the works it cites.
Chain of thought empowers transformers to solve inherently serial problems, 2024
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Later among the works it cites.
Dakota Mahan, Duy Van Phung, Rafael Rafailov, Chase Blagden, Nathan Lile, Louis Castricato, Jan-Philipp Fränken, Chelsea Finn, and Alon Albalak · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback, 2023
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Cited alongside, same era.
The expresssive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Training chain-of-thought via latent-variable inference, 2023
Du Phan, Matthew D. Hoffman, David Dohan, Sholto Douglas, Tuan Anh Le, Aaron Parisi, Pavel Sountsov, Charles Sutton, Sharad Vikram, and Rif A. Saurous · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning, 2023
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Cited alongside, same era.
A long way to go: Investigating length correlations in rlhf, 2023
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Cited alongside, same era.
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar · 2024
Later among the works it cites.
Orca-math: Unlocking the potential of slms in grade school math, 2024
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah · 2024
Later among the works it cites.
Evolve: Evaluating and optimizing llms for exploration, 2024
Allen Nie, Yi Su, Bo Chang, Jonathan N. Lee, Ed H. Chi, Quoc V. Le, and Minmin Chen · 2024
Later among the works it cites.
Asynchronous rlhf: Faster and more efficient off-policy rl for language models, 2024
Michael Noukhovitch, Shengyi Huang, Sophie Xhonneux, Arian Hosseini, Rishabh Agarwal, and Aaron Courville · 2024
Later among the works it cites.
On the representational capacity of neural language models with chain-of-thought reasoning
Franz Nowak, Anej Svete, Alexandra Butoi, and Ryan Cotterell · 2024
Later among the works it cites.
Learning to reason with llms
OpenAI · 2024
Later among the works it cites.
Disentangling length from quality in direct preference optimization, 2024
Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
Why think step by step? reasoning emerges from the locality of experience
Ben Prystawski, Michael Li, and Noah Goodman · 2024
Later among the works it cites.
Agent q: Advanced reasoning and learning for autonomous ai agents
Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov · 2024
Later among the works it cites.
Recursive introspection: Teaching language model agents how to self-improve, 2024
Yuxiao Qu, Tianjun Zhang, Naman Garg, and Aviral Kumar · 2024
Later among the works it cites.
Iteration of thought: Leveraging inner dialogue for autonomous large language model reasoning, 2024
Santosh Kumar Radha, Yasamin Nouri Jelyani, Ara Ghukasyan, and Oktay Goktas · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Later among the works it cites.
Mastering board games by external and internal planning with language models, 2024
John Schultz, Jakub Adamek, Matej Jusup, Marc Lanctot, Michael Kaisers, Sarah Perrin, Daniel Hennes, Jeremy Shar, Cannada Lewis, Anian Ruoss, Tom Zahavy, Petar Veličković, Laurel Prince, Satinder Singh, Eric Malmi, and Nenad Tomašev · 2024
Later among the works it cites.
Algorithm of thoughts: Enhancing exploration of ideas in large language models, 2024
Bilgehan Sel, Ahmad Al-Tawaha, Vanshaj Khattar, Ruoxi Jia, and Ming Jin · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Later among the works it cites.
Beyond human data: Scaling self-training for problem-solving with language models, 2024
Avi Singh, John D. Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J. Liu, James Harrison, Jaehoon Lee, Kelvin Xu, Aaron Parisi, Abhishek Kumar, Alex Alemi, Alex Rizkowsky, Azade Nova, Ben Adlam, Bernd Bohnet, Gamaleldin Elsayed, Hanie Sedghi, Igor Mordatch, Isabelle Simpson, Izzeddin Gur, Jasper Snoek, Jeffrey Pennington, Jiri Hron, Kathleen Kenealy, Kevin Swersky, Kshiteej Mahajan, Laura Culp, Lechao Xiao, Maxwell L. Bileschi, Noah Constant, Roman Novak, Rosanne Liu, Tris Warkentin, Yundi Qian, Yamini Bansal, Ethan Dyer, Behnam Neyshabur, Jascha Sohl-Dickstein, and Noah Fiedel · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Fahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov, Jeff Schneider, Tengyang Xie, Stefano Ermon, Chelsea Finn, and Aviral Kumar · 2024
Later among the works it cites.
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Qwen Team · 2024
Later among the works it cites.
Toward self-improvement of llms via imagination, searching, and criticizing, 2024
Ye Tian, Baolin Peng, Linfeng Song, Lifeng Jin, Dian Yu, Haitao Mi, and Dong Yu · 2024
Later among the works it cites.
Openmathinstruct-2: Accelerating ai for math with massive open-source instruction data
Shubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin, Alexan Ayrapetyan, and Igor Gitman · 2024
Later among the works it cites.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui · 2024
Later among the works it cites.
Monte carlo tree search boosts reasoning via iterative preference learning, 2024
Yuxi Xie, Anirudh Goyal, Wenyue Zheng, Min-Yen Kan, Timothy P. Lillicrap, Kenji Kawaguchi, and Michael Shieh · 2024
Later among the works it cites.
Shuo Yin, Weihao You, Zhilong Ji, Guoqiang Zhong, and Jinfeng Bai · 2024
Later among the works it cites.
Exact: Teaching ai agents to explore with reflective-mcts and exploratory learning, 2024
Xiao Yu, Baolin Peng, Vineeth Vajipey, Hao Cheng, Michel Galley, Jianfeng Gao, and Zhou Yu · 2024
Later among the works it cites.
HARP: A challenging human-annotated math reasoning benchmark, 2024
Albert S. Yue, Lovish Madaan, Ted Moskovitz, DJ Strouse, and Aaditya K. Singh · 2024
Later among the works it cites.
Quiet-star: Language models can teach themselves to think before speaking
Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber, and Noah D Goodman · 2024
Later among the works it cites.
Optimizing llm test-time compute involves solving a meta-rl problem
Amrith Setlur, Yuxiao Qu, Lunjun Zhang, Matthew Yang, Virginia Smith, and Aviral Kumar · 2025
Closest in time.