Fetching the paper…
Reading the bibliography…
The technical report introduces O1-CODER, an attempt to replicate OpenAI's o1 model with a focus on coding tasks.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Thinking fast and slow in ai: the role of metacognition, 2021
Marianna Bergamaschi Ganapini, Murray Campbell, Francesco Fabiano, Lior Horesh, Jon Lenchner, Andrea Loreggia, Nicholas Mattei, Francesca Rossi, Biplav Srivastava, and Kristen Brent Venable · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Grounding large language models in interactive environments with online reinforcement learning
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2023
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Earlier work this paper cites.
Alphazero-like tree-search can guide large language model decoding and training
Xidong Feng, Ziyu Wan, Muning Wen, Stephen Marcus McAleer, Ying Wen, Weinan Zhang, and Jun Wang · 2023
Earlier work this paper cites.
Taco: Topics in algorithmic code generation dataset
Rongao Li, Jie Fu, Bo-Wen Zhang, Tao Huang, Zhihong Sun, Chen Lyu, Guang Liu, Zhi Jin, and Ge Li · 2023
Earlier work this paper cites.
https://github.com/Open-Source-O1/Open-O1/ , 2024
Open o1: A model matching proprietary power with open-source innovation · 2024
Earlier work this paper cites.
g1: Using Llama-3.1 70b on Groq to create o1-like reasoning chains
Benjamin Klieger · 2024
Earlier work this paper cites.
Don’t teach. incentivize: Scale-first view of large language models, 2024
Hyung Won Chung · 2024
Earlier work this paper cites.
Language is primarily a tool for communication rather than thought
Evelina Fedorenko, Steven T. Piantadosi, and Edward A. Gibson · 2024
Earlier work this paper cites.
O1 replication journey: A strategic progress report, 2024
GAIR-NLP · 2024
Cited alongside, same era.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al · 2024
Cited alongside, same era.
Are large vision language models up to the challenge of chart comprehension and reasoning?
Mohammed Saidul Islam, Raian Rahman, Ahmed Masry, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, , and Enamul Hoque · 2024
Cited alongside, same era.
A small step towards reproducing openai o1: Progress report on the steiner open source models, October 2024
Yichao Ji · 2024
Cited alongside, same era.
Equitable access to justice: Logical llms show promise
Manuj Kant, Marzieh Nabi, Manav Kant, Preston Carlson, and Megan Ma · 2024
Thinking Claude
Richards Tu · 2024
Closest in time.
A note on generative games: Positioning, progress and prospects
Jitao Sang · 2024
Closest in time.
InternThinker
Shanghai AI Lab · 2024
Closest in time.
Llama-o1: Open large reasoning model frameworks for training, inference and evaluation with pytorch and huggingface
SimpleBerry · 2024
Closest in time.
Openr: An open source framework for advanced reasoning with large language models
OpenR Team · 2024
Closest in time.
Planning in strawberry fields: Evaluating and improving the planning and scheduling capabilities of lrm o1, 2024
Karthik Valmeekam, Kaya Stechly, Atharva Gundawar, and Subbarao Kambhampati · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Alr2: A retrieve-then-reason framework for long-context question answering
Huayang Li, Pat Verga, Priyanka Sen, Bowen Yang, Vijay Viswanathan, Patrick Lewis, Taro Watanabe, and Yixuan Su · 2024
Cited alongside, same era.
System 2 reasoning capabilities are nigh, 2024
Scott C. Lowe · 2024
Cited alongside, same era.
Generative reward models, 2024
Dakota Mahan, Duy Van Phung, Rafael Rafailov, Chase Blagden, Nathan Lile, Louis Castricato, Jan-Philipp Fränken, Chelsea Finn, and Alon Albalak · 2024
Cited alongside, same era.
Learning to reason with large language models
OpenAI · 2024
Cited alongside, same era.
Genie 2: A large-scale foundation world model
Jack Parker-Holder, Stephen Spencer, Philip Ball, Jake Bruce, Vibhavari Dasagi, Kristian Holsheimer, Christos Kaplanis, Alexandre Moufarek, Guy Scully, Jeremy Shar, Jimmy Shi, Jessica Yung, Michael Dennis, Sultan Kenjeyev, Shangbang Long, Yusuf Aytar, Jeff Clune, Sander Dieleman, Doug Eck, Shlomi Fruchter, Raia Hadsell, Demis Hassabis, Georg Ostrovski, Pieter-Jan Kindermans, Nicolas Heess, Charles Blundell, Simon Osindero, Rushil Mistry, et al · 2024
Cited alongside, same era.
Mutual reasoning makes smaller llms stronger problem-solvers
Zhenting Qi, Mingyuan Ma, Jiahang Xu, Li Lyna Zhang, Fan Yang, and Mao Yang · 2024
Cited alongside, same era.
QwQ-32b-preview
Qwen Team · 2024
Cited alongside, same era.
Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations
Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui · 2024
Closest in time.
Don’t command, cultivate: An exploratory study of system-2 alignment, 2024
Yuhang Wang and Jitao Sang · 2024
Closest in time.
Llava-o1: Let vision language models reason step-by-step, 2024
Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, and Li Yuan · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking
Eric Zelikman et al · 2024
Closest in time.
Llama-berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning
Di Zhang, Jianbo Wu, Jingdi Lei, Tong Che, Jiatong Li, Tong Xie, Xiaoshui Huang, Shufei Zhang, Marco Pavone, Yuqiang Li, Wanli Ouyang, and Dongzhan Zhou · 2024
Closest in time.
Marco-o1: Towards open reasoning models for open-ended solutions, 2024
Yu Zhao, Huifeng Yin, Bo Zeng, Hao Wang, Tianqi Shi, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang · 2024
Closest in time.