Fetching the paper…
Reading the bibliography…
We introduce Olmo 3, a family of state-of-the-art, fully-open language models at the 7B and 32B parameter scales.
Hierarchical grouping to optimize an objective function
J. H. Ward Jr · 1963
Earlier work this paper cites.
Causal reasoning and backtracking
J. M. Joyce · 2009
Earlier work this paper cites.
Wikiextractor
G. Attardi · 2015
Earlier work this paper cites.
Metacognition and abstract reasoning
H. Markovits, V. A. Thompson, and J. Brisson · 2015
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Meta-reasoning: Monitoring and control of thinking and reasoning
R. Ackerman and V. A. Thompson · 2017
Earlier work this paper cites.
Self-evaluation of decision-making: A general bayesian framework for metacognitive computation
S. Fleming and N. Daw · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
J. Welbl, N. F. Liu, and M. Gardner · 2017
Earlier work this paper cites.
Elastic ChatNoir: Search Engine for the ClueWeb and the Common Crawl
J. Bevendorff, B. Stein, M. Hagen, and M. Potthast · 2018
Earlier work this paper cites.
Think you have solved question answering? Try ARC, the AI2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Hierarchical thinking: a cognitive tool for guiding coherent decision making in design problem solving
G. Haupt · 2018
Earlier work this paper cites.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. Van Hasselt, and D. Silver · 2018
Earlier work this paper cites.
Deep reinforcement learning and the deadly triad
H. Van Hasselt, Y. Doron, F. Strub, M. Hessel, N. Sonnerat, and J. Modayil · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Eli5: Long form question answering
A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli · 2019
Earlier work this paper cites.
Doing more with less: meta-reasoning and meta-learning in humans and machines
T. L. Griffiths, F. Callaway, M. B. Chang, E. Grant, P. M. Krueger, and F. Lieder · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Earlier work this paper cites.
CoQA: A conversational question answering challenge
S. Reddy, D. Chen, and C. D. Manning · 2019
Earlier work this paper cites.
Social IQa: Commonsense reasoning about social interactions
M. Sap, H. Rashkin, D. Chen, R. Le Bras, and Y. Choi · 2019
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
D. Saxton, E. Grefenstette, F. Hill, and P. Kohli · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Earlier work this paper cites.
PIQA: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. Le bras, J. Gao, and Y. Choi · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
WinoGrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. Le Bras, C. Bhagavatula, and Y. Choi · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
P. Dasigi, K. Lo, I. Beltagy, A. Cohan, N. A. Smith, and M. Gardner · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Earlier work this paper cites.
Efficient training of language models to fill in the middle
M. Bavarian, H. Jun, N. Tezak, J. Schulman, C. McLeavey, J. Tworek, and M. Chen · 2022
Earlier work this paper cites.
Rational use of cognitive resources in human planning
F. Callaway, B. {van Opheusden}, S. Gul, P. Das, P. Krueger, T. Griffiths, and F. Lieder · 2022
Earlier work this paper cites.
Multipl-e: A scalable and extensible approach to benchmarking neural code generation
F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, et al · 2022
Earlier work this paper cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
T. Hartvigsen, S. Gabriel, H. Palangi, M. Sap, D. Ray, and E. Kamar · 2022
Earlier work this paper cites.
Less is more: Parameter-free text classification with gzip, 2022
Z. Jiang, M. Y. R. Yang, M. Tsirlin, R. Tang, and J. Lin · 2022
Earlier work this paper cites.
Ds-1000: A natural and reliable benchmark for data science code generation
Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, W.-T. Yih, D. Fried, S. Wang, and T. Yu · 2022
Earlier work this paper cites.
Deduplicating training data makes language models better, 2022
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al · 2022
Earlier work this paper cites.
Data contamination: From memorization to exploitation
I. Magar and R. Schwartz · 2022
Earlier work this paper cites.
When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories
A. Mallen, A. Asai, V. Zhong, R. Das, H. Hajishirzi, and D. Khashabi · 2022
Earlier work this paper cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
S. Mindermann, J. M. Brauner, M. T. Razzak, M. Sharma, A. Kirsch, W. Xu, B. Höltgen, A. N. Gomez, A. Morisot, S. Farquhar, et al · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
A. Pal, L. K. Umapathi, and M. Sankarasubbu · 2022
Earlier work this paper cites.
BBQ: A hand-built bias benchmark for question answering
A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. Bowman · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Earlier work this paper cites.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al · 2022
Earlier work this paper cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, et al · 2022
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai · 2023
Earlier work this paper cites.
A general theoretical paradigm to understand learning from human preferences, 2023
M. G. Azar, M. Rowland, B. Piot, D. Guo, D. Calandriello, M. Valko, and R. Munos · 2023
Earlier work this paper cites.
Llemma: An open language model for mathematics, 2023
Z. Azerbayev, H. Schoelkopf, K. Paster, M. D. Santos, S. McAleer, A. Q. Jiang, J. Deng, S. Biderman, and S. Welleck · 2023
Earlier work this paper cites.
Extending context window of large language models via positional interpolation, 2023
S. Chen, S. Wong, L. Chen, and Y. Tian · 2023
Earlier work this paper cites.
UltraFeedback: Boosting language models with scaled ai feedback
G. Cui, L. Yuan, N. Ding, G. Yao, W. Zhu, Y. Ni, G. Xie, Z. Liu, and M. Sun · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
N. Ding, Y. Chen, B. Xu, Y. Qin, S. Hu, Z. Liu, M. Sun, and B. Zhou · 2023
Earlier work this paper cites.
Aligning large language models through synthetic feedback
S. Kim, S. Bae, J. Shin, S. Kang, D. Kwak, K. Yoo, and M. Seo · 2023
Earlier work this paper cites.
Madlad-400: A multilingual and document-level large audited dataset, 2023
S. Kudugunta, I. Caswell, B. Zhang, X. Garcia, C. A. Choquette-Choo, K. Lee, D. Xin, A. Kusupati, R. Stella, A. Bapna, and O. Firat · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Earlier work this paper cites.
Entangled preferences: The history and risks of reinforcement learning and human feedback
N. Lambert, T. K. Gilbert, and T. Zick · 2023
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Earlier work this paper cites.
The flan collection: Designing data and methods for effective instruction tuning
S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, et al · 2023
Earlier work this paper cites.
Wizardcoder: Empowering code large language models with evol-instruct, 2023
Z. Luo, C. Xu, P. Zhao, Q. Sun, X. Geng, W. Hu, C. Tao, J. Ma, Q. Lin, and D. Jiang · 2023
Earlier work this paper cites.
Openwebmath: An open dataset of high-quality mathematical web text, 2023
K. Paster, M. D. Santos, Z. Azerbayev, and J. Ba · 2023
Earlier work this paper cites.
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay · 2023
Earlier work this paper cites.
Yarn: Efficient context window extension of large language models, 2023
B. Peng, J. Quesnelle, H. Fan, and E. Shippole · 2023
Earlier work this paper cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
P. Röttger, H. R. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy · 2023
Cited alongside, same era.
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo · 2023
Cited alongside, same era.
peS2o (Pretraining Efficiently on S2ORC) Dataset, 2023
L. Soldaini and K. Lo · 2023
Cited alongside, same era.
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
Teknium · 2023
Cited alongside, same era.
RedPajama: An open source recipe to reproduce LLaMA training dataset, 2023
Together AI · 2023
Cited alongside, same era.
Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models, 2025
C. An, Z. Xie, X. Li, L. Li, J. Zhang, S. Gong, M. Zhong, J. Xu, X. Qiu, M. Wang, and L. Kong · 2025
Closest in time.
System card: Claude opus 4 & claude sonnet 4
Anthropic · 2025
Closest in time.
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Apertus Team · 2025
Closest in time.
LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks
Y. Bai, S. Tu, J. Zhang, H. Peng, X. Wang, X. Lv, S. Cao, J. Xu, L. Hou, Y. Dong, J. Tang, and J. Li · 2025
Closest in time.
SmolLM3: smol, multilingual, long-context reasoner
E. Bakouch, L. Ben Allal, A. Lozhkov, N. Tazi, L. Tunstall, C. M. Patiño, E. Beeching, A. Roucher, A. J. Reedi, Q. Gallouédec, K. Rasul, N. Habib, C. Fourrier, H. Kydlicek, G. Penedo, H. Larcher, M. Morlon, V. Srivastav, J. Lochner, X.-S. Nguyen, C. Raffel, L. von Werra, and T. Wolf · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Xiong, J. Liu, I. Molybog, H. Zhang, P. Bhargava, R. Hou, L. Martin, R. Rungta, K. A. Sankararaman, B. Oguz, M. Khabsa, H. Fang, Y. Mehdad, S. Narang, K. Malik, A. Fan, S. Bhosale, S. Edunov, M. Lewis, S. Wang, and H. Ma · 2023
Cited alongside, same era.
Tablegpt: Towards unifying tables, natural language and commands into one gpt
L. Zha, J. Zhou, L. Li, R. Wang, Q. Huang, S. Yang, J. Yuan, C. Su, X. Li, A. Su, T. Zhang, C. Zhou, K. Shou, M. Wang, W. Zhu, G. Lu, C. Ye, Y. Ye, W. Ye, Y. Zhang, X. Deng, J. Xu, H. Wang, G. Chen, and J. Zhao · 2023
Cited alongside, same era.
Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023
Y. Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Cited alongside, same era.
Agieval: A human-centric benchmark for evaluating foundation models
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Cited alongside, same era.
Instruction-following evaluation for large language models, 2023
J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou · 2023
Cited alongside, same era.
M. Abdin, J. Aneja, H. S. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kauffmann, J. R. Lee, Y. T. Lee, Y. Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y. Wu, D. Yu, C. Zhang, and Y. Zhang · 2024
Cited alongside, same era.
A. Bercovich, I. Levy, I. Golan, M. Dabbah, R. El-Yaniv, O. Puny, I. Galil, Z. Moshe, T. Ronen, N. Nabwani, I. Shahaf, O. Tropp, E. Karpas, R. Zilberstein, J. Zeng, S. Singhal, A. Bukharin, Y. Zhang, T. Konuk, G. Shen, A. S. Mahabaleshwarkar, B. Kartal, Y. Suhara, O. Delalleau, Z. Chen, Z. Wang, D. Mosallanezhad, A. Renduchintala, H. Qian, D. Rekesh, F. Jia, S. Majumdar, V. Noroozi, W. U. Ahmad, S. Narenthiran, A. Ficek, M. Samadi, J. Huang, S. Jain, I. Gitman, I. Moshkov, W. Du, S. Toshniwal, G. Armstrong, B. Kisacanin, M. Novikov, D. Gitman, E. Bakhturina, J. P. Scowcroft, J. Kamalu, D. Su, K. Kong, M. Kliegl, R. Karimi, Y. Lin, S. Satheesh, J. Parmar, P. Gundecha, B. Norick, J. Jennings, S. Prabhumoye, S. N. Akter, M. Patwary, A. Khattar, D. Narayanan, R. Waleffe, J. Zhang, B.-Y. Su, G. Huang, T. Kong, P. Chadha, S. Jain, C. Harvey, E. Segal, J. Huang, S. Kashirsky, R. McQueen, I. Putterman, G. Lam, A. Venkatesan, S. Wu, V. Nguyen, M. Kilaru, A. Wang, A. Warno, A. Somasamudramath, S. Bhaskar, M. Dong, N. Assaf, S. Mor, O. U. Argov, S. Junkin, O. Romanenko, P. Larroy, M. Katariya, M. Rovinelli, V. Balas, N. Edelman, A. Bhiwandiwalla, M. Subramaniam, S. Ithape, K. Ramamoorthy, Y. Wu, S. V. Velury, O. Almog, J. Daw, D. Fridman, E. Galinkin, M. Evans, K. Luna, L. Derczynski, N. Pope, E. Long, S. Schneider, G. Siman, T. Grzegorzek, P. Ribalta, M. Katariya, J. Conway, T. Saar, A. Guan, K. Pawelec, S. Prayaga, O. Kuchaiev, B. Ginsburg, O. Olabiyi, K. Briski, J. Cohen, B. Catanzaro, J. Alben, Y. Geifman, E. Chung, and C. Alexiuk · 2025
Closest in time.
Astabench: Rigorous benchmarking of ai agents with a scientific research suite
J. Bragg, M. D’Arcy, N. Balepur, D. Bareket, B. Dalvi, S. Feldman, D. Haddad, J. D. Hwang, P. Jansen, V. Kishore, et al · 2025
Closest in time.
Aegisllm: Scaling agentic systems for self-reflective defense in llm security
Z. Cai, S. Shabihi, B. An, Z. Che, B. R. Bartoldson, B. Kailkhura, T. Goldstein, and F. Huang · 2025
Closest in time.
How people use ChatGPT
A. Chatterji, T. Cunningham, D. Deming, Z. Hitzig, C. Ong, C. Y. Shan, and K. Wadman · 2025
Closest in time.
Acereason-nemotron: Advancing math and code reasoning through reinforcement learning
Y. Chen, Z. Yang, Z. Liu, C. Lee, P. Xu, M. Shoeybi, B. Catanzaro, and W. Ping · 2025
Closest in time.
Revisiting reinforcement learning for llm reasoning from a cross-domain perspective, 2025
Z. Cheng, S. Hao, T. Liu, F. Zhou, Y. Xie, F. Yao, Y. Bian, Y. Zhuang, N. Dey, Y. Zha, Y. Gu, K. Zhou, Y. Wang, Y. Li, R. Fan, J. She, C. Gao, A. Saparov, H. Li, T. W. Killian, M. Yurochkin, Z. Liu, E. P. Xing, and Z. Hu · 2025
Closest in time.
Scaling llama 3 training with efficient parallelism strategies
W. Chu, X. Xie, J. Yu, J. Wang, A. Phanishayee, C. Tang, Y. Hao, J. Huang, M. Ozdal, J. Wang, V. Goswami, N. Goyal, A. Kadian, A. Gu, C. Cai, F. Tian, X. Wang, M. Si, P. Balaji, C.-H. Chu, and J. Park · 2025
Closest in time.
DeepSeek-V3.1 release
DeepSeek-AI · 2025
Closest in time.
Deepseek-v3 technical report, 2025
DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, J. Qiu, J. Li, J. Song, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Xu, L. Xia, L. Zhao, L. Wang, L. Zhang, M. Li, M. Wang, M. Zhang, M. Zhang, M. Tang, M. Li, N. Tian, P. Huang, P. Wang, P. Zhang, Q. Wang, Q. Zhu, Q. Chen, Q. Du, R. J. Chen, R. L. Jin, R. Ge, R. Zhang, R. Pan, R. Wang, R. Xu, R. Zhang, R. Chen, S. S. Li, S. Lu, S. Zhou, S. Chen, S. Wu, S. Ye, S. Ye, S. Ma, S. Wang, S. Zhou, S. Yu, S. Zhou, S. Pan, T. Wang, T. Yun, T. Pei, T. Sun, W. L. Xiao, W. Zeng, W. Zhao, W. An, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, X. Q. Li, X. Jin, X. Wang, X. Bi, X. Liu, X. Wang, X. Shen, X. Chen, X. Zhang, X. Chen, X. Nie, X. Sun, X. Wang, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Song, X. Shan, X. Zhou, X. Yang, X. Li, X. Su, X. Lin, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. X. Zhu, Y. Zhang, Y. Xu, Y. Xu, Y. Huang, Y. Li, Y. Zhao, Y. Sun, Y. Li, Y. Wang, Y. Yu, Y. Zheng, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Tang, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Wu, Y. Ou, Y. Zhu, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Zha, Y. Xiong, Y. Ma, Y. Yan, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Huang, Z. Zhang, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Xu, Z. Wu, Z. Zhang, Z. Li, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Gao, and Z. Pan · 2025
Closest in time.
Climb: Clustering-based iterative data mixture bootstrapping for language model pre-training, 2025
S. Diao, Y. Yang, Y. Fu, X. Dong, D. Su, M. Kliegl, Z. Chen, P. Belcak, Y. Suhara, H. Yin, M. Patwary, Yingyan, Lin, J. Kautz, and P. Molchanov · 2025
Closest in time.
Anchored preference optimization and contrastive revisions: Addressing underspecification in alignment
K. D’Oosterlinck, W. Xu, C. Develder, T. Demeester, A. Singh, C. Potts, D. Kiela, and S. Mehri · 2025
Closest in time.
Rewriting pre-training data boosts llm performance in math and code
K. Fujii, Y. Tajima, S. Mizuki, H. Shimada, T. Shiotani, K. Saito, M. Ohi, M. Kawamura, T. Nakamura, T. Okamoto, et al · 2025
Closest in time.
Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars
K. Gandhi, A. Chakravarthy, A. Singh, N. Lile, and N. D. Goodman · 2025
Closest in time.
How to train long-context language models (effectively)
T. Gao, A. Wettig, H. Yen, and D. Chen · 2025
Closest in time.
Gemma 3 technical report, 2025
Gemma 3 Team · 2025
Closest in time.
The delta learning hypothesis: Preference tuning on weak data can yield strong gains
S. Geng, H. Ivison, C.-L. Li, M. Sap, J. Li, R. Krishna, and P. W. Koh · 2025
Closest in time.
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models, 2025
GLM-4.5 Team, A. Zeng, X. Lv, Q. Zheng, Z. Hou, B. Chen, C. Xie, C. Wang, D. Yin, H. Zeng, J. Zhang, K. Wang, L. Zhong, M. Liu, R. Lu, S. Cao, X. Zhang, X. Huang, Y. Wei, Y. Cheng, Y. An, Y. Niu, Y. Wen, Y. Bai, Z. Du, Z. Wang, Z. Zhu, B. Zhang, B. Wen, B. Wu, B. Xu, C. Huang, C. Zhao, C. Cai, C. Yu, C. Li, C. Ge, C. Huang, C. Zhang, C. Xu, C. Zhu, C. Li, C. Yin, D. Lin, D. Yang, D. Jiang, D. Ai, E. Zhu, F. Wang, G. Pan, G. Wang, H. Sun, H. Li, H. Li, H. Hu, H. Zhang, H. Peng, H. Tai, H. Zhang, H. Wang, H. Yang, H. Liu, H. Zhao, H. Liu, H. Yan, H. Liu, H. Chen, J. Li, J. Zhao, J. Ren, J. Jiao, J. Zhao, J. Yan, J. Wang, J. Gui, J. Zhao, J. Liu, J. Li, J. Li, J. Lu, J. Wang, J. Yuan, J. Li, J. Du, J. Du, J. Liu, J. Zhi, J. Gao, K. Wang, L. Yang, L. Xu, L. Fan, L. Wu, L. Ding, L. Wang, M. Zhang, M. Li, M. Xu, M. Zhao, M. Zhai, P. Du, Q. Dong, S. Lei, S. Tu, S. Yang, S. Lu, S. Li, S. Li, Shuang-Li, S. Yang, S. Yi, T. Yu, W. Tian, W. Wang, W. Yu, W. L. Tam, W. Liang, W. Liu, X. Wang, X. Jia, X. Gu, X. Ling, X. Wang, X. Fan, X. Pan, X. Zhang, X. Zhang, X. Fu, X. Zhang, Y. Xu, Y. Wu, Y. Lu, Y. Wang, Y. Zhou, Y. Pan, Y. Zhang, Y. Wang, Y. Li, Y. Su, Y. Geng, Y. Zhu, Y. Yang, Y. Li, Y. Wu, Y. Li, Y. Liu, Y. Wang, Y. Li, Y. Zhang, Z. Liu, Z. Yang, Z. Zhou, Z. Qiao, Z. Feng, Z. Liu, Z. Zhang, Z. Wang, Z. Yao, Z. Wang, Z. Liu, Z. Chai, Z. Li, Z. Zhao, W. Chen, J. Zhai, B. Xu, M. Huang, H. Wang, J. Li, Y. Dong, and J. Tang · 2025
Closest in time.
Extending AFM-4.5B to 64K context length
C. Goddard · 2025
Closest in time.
Gaperon: A peppered english-french generative language model suite, 2025
N. Godey, W. Antoun, R. Touchent, R. Bawden, Éric de la Clergerie, B. Sagot, and D. Seddah · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
marin-community/marin
D. Hall, C. Chou, A. Garg, N. Ravi, N. Liu, H. Shandilya, A. Ahmed, P. Liang, R. Kuditipudi, J38, T. Lee, R. Power, K. Salahi, W. Held, J. Wang, chiheem, J. Niklaus, Y. Mai, dependabot[bot], I. Zhou, K. X. Li, S. Yang, S. Karamcheti, R. Williams, C. Zhou, A. Ramaswami, whenwen, S. Kotha, G. Miguel, and C. Xu · 2025
Closest in time.
Signal and noise: A framework for reducing uncertainty in language model evaluation, 2025
D. Heineman, V. Hofmann, I. Magnusson, Y. Gu, N. A. Smith, H. Hajishirzi, K. Lo, and J. Dodge · 2025
Closest in time.
Liger kernel: Efficient triton kernels for llm training, 2025
P.-L. Hsu, Y. Dai, V. Kothapalli, Q. Song, S. Tang, S. Zhu, S. Shimizu, S. Sahni, H. Ning, and Y. Chen · 2025
Closest in time.
Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model
J. Hu, Y. Zhang, Q. Han, D. Jiang, X. Zhang, and H.-Y. Shum · 2025
Closest in time.
Cognitive foundations for reasoning and their manifestation in llms
P. Kargupta, S. S. Li, H. Wang, J. Lee, S. Chen, O. Ahia, D. Light, T. L. Griffiths, M. Kleiman-Weiner, J. Han, A. Celikyilmaz, and Y. Tsvetkov · 2025
Closest in time.
Gemini 2.5: Our most intelligent ai model
K. Kavukcuoğlu and G. DeepMind · 2025
Closest in time.
A systematic examination of preference learning through the lens of instruction-following
J. Kim, A. Goyal, A. Zhang, B. Xiong, R. Hou, M. Kambadur, D. Mahajan, H. Hajishirzi, and L. Tan · 2025
Closest in time.
Kimi k2: Open agentic intelligence, 2025
Kimi Team, Y. Bai, Y. Bao, G. Chen, J. Chen, N. Chen, R. Chen, Y. Chen, Y. Chen, Y. Chen, Z. Chen, J. Cui, H. Ding, M. Dong, A. Du, C. Du, D. Du, Y. Du, Y. Fan, Y. Feng, K. Fu, B. Gao, H. Gao, P. Gao, T. Gao, X. Gu, L. Guan, H. Guo, J. Guo, H. Hu, X. Hao, T. He, W. He, W. He, C. Hong, Y. Hu, Z. Hu, W. Huang, Z. Huang, Z. Huang, T. Jiang, Z. Jiang, X. Jin, Y. Kang, G. Lai, C. Li, F. Li, H. Li, M. Li, W. Li, Y. Li, Y. Li, Z. Li, Z. Li, H. Lin, X. Lin, Z. Lin, C. Liu, C. Liu, H. Liu, J. Liu, J. Liu, L. Liu, S. Liu, T. Y. Liu, T. Liu, W. Liu, Y. Liu, Y. Liu, Y. Liu, Y. Liu, Z. Liu, E. Lu, L. Lu, S. Ma, X. Ma, Y. Ma, S. Mao, J. Mei, X. Men, Y. Miao, S. Pan, Y. Peng, R. Qin, B. Qu, Z. Shang, L. Shi, S. Shi, F. Song, J. Su, Z. Su, X. Sun, F. Sung, H. Tang, J. Tao, Q. Teng, C. Wang, D. Wang, F. Wang, H. Wang, J. Wang, J. Wang, J. Wang, S. Wang, S. Wang, Y. Wang, Y. Wang, Y. Wang, Y. Wang, Y. Wang, Z. Wang, Z. Wang, Z. Wang, C. Wei, Q. Wei, W. Wu, X. Wu, Y. Wu, C. Xiao, X. Xie, W. Xiong, B. Xu, J. Xu, J. Xu, L. H. Xu, L. Xu, S. Xu, W. Xu, X. Xu, Y. Xu, Z. Xu, J. Yan, Y. Yan, X. Yang, Y. Yang, Z. Yang, Z. Yang, Z. Yang, H. Yao, X. Yao, W. Ye, Z. Ye, B. Yin, L. Yu, E. Yuan, H. Yuan, M. Yuan, H. Zhan, D. Zhang, H. Zhang, W. Zhang, X. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, Z. Zhang, H. Zhao, Y. Zhao, H. Zheng, S. Zheng, J. Zhou, X. Zhou, Z. Zhou, Z. Zhu, W. Zhuang, and X. Zu · 2025
Closest in time.
Reinforcement Learning from Human Feedback
N. Lambert · 2025
Closest in time.
Model merging in pre-training of large language models
Y. Li, Y. Ma, S. Yan, C. Zhang, J. Liu, J. Lu, Z. Xu, M. Chen, M. Wang, S. Zhan, J. Ma, X. Lai, D. Liu, Y. Luo, X. Bin, H. Ren, M. Han, W. Hao, B. Yi, L. Liu, B. Ma, X. Jia, X. Zhou, S. Qiao, L. Xiang, and Y. Wu · 2025
Closest in time.
Zebralogic: On the scaling limits of llms for logical reasoning
B. Y. Lin, R. L. Bras, K. Richardson, A. Sabharwal, R. Poovendran, P. Clark, and Y. Choi · 2025
Closest in time.
Datadecide: How to predict best pretraining data with small experiments, 2025
I. Magnusson, N. Tai, B. Bogin, D. Heineman, J. D. Hwang, L. Soldaini, A. Bhagia, J. Liu, D. Groeneveld, O. Tafjord, N. A. Smith, P. W. Koh, and J. Dodge · 2025
Closest in time.
Deepseek-r1 thoughtology: Let’s think about llm reasoning, 2025
S. V. Marjanović, A. Patel, V. Adlakha, M. Aghajohari, P. BehnamGhader, M. Bhatia, A. Khandelwal, A. Kraft, B. Krojer, X. H. Lù, N. Meade, D. Shin, A. Kazemnejad, G. Kamath, M. Mosbach, K. Stańczak, and S. Reddy · 2025
Closest in time.
Nemotron-Personas-USA: Synthetic personas aligned to real-world distributions, June 2025
Y. Meyer and D. Corneil · 2025
Closest in time.
Search arena: Analyzing search-augmented llms
M. Miroyan, T.-H. Wu, L. King, T. Li, J. Pan, X. Hu, W.-L. Chiang, A. N. Angelopoulos, T. Darrell, N. Norouzi, and J. Gonzalez · 2025
Closest in time.
I. Moshkov, D. Hanley, I. Sorokin, S. Toshniwal, C. Henkel, B. Schifferer, W. Du, and I. Gitman · 2025
Closest in time.
Nemotron-Post-Training-Dataset-v1, 2025
D. Nathawani, I. Gitman, S. Majumdar, E. Bakhturina, A. Sunil Mahabaleshwarkar, , J. Zhang, and J. Polak Scowcroft · 2025
Closest in time.
NVIDIA Nemotron Nano 2: An accurate and efficient hybrid mamba-transformer reasoning model, 2025
NVIDIA, , A. Basant, A. Khairnar, A. Paithankar, A. Khattar, A. Renduchintala, A. Malte, A. Bercovich, A. Hazare, A. Rico, A. Ficek, A. Kondratenko, A. Shaposhnikov, A. Bukharin, A. Taghibakhshi, A. Barton, A. S. Mahabaleshwarkar, A. Shen, A. Tao, A. Guan, A. Shors, A. Mandarwal, A. Mehta, A. Venkatesan, A. Sharabiani, A. Aithal, A. Poojary, A. Dattagupta, B. Buddharaju, B. Zhu, B. Simkin, B. Kartal, B. D. Rouhani, B. Chen, B. Ginsburg, B. Norick, B. Yu, B. Catanzaro, C. Wang, C. Truong, C. Mungekar, C. Patel, C. Alexiuk, C. Munley, C. Parisien, D. Su, D. Afrimi, D. Korzekwa, D. Rohrer, D. Gitman, D. Mosallanezhad, D. Narayanan, D. Rekesh, D. Yared, D. Pykhtar, D. Ahn, D. Riach, E. Long, E. Ning, E. Chung, E. Galinkin, E. Bakhturina, G. Prasad, G. Shen, H. Qian, H. Elisha, H. Sharma, H. Ross, H. Ngo, H. Sahota, H. Wang, H. C. Shin, H. Huang, I. Cunningham, I. Gitman, I. Moshkov, J. Jung, J. Kautz, J. P. Scowcroft, J. Casper, J. Zhang, J. Zeng, J. Zhang, J. Xue, J. Huang, J. Conway, J. Kamalu, J. Cohen, J. Jennings, J. V. Vialard, J. Yi, J. Parmar, K. Briski, K. Cheung, K. Luna, K. Wyss, K. Santhanam, K. Kong, K. Pawelec, K. Anik, K. Li, K. Ahmadian, L. McAfee, L. Sleiman, L. Derczynski, L. Vega, M. R. de Melo, M. N. Sreedhar, M. Chochowski, M. Cai, M. Kliegl, M. Stepniewska-Dziubinska, M. Novikov, M. Samadi, M. Price, M. Boubdir, M. Boone, M. Evans, M. Bien, M. Zawalski, M. Martinez, M. Chrzanowski, M. Shoeybi, M. Patwary, N. Dhameja, N. Assaf, N. Habibi, N. Bhatia, N. Pope, N. Tajbakhsh, N. K. Juluru, O. Rybakov, O. Hrinchuk, O. Kuchaiev, O. Olabiyi, P. Ribalta, P. Subramanian, P. Chadha, P. Molchanov, P. Dykas, P. Jin, P. Bialecki, P. Januszewski, P. Thalasta, P. Gaikwad, P. Varshney, P. Gundecha, P. Tredak, R. K. Mahabadi, R. Patel, R. El-Yaniv, R. Rajan, R. Cheruvu, R. Shahbazyan, R. Borkar, R. Gala, R. Waleffe, R. Zhang, R. J. Hewett, R. Prenger, S. Jain, S. Kriman, S. Satheesh, S. Kaji, S. Yurick, S. Muralidharan, S. Narenthiran, S. Bak, S. Sameni, S. Han, S. Ramasamy, S. Ghosh, S. T. Sreenivas, S. Thomas, S. Diao, S. Gopal, S. Prabhumoye, S. Toshniwal, S. Ding, S. Singh, S. Jain, S. Majumdar, S. Singhal, S. Alborghetti, S. N. Akter, T. Kong, T. Moon, T. Hliwiak, T. Asida, T. Wang, T. Konuk, T. Vashishth, T. Poon, U. Karpas, V. Noroozi, V. Srinivasan, V. Korthikanti, V. Fugro, V. Kalluru, V. Kurin, V. Lavrukhin, W. U. Ahmad, W. Du, W. Byeon, X. Lu, X. Dong, Y. Karnati, Y. Choi, Y. Zhang, Y. Lin, Y. Fu, Y. Suhara, Z. Dong, Z. Li, Z. Zhu, and Z. Chen · 2025
Closest in time.
Nemotron-post-training-dataset-v1
NVIDIA AI · 2025
Closest in time.
Gpt-5 system card
OpenAI · 2025
Closest in time.
The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models
S. G. Patil, H. Mao, C. Cheng-Jie Ji, F. Yan, V. Suresh, I. Stoica, and J. E. Gonzalez · 2025
Closest in time.
Clipper: Compression enables long-context synthetic data generation, 2025
C. M. Pham, Y. Chang, and M. Iyyer · 2025
Closest in time.
Pipelinerl: Faster on-policy reinforcement learning for long sequence generatio
A. Piché, E. Kamaloo, R. Pardinas, and D. Bahdanau · 2025
Closest in time.
Synthetic-2
PrimeIntellect · 2025
Closest in time.
Generalizing verifiable instruction following
V. Pyatkin, S. Malik, V. Graf, H. Ivison, S. Huang, P. Dasigi, N. Lambert, and H. Hajishirzi · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning
Qwen Team · 2025
Closest in time.
Taskcraft: Automated generation of agentic tasks
D. Shi, J. Cao, Q. Chen, W. Sun, W. Li, H. Lu, F. Dong, T. Qin, K. Zhu, M. Liu, J. Yang, G. Zhang, J. Liu, C. Zhang, J. Wang, Y. E. Jiang, and W. Zhou · 2025
Closest in time.
IBM Granite 3.3: Speech recognition, refined reasoning, and RAG LoRAs, Apr. 2025
K. Soule and D. Bergmann · 2025
Closest in time.
Y. Sun, S. Hu, G. Zhou, K. Zheng, H. Hajishirzi, N. Dziri, and D. X. Song · 2025
Closest in time.
K2-v2: A 360-open, reasoning-enhanced llm, 2025
K. Team, Z. Liu, L. Tang, L. Jin, H. Li, N. Ranjan, D. Fan, S. Rohatgi, R. Fan, O. Pangarkar, H. Wang, Z. Cheng, S. Sun, S. Han, B. Tan, G. Gosal, X. Han, V. Pimpalkhute, S. Hao, M. S. Hee, J. Hestness, H. Jia, L. Ma, A. Singh, D. Soboleva, N. Vassilieva, R. Wang, Y. Wu, Y. Sun, T. Killian, A. Moreno, J. Maggs, H. Ren, G. He, H. Wang, X. Ma, Y. Wang, M. Yurochkin, and E. P. Xing · 2025
Closest in time.
The algorithms – python
The Algorithms · 2025
Closest in time.
Announcing rnj-1: Building instruments of intelligence, Dec. 2025
A. Vaswani · 2025
Closest in time.
Do large language model benchmarks test reliability?
J. Vendrow, E. Vendrow, S. Beery, and A. Madry · 2025
Closest in time.
Octothinker: Mid-training incentivizes reinforcement learning scaling
Z. Wang, F. Zhou, X. Li, and P. Liu · 2025
Closest in time.
Organize the web: Constructing domains enhances pre-training data curation, 2025
A. Wettig, K. Lo, S. Min, H. Hajishirzi, D. Chen, and L. Soldaini · 2025
Closest in time.
WebWalker: Benchmarking LLMs in web traversal
J. Wu, W. Yin, Y. Jiang, Z. Wang, Z. Xi, R. Fang, L. Zhang, Y. He, D. Zhou, P. Xie, and F. Huang · 2025
Closest in time.
MiMo: Unlocking the reasoning potential of language model – from pretraining to posttraining, 2025
L.-C. Xiaomi, :, B. Xia, B. Shen, Cici, D. Zhu, D. Zhang, G. Wang, H. Zhang, H. Liu, J. Xiao, J. Dong, L. Zhao, P. Li, P. Wang, S. Yu, S. Chen, W. Wang, W. Ma, X. Deng, Y. Huang, Y. Song, Z. Jiang, B. Ye, C. Cai, C. He, D. Zhang, D. Zhang, G. Wang, H. Tian, H. Zhao, H. Qu, H. Xu, J. Shi, K. Bao, K. Fang, K. Zhou, K. Zhou, L. Li, M. Zhu, N. Chen, Q. Wang, S. Liu, S. Li, S. Gu, S. Ren, S. Liu, S. Deng, W. Zhuang, W. Lv, W. Yang, X. Zhang, X. Yong, X. Zhang, X. Song, X. Xu, X. Wang, Y. Yan, Y. Tu, Y. Tian, Y. Wang, Y. Yu, Z. Lin, Z. Song, and Z. Yue · 2025
Closest in time.
Your efficient rl framework secretly brings you off-policy rl training, Aug. 2025
F. Yao, L. Liu, D. Zhang, C. Dong, J. Shang, and J. Gao · 2025
Closest in time.
Data mixing laws: Optimizing data mixtures by predicting language modeling performance, 2025
J. Ye, P. Liu, T. Sun, J. Zhan, Y. Zhou, and X. Qiu · 2025
Closest in time.
HELMET: How to evaluate long-context models effectively and thoroughly
H. Yen, T. Gao, M. Hou, K. Ding, D. Fleischer, P. Izsak, M. Wasserblat, and D. Chen · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, W. Dai, T. Fan, G. Liu, L. Liu, et al · 2025
Closest in time.
Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?
Y. Yue, Z. Chen, R. Lu, A. Zhao, Z. Wang, Y. Yue, S. Song, and G. Huang · 2025
Closest in time.
Megamath: Pushing the limits of open math corpora
F. Zhou, Z. Wang, N. Ranjan, Z. Cheng, L. Tang, G. He, Z. Liu, and E. P. Xing · 2025
Closest in time.
Cracks in the foundation: Architectural choices impact long context extension, 2026
A. Bertsch, L. Soldaini, M. Gormley, G. Neubig, H. Hajishirzi, K. Lo, and D. Groeneveld · 2026
Closest in time.
Olmix: Efficient mixture recomputation for evolving lm datasets, 2026
M. F. Chen, T. Murray, D. Heineman, M. Jordan, H. Hajishirzi, C. Ré, L. Soldaini, and K. Lo · 2026
Closest in time.