Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable capabilities in various natural language processing tasks.
Boolq: Exploring the surprising difficulty of natural yes/no questions, 2019
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 1905
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 1905
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021a
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer, 2021
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2010
Earlier work this paper cites.
Fasttext.zip: Compressing text classification models
Joulin, A., Grave, E., Bojanowski, P., Douze, M., Jégou, H., and Mikolov, T · 2016
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension, 2017
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Senteval: An evaluation toolkit for universal sentence representations, 2018
Conneau, A. and Kiela, D · 2018
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models, 2022
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher, 2022
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2022
Earlier work this paper cites.
What language model to train if you have one million gpu hours?, 2022
Scao, T. L., Wang, T., Hesslow, D., Saulnier, L., Bekman, S., Bari, M. S., Biderman, S., Elsahar, H., Muennighoff, N., Phang, J., Press, O., Raffel, C., Sanh, V., Shen, S., Sutawika, L., Tae, J., Yong, Z. X., Launay, J., and Beltagy, I · 2022
Earlier work this paper cites.
Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model
Soltan, S., Ananthakrishnan, S., FitzGerald, J., Gupta, R., Hamza, W., Khan, H., Peris, C., Rawls, S., Rosenbaum, A., Rumshisky, A., et al · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., , and Wei, J · 2022
Earlier work this paper cites.
Finetuned language models are zero-shot learners, 2022
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Earlier work this paper cites.
Llemma: An open language model for mathematics, 2023
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2023
Earlier work this paper cites.
When do program-of-thoughts work for reasoning?, 2023
Bi, Z., Zhang, N., Jiang, Y., Deng, S., Zheng, G., and Chen, H · 2023
Cited alongside, same era.
Theoremqa: A theorem-driven question answering dataset, 2023
Chen, W., Yin, M., Ku, M., Lu, P., Wan, Y., Ma, X., Xu, J., Wang, X., and Xia, T · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
Dao, T · 2023
Cited alongside, same era.
Lighteval: A lightweight framework for llm evaluation, 2023
Fourrier, C., Habib, N., Kydlíček, H., Wolf, T., and Tunstall, L · 2023
Cited alongside, same era.
Complexity-based prompting for multi-step reasoning, 2023
Fu, Y., Peng, H., Sabharwal, A., Clark, P., and Khot, T · 2023
Cited alongside, same era.
OLMo, T., Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., Lambert, N., Schwenk, D., Tafjord, O., Anderson, T., Atkinson, D., Brahman, F., Clark, C., Dasigi, P., Dziri, N., Guerquin, M., Ivison, H., Koh, P. W., Liu, J., Malik, S., Merrill, W., Miranda, L. J. V., Morrison, J., Murray, T., Nam, C., Pyatkin, V., Rangapur, A., Schmitz, M., Skjonsberg, S., Wadden, D., Wilhelm, C., Wilson, M., Zettlemoyer, L., Farhadi, A., Smith, N. A., and Hajishirzi, H · 2024
Later among the works it cites.
Learning to reason with llms, September 2024
OpenAI · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y. K., Wu, Y., and Guo, D · 2024
Later among the works it cites.
Singh, S., Romanou, A., Fourrier, C., Adelani, D. I., Ngui, J. G., Vila-Suero, D., Limkonchotiwat, P., Marchisio, K., Leong, W. Q., Susanto, Y., Ng, R., Longpre, S., Ko, W.-Y., Smith, M., Bosselut, A., Oh, A., Martins, A. F. T., Choshen, L., Ippolito, D., Ferrante, E., Fadaee, M., Ermis, B., and Hooker, S · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neftune: Noisy embeddings improve instruction finetuning, 2023
Jain, N., yeh Chiang, P., Wen, Y., Kirchenbauer, J., Chu, H.-M., Somepalli, G., Bartoldson, B. R., Kailkhura, B., Schwarzschild, A., Saha, A., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Madlad-400: A multilingual and document-level large audited dataset, 2023
Kudugunta, S., Caswell, I., Zhang, B., Garcia, X., Choquette-Choo, C. A., Lee, K., Xin, D., Kusupati, A., Stella, R., Bapna, A., and Firat, O · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Lai, V. D., Nguyen, C. V., Ngo, N. T., Nguyen, T., Dernoncourt, F., Rossi, R. A., and Nguyen, T. H · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., et al · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Cited alongside, same era.
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations, 2024
Wang, P., Li, L., Shao, Z., Xu, R. X., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z · 2024
Later among the works it cites.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing
Xu, Z., Jiang, F., Niu, L., Deng, Y., Poovendran, R., Choi, Y., and Lin, B. Y · 2024
Later among the works it cites.
Qwen2.5-math technical report: Toward mathematical expert model via self-improvement
Yang, A., Zhang, B., Hui, B., Gao, B., Yu, B., Li, C., Liu, D., Tu, J., Zhou, J., Lin, J., Lu, K., Xue, M., Lin, R., Liu, T., Ren, X., and Zhang, Z · 2024
Later among the works it cites.
Language agent tree search unifies reasoning acting and planning in language models, 2024
Zhou, A., Yan, K., Shlapentokh-Rothman, M., Wang, H., and Wang, Y.-X · 2024
Later among the works it cites.
Dolphin-r1 dataset
Computations, C · 2025
Closest in time.
Liger kernel: Efficient triton kernels for llm training, 2025
Hsu, P.-L., Dai, Y., Kothapalli, V., Song, Q., Tang, S., Zhu, S., Shimizu, S., Sahni, S., Ning, H., and Chen, Y · 2025
Closest in time.
Hyperbolic: Access ai cloud
Hyperbolic · 2025
Closest in time.
Menlo research
Ltd., M. R. P · 2025
Closest in time.
Packing inputs without cross-contamination attention
MeetKai · 2025
Closest in time.
Modal: High-performance ai infrastructure
Modal · 2025
Closest in time.
s1: Simple test-time scaling, 2025
Muennighoff, N., Yang, Z., Shi, W., Li, X. L., Fei-Fei, L., Hajishirzi, H., Zettlemoyer, L., Liang, P., Candès, E., and Hashimoto, T · 2025
Closest in time.
Openr1-math-220k dataset
R1, O · 2025
Closest in time.
Kimi k1.5: Scaling reinforcement learning with llms, 2025
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., Tang, C., Wang, C., Zhang, D., Yuan, E., Lu, E., Tang, F., Sung, F., Wei, G., Lai, G., Guo, H., Zhu, H., Ding, H., Hu, H., Yang, H., Zhang, H., Yao, H., Zhao, H., Lu, H., Li, H., Yu, H., Gao, H., Zheng, H., Yuan, H., Chen, J., Guo, J., Su, J., Wang, J., Zhao, J., Zhang, J., Liu, J., Yan, J., Wu, J., Shi, L., Ye, L., Yu, L., Dong, M., Zhang, N., Ma, N., Pan, Q., Gong, Q., Liu, S., Ma, S., Wei, S., Cao, S., Huang, S., Jiang, T., Gao, W., Xiong, W., He, W., Huang, W., Wu, W., He, W., Wei, X., Jia, X., Wu, X., Xu, X., Zu, X., Zhou, X., Pan, X., Charles, Y., Li, Y., Hu, Y., Liu, Y., Chen, Y., Wang, Y., Liu, Y., Qin, Y., Liu, Y., Yang, Y., Bao, Y., Du, Y., Wu, Y., Wang, Y., Zhou, Z., Wang, Z., Li, Z., Zhu, Z., Zhang, Z., Wang, Z., Yang, Z., Huang, Z., Huang, Z., Xu, Z., and Yang, Z · 2025
Closest in time.
Open thoughts, January 2025
Team, O. T · 2025
Closest in time.
Towards system 2 reasoning in llms: Learning how to think with meta chain-of-thought, 2025
Xiang, V., Snell, C., Gandhi, K., Albalak, A., Singh, A., Blagden, C., Phung, D., Rafailov, R., Lile, N., Mahan, D., Castricato, L., Franken, J.-P., Haber, N., and Finn, C · 2025
Closest in time.
Limo: Less is more for reasoning, 2025
Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., and Liu, P · 2025
Closest in time.