Fetching the paper…
Reading the bibliography…
Recent advancements have demonstrated that the performance of large language models (LLMs) can be significantly enhanced by scaling computational resources at test time.
Transformers are rnns: Fast autoregressive transformers with linear attention, 2020
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2006
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., and et. al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Finetuning pretrained transformers into rnns, 2021
Kasai, J., Peng, H., Zhang, Y., Yogatama, D., Ilharco, G., Pappas, N., Mao, Y., Chen, W., and Smith, N. A · 2021
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces, 2022
Gu, A., Goel, K., and Ré, C · 2022
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback, 2022
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Emergent abilities of large language models, 2022
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
Dao, T · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Textbooks are all you need ii: phi-1.5 technical report, 2023
Li, Y., Bubeck, S., Eldan, R., Giorno, A. D., Gunasekar, S., and Lee, Y. T · 2023
Earlier work this paper cites.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
Rwkv: Reinventing rnns for the transformer era, 2023
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., He, X., Hou, H., Lin, J., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., and et. al · 2023
Earlier work this paper cites.
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
Teknium · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models, 2023
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2023
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
Planning with large language models for code generation, 2023
Zhang, S., Chen, Z., Shen, Y., Ding, M., Tenenbaum, J. B., and Gan, C · 2023
Cited alongside, same era.
xlstm: Extended long short-term memory, 2024
Beck, M., Pöppel, K., Spanring, M., Auer, A., Prudnikova, O., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2024
Cited alongside, same era.
Scaling test-time compute with open models, 2024
Beeching, E., Tunstall, L., and Rush, S · 2024
Cited alongside, same era.
Transformers to ssms: Distilling quadratic knowledge to subquadratic models, 2024
Let’s think dot by dot: Hidden computation in transformer language models, 2024
Pfau, J., Merrill, W., and Bowman, S. R · 2024
Later among the works it cites.
Scavenging hyena: Distilling transformers into long convolution models, 2024
Ralambomihanta, T. R., Mohammadzadeh, S., Islam, M. S. N., Jabbour, W., and Liang, L · 2024
Later among the works it cites.
Samba: Simple hybrid state space models for efficient unlimited context language modeling, 2024
Ren, L., Liu, Y., Lu, Y., Shen, Y., Liang, C., and Chen, W · 2024
Later among the works it cites.
The effect of sampling temperature on problem solving in large language models, 2024
Renze, M. and Guven, E · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bick, A., Li, K. Y., Xing, E. P., Kolter, J. Z., and Gu, A · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling, 2024
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Dao, T. and Gu, A · 2024
Cited alongside, same era.
Hymba: A hybrid-head architecture for small language models, 2024
Dong, X., Fu, Y., Diao, S., Byeon, W., Chen, Z., Mahabaleshwarkar, A. S., Liu, S.-Y., Keirsbilck, M. V., Chen, M.-H., Suhara, Y., Lin, Y., Kautz, J., and Molchanov, P · 2024
Cited alongside, same era.
Think before you speak: Training language models with pause tokens, 2024
Goyal, S., Ji, Z., Rawat, A. S., Menon, A. K., Kumar, S., and Nagarajan, V · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., and et. al · 2024
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2024
Gu, A. and Dao, T · 2024
Cited alongside, same era.
Minillm: Knowledge distillation of large language models, 2024
Gu, Y., Dong, L., Wei, F., and Huang, M · 2024
Cited alongside, same era.
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Openmathinstruct-2: Accelerating ai for math with massive open-source instruction data, 2024
Toshniwal, S., Du, W., Moshkov, I., Kisacanin, B., Ayrapetyan, A., and Gitman, I · 2024
Later among the works it cites.
The mamba in the llama: Distilling and accelerating hybrid models
Wang, J., Paliotta, D., May, A., Rush, A. M., and Dao, T · 2024
Later among the works it cites.
Wu, Y., Sun, Z., Li, S., Welleck, S., and Yang, Y · 2024
Later among the works it cites.
An implementation of generative prm
Xiong, W., Zhang, H., Jiang, N., and Zhang, T · 2024
Later among the works it cites.
A survey on knowledge distillation of large language models, 2024
Xu, X., Li, M., Tao, C., Shen, T., Cheng, R., Li, J., Xu, C., Tao, D., and Zhou, T · 2024
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training, 2024
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2024
Later among the works it cites.
Llamba: Scaling distilled recurrent models for efficient language processing, 2025
Bick, A., Katsch, T., Sohoni, N., Desai, A., and Gu, A · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., and et. al · 2025
Closest in time.
Bespoke-stratos: The unreasonable effectiveness of reasoning distillation
Labs, B · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., and et. al · 2025
Closest in time.
The mamba in the llama: Distilling and accelerating hybrid models
Wang, J., Paliotta, D., May, A., Rush, A., and Dao, T · 2025
Closest in time.
Towards large reasoning models: A survey of reinforced reasoning with large language models, 2025
Xu, F., Hao, Q., Zong, Z., Wang, J., Zhang, Y., Wang, J., Lan, X., Gong, J., Ouyang, T., Meng, F., Shao, C., Yan, Y., Yang, Q., Song, Y., Ren, S., Hu, X., Li, Y., Feng, J., Gao, C., and Li, Y · 2025
Closest in time.