Fetching the paper…
Reading the bibliography…
Large language models (LLMs) face inherent performance bottlenecks under parameter constraints, particularly in processing critical tokens that demand complex reasoning.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 1905
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 1909
Earlier work this paper cites.
Vehicles: Experiments in Synthetic Psychology
Valentino Braitenberg. 1986 · 1986
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Felix Alexander Gers, Jürgen Schmidhuber, and Fred Cummins. 2000 · 2000
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh. 2006 · 2006
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020 · 2007
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016 · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017 · 2017
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Understanding generalization in recurrent neural networks
Zhuozhuo Tu, Fengxiang He, and Dacheng Tao. 2020 · 2020
Earlier work this paper cites.
Introducing pathways: A next-generation ai architecture
Jeff Dean. 2021 · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Mario Neumann, Rodolphe Jenatton, António Susano Pinto, Daniel Keysers, and Neil Houlsby. 2021 · 2021
Earlier work this paper cites.
composer
The Mosaic ML Team. 2021 · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2022 · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022 · 2022
Cited alongside, same era.
Mixture-of-experts with expert choice routing
Yutian Zhou, Tao Lei, Henry Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew M Dai, Quoc V Le, James Laudon, et al. 2022 · 2022
Cited alongside, same era.
Introducing claude
Anthropic. 2023 · 2023
Cited alongside, same era.
Implicit chain of thought reasoning via knowledge distillation
Yuntian Deng, Kiran Prasad, Roland Fernandez, Paul Smolensky, Vishrav Chaudhary, and Stuart Shieber. 2023 · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2023 · 2023
What happened in llms layers when trained for fast vs. slow thinking: A gradient perspective
Ming Li, Yanhong Li, and Tianyi Zhou. 2024 · 2024
Later among the works it cites.
Cross-layer attention sharing for large language models
Yongyu Mu, Yuzhang Wu, Yuchun Fan, Chenglong Wang, Hengyu Li, Qiaozhi He, Murun Yang, Tong Xiao, and Jingbo Zhu. 2024 · 2024
Later among the works it cites.
Loop neural networks for parameter sharing
Kei-Sing Ng and Qingchen Wang. 2024 · 2024
Later among the works it cites.
Mixture-of-depths: Dynamically allocating compute in transformer-based language models
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, and Adam Santoro. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
OpenAI. 2023 · 2023
Cited alongside, same era.
Deep Thinking Systems: Logical Extrapolation with Recurrent Neural Networks
A. Schwarzschild. 2023 · 2023
Cited alongside, same era.
Lessons on parameter sharing across layers in transformers
Sho Takase and Shun Kiyono. 2023 · 2023
Cited alongside, same era.
Redpajama: An open source recipe to reproduce llama training dataset
TogetherAI. 2023 · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Cited alongside, same era.
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2023 · 2023
Cited alongside, same era.
Duo-llm: A framework for studying adaptive computation in large language models
Keivan Alizadeh, Iman Mirzadeh, Hooman Shahrokhi, Dmitry Belenko, Frank Sun, Minsik Cho, Mohammad Hossein Sekhavat, Moin Nabi, and Mehrdad Farajtabar. 2024 · 2024
Cited alongside, same era.
Yuval Shalev, Amir Feder, and Ariel Goldstein. 2024 · 2024
Later among the works it cites.
Exposing the achilles’ heel: Evaluating llms ability to handle mistakes in mathematical reasoning
Joykirat Singh, Akshay Nambi, and Vibhav Vineet. 2024 · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Later among the works it cites.
X. Wu, S. Huang, and F. Wei. 2024 · 2024
Later among the works it cites.
Beyond model adaptation at test time: A survey
Zehao Xiao and Cees G. M. Snoek. 2024 · 2024
Later among the works it cites.
Openmoe: An early effort on open mixture-of-experts language models
F. Xue, Z. Zheng, Y. Fu, J. Ni, and W. Zhou. 2024 · 2024
Later among the works it cites.
p-mod: Building mixture-of-depths mllms via progressive ratio decay
Jun Zhang, Desen Meng, Ji Qi, Zhenpeng Huang, Tao Wu, and Limin Wang. 2024 · 2024
Later among the works it cites.
Scaling up test-time compute with latent reasoning: A recurrent depth approach
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein. 2025 · 2025
Closest in time.
Critical tokens matter: Token-level contrastive estimation enhances llm’s reasoning capability
Zicheng Lin, Tian Liang, Jiahao Xu, Qiuzhi Lin, Xing Wang, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang, and Zhaopeng Tu. 2025 · 2025
Closest in time.
Inference-time scaling for diffusion models beyond scaling denoising steps
Nanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu, Yu-Chuan Su, Mingda Zhang, Xuan Yang, Yandong Li, Tommi Jaakkola, Xuhui Jia, and Saining Xie. 2025 · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025 · 2025
Closest in time.
Transformer layers as painters
Qi Sun, Marc Pickett, Aakash Kumar Nain, and Llion Jones. 2025 · 2025
Closest in time.