Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, especially when guided by explicit chain-of-thought (CoT) reasoning that verbalizes intermediate steps.
An outsider’s view of neural nets
Michael I. Jordan · 1986
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
On the computational power of neural nets
Hava T. Siegelmann and Eduardo D. Sontag · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan · 2019
Earlier work this paper cites.
On the turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló · 2019
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tarlow, and Rianne Van Den Berg · 2021
Earlier work this paper cites.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein · 2021
Earlier work this paper cites.
Recurrent memory transformer
Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev · 2022
Earlier work this paper cites.
Continuous diffusion for categorical data, 2022
Sander Dieleman, Laurent Sartran, Arman Roshannai, Nikolay Savinov, Yaroslav Ganin, Pierre H. Richemond, Arnaud Doucet, Robin Strudel, Chris Dyer, Conor Durkan, Curtis Hawthorne, Rémi Leblond, Will Grathwohl, and Jonas Adler · 2022
Earlier work this paper cites.
micse: Mutual information contrastive learning for low-shot sentence embeddings
Tassilo Klein and Moin Nabi · 2022
Earlier work this paper cites.
Diffusion-lm improves controllable text generation, 2022
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B. Hashimoto · 2022
Earlier work this paper cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Analog bits: Generating discrete data using diffusion models with self-conditioning, 2023
Ting Chen, Ruixiang Zhang, and Geoffrey Hinton · 2023
Earlier work this paper cites.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Earlier work this paper cites.
Likelihood-based diffusion language models, 2023
Ishaan Gulrajani and Tatsunori B. Hashimoto · 2023
Earlier work this paper cites.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Earlier work this paper cites.
Towards a mechanistic interpretation of multi-step reasoning capabilities of language models
Yifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, and Mrinmaya Sachan · 2023
Earlier work this paper cites.
Cotformer: A chain-of-thought driven architecture with budget-adaptive computation cost at inference
Amirkeivan Mohtashami, Matteo Pagliardini, and Martin Jaggi · 2023
Earlier work this paper cites.
Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan · 2023
Earlier work this paper cites.
Retentive network: A successor to transformer for large language models
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei · 2023
Earlier work this paper cites.
Gated linear attention transformers with hardware-efficient training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim · 2023
Earlier work this paper cites.
Relaxed recursive transformers: Effective parameter sharing with layer-wise lora
Sangmin Bae, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji, Seungyeon Kim, and Tal Schuster · 2024
Earlier work this paper cites.
Titans: Learning to memorize at test time
Ali Behrouz, Peilin Zhong, and Vahab Mirrokni · 2024
Earlier work this paper cites.
Transformers to ssms: Distilling quadratic knowledge to subquadratic models
Aviv Bick, Kevin Li, Eric Xing, J Zico Kolter, and Albert Gu · 2024
Earlier work this paper cites.
Hopping too late: Exploring the limitations of large language models on multi-hop queries
Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, and Amir Globerson · 2024
Earlier work this paper cites.
Iteration head: A mechanistic study of chain-of-thought
Vivien Cabannes, Charles Arnal, Wassim Bouaziz, Xingyu Yang, Francois Charton, and Julia Kempe · 2024
Earlier work this paper cites.
Compressed chain of thought: Efficient reasoning through dense representations
Jeffrey Cheng and Benjamin Van Durme · 2024
Earlier work this paper cites.
Investigating recurrent transformers with dynamic halt
Jishnu Ray Chowdhury and Cornelia Caragea · 2024
Earlier work this paper cites.
Tri Dao and Albert Gu · 2024
Cited alongside, same era.
Simulation of graph algorithms with looped transformers
Artur Back De Luca and Kimon Fountoulakis · 2024
Cited alongside, same era.
From explicit cot to implicit cot: Learning to internalize cot step by step
Yuntian Deng, Yejin Choi, and Stuart Shieber · 2024
Cited alongside, same era.
Learning iterative reasoning through energy diffusion
Yilun Du, Jiayuan Mao, and Joshua B. Tenenbaum · 2024
Cited alongside, same era.
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
Subhabrata Dutta, Joykirat Singh, Soumen Chakrabarti, and Tanmoy Chakraborty · 2024
Why lift so heavy? slimming large language models by cutting off the layers
Shuzhou Yuan, Ercong Nie, Bolei Ma, and Michael Färber · 2024
Later among the works it cites.
Quiet-star: Language models can teach themselves to think before speaking
Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber, and Noah D Goodman · 2024
Later among the works it cites.
Emergent abilities in large language models: A survey
Leonardo Berti, Flavio Giorgi, and Gjergji Kasneci · 2025
Closest in time.
Llamba: Scaling distilled recurrent models for efficient language processing
Aviv Bick, Tobias Katsch, Nimit Sohoni, Arjun Desai, and Albert Gu · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Algoformer: An efficient transformer framework with algorithmic structures
Yihang Gao, Chuanyang Zheng, Enze Xie, Han Shi, Tianyang Hu, Yu Li, Michael K Ng, Zhenguo Li, and Zhaoqiang Liu · 2024
Cited alongside, same era.
Can looped transformers learn to implement multi-step gradient descent for in-context learning?
Khashayar Gatmiry, Nikunj Saunshi, Sashank J Reddi, Stefanie Jegelka, and Sanjiv Kumar · 2024
Cited alongside, same era.
Think before you speak: Training language models with pause tokens
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, and Vaishnavh Nagarajan · 2024
Cited alongside, same era.
The unreasonable ineffectiveness of the deeper layers
Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A. Roberts · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian · 2024
Cited alongside, same era.
Opencoder: The open cookbook for top-tier code large language models
Siming Huang, Tianhao Cheng, Jason Klein Liu, Jiaran Hao, Liuyihan Song, Yang Xu, J. Yang, J. H. Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Zhaoxiang Zhang, Jie Fu, Qian Liu, Ge Zhang, Zili Wang, Yuan Qi, Yinghui Xu, and Wei Chu · 2024
Cited alongside, same era.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Cited alongside, same era.
Edoardo Cetin, Tianyu Zhao, and Yujin Tang · 2025
Closest in time.
Do language models use their depth efficiently?
Róbert Csordás, Christopher D Manning, and Christopher Potts · 2025
Closest in time.
Qingxiu Dong, Li Dong, Yao Tang, Tianzhu Ye, Yutao Sun, Zhifang Sui, and Furu Wei · 2025
Closest in time.
Scaling up test-time compute with latent reasoning: A recurrent depth approach
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein · 2025
Closest in time.
Scaling diffusion language models via adaptation from autoregressive models
Shansan Gong, Shivam Agarwal, Yizhe Zhang, Jiacheng Ye, Lin Zheng, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han, Hao Peng, and Lingpeng Kong · 2025
Closest in time.
Alex Graves, Rupesh Kumar Srivastava, Timothy Atkinson, and Faustino Gomez · 2025
Closest in time.
Reinforcing the diffusion chain of lateral thought with diffusion language models, 2025
Zemin Huang, Zhiyang Chen, Zijun Wang, Tiancheng Li, and Guo-Jun Qi · 2025
Closest in time.
Lattice: Learning to efficiently compress the memory
Mahdi Karami and Vahab Mirrokni · 2025
Closest in time.
Mercury: Ultra-fast language models based on diffusion, 2025
Inception Labs, Samar Khanna, Siddhant Kharbanda, Shufan Li, Harshit Varma, Eric Wang, Sawyer Birnbaum, Ziyang Luo, Yanis Miraoui, Akash Palrecha, Stefano Ermon, Aditya Grover, and Volodymyr Kuleshov · 2025
Closest in time.
Liger: Linearizing large language models to gated recurrent structures
Disen Lan, Weigao Sun, Jiaxi Hu, Jusen Du, and Yu Cheng · 2025
Closest in time.
Ge Lei and Samuel J Cooper · 2025
Closest in time.
Seek in the dark: Reasoning via test-time instance-level policy gradient in latent space
Hengli Li, Chenxi Li, Tong Wu, Xuekai Zhu, Yuxuan Wang, Zhaoxin Yu, Eric Hanchen Jiang, Song-Chun Zhu, Zixia Jia, Ying Nian Wu, et al · 2025
Closest in time.
Constant bit-size transformers are turing complete
Qian Li and Yuyi Wang · 2025
Closest in time.
dkv-cache: The cache for diffusion language models
Xinyin Ma, Runpeng Yu, Gongfan Fang, and Xinchao Wang · 2025
Closest in time.
A little depth goes a long way: The expressive power of log-depth transformers
William Merrill and Ashish Sabharwal · 2025
Closest in time.
Large language diffusion models
Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, JUN ZHOU, Yankai Lin, Ji-Rong Wen, and Chongxuan Li · 2025
Closest in time.
Reasoning with latent thoughts: On the power of looped transformers
Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J Reddi · 2025
Closest in time.
Implicit language models are RNNs: Balancing parallelization and expressivity
Mark Schöne, Babak Rahmani, Heiner Kremer, Fabian Falck, Hitesh Ballani, and Jannes Gladrow · 2025
Closest in time.
Mani Shemiranifar · 2025
Closest in time.
Codi: Compressing chain-of-thought into continuous space via self-distillation
Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, and Yulan He · 2025
Closest in time.
Layer by layer: Uncovering hidden representations in language models
Oscar Skean, Md Rifat Arefin, Dan Zhao, Niket Patel, Jalal Naghiyev, Yann LeCun, and Ravid Shwartz-Ziv · 2025
Closest in time.
Token assorted: Mixing latent and text tokens for improved language model reasoning
DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, and Qinqing Zheng · 2025
Closest in time.
The curse of depth in large language models
Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, and Shiwei Liu · 2025
Closest in time.
Tess 2: A large-scale generalist diffusion language model
Jaesung Tae, Hamish Ivison, Sachin Kumar, and Arman Cohan · 2025
Closest in time.
An explainable transformer circuit for compositional generalization
Cheng Tang, Brenden Lake, and Mehrdad Jazayeri · 2025
Closest in time.
Unpacking robustness in inflectional languages: Adversarial evaluation and mechanistic insights
PaweĹ Walkowiak, Marek Klonowski, Marcin Oleksy, and Arkadiusz Janz · 2025
Closest in time.
System-1.5 reasoning: Traversal in language and latent spaces with dynamic shortcuts
Xiaoqiang Wang, Suyuchen Wang, Yun Zhu, and Bang Liu · 2025
Closest in time.
Parallel continuous chain-of-thought with jacobi iteration
Haoyi Wu, Zhihao Teng, and Kewei Tu · 2025
Closest in time.
Dream 7b, 2025b
Jiacheng Ye, Zhihui Xie, Lin Zheng, Jiahui Gao, Zirui Wu, Xin Jiang, Zhenguo Li, and Lingpeng Kong · 2025
Closest in time.
Pretraining language models to ponder in continuous space
Boyi Zeng, Shixiang Song, Siyuan Huang, Yixuan Wang, He Li, Ziwei He, Xinbing Wang, Zhiyu Li, and Zhouhan Lin · 2025
Closest in time.
d1: Scaling reasoning in diffusion large language models via reinforcement learning, 2025
Siyan Zhao, Devaansh Gupta, Qinqing Zheng, and Aditya Grover · 2025
Closest in time.