Fetching the paper…
Reading the bibliography…
Looped Transformers provide advantages in parameter efficiency, computational capabilities, and generalization for reasoning tasks.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R · 1993
Earlier work this paper cites.
Almost linear vc dimension bounds for piecewise polynomial networks
Bartlett, P., Maiorov, V., and Meir, R · 1998
Earlier work this paper cites.
Ha, D., Dai, A., and Le, Q. V · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Fixing weight decay regularization in adam, 2018
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Optimal approximation of continuous functions by very deep relu networks
Yarotsky, D · 2018
Earlier work this paper cites.
Deep equilibrium models
Bai, S., Kolter, J. Z., and Koltun, V · 2019
Earlier work this paper cites.
Universal transformers
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Earlier work this paper cites.
3 million sudoku puzzles with ratings, 2020
Radcliffe, D. G · 2020
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Cited alongside, same era.
Lessons on parameter sharing across layers in transformers
Takase, S. and Kiyono, S · 2021
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P., and Valiant, G · 2022
Cited alongside, same era.
Efficiently scaling transformer inference, 2022
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2022
Cited alongside, same era.
On the optimal memorization power of reLU neural networks
Vardi, G., Yehudai, G., and Shamir, O · 2022
Cited alongside, same era.
Learning to solve constraint satisfaction problems with recurrent transformer
Yang, Z., Ishay, A., and Lee, J · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. R · 2023
Later among the works it cites.
On enhancing expressive power via compositions of single fixed-size relu network
Zhang, S., Lu, J., and Zhao, H · 2023
Later among the works it cites.
MoEUT: Mixture-of-experts universal transformers
Csordás, R., Irie, K., Schmidhuber, J., Potts, C., and Manning, C. D · 2024
Closest in time.
Stream of search (sos): Learning to search in language
Gandhi, K., Lee, D. H. J., Grand, G., Liu, M., Cheng, W., Sharma, A., and Goodman, N · 2024
Closest in time.
How well can transformers emulate in-context newton’s method?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Cited alongside, same era.
Neural networks and the chomsky hierarchy
Deletang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L. K., Catt, E., Cundy, C., Hutter, M., Legg, S., Veness, J., and Ortega, P. A · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
et al., H. T · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Feng, G., Zhang, B., Gu, Y., Ye, H., He, D., and Wang, L · 2023
Cited alongside, same era.
Looped transformers as programmable computers
Giannou, A., Rajput, S., Sohn, J.-Y., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Cited alongside, same era.
Provable memorization capacity of transformers
Kim, J., Kim, M., and Mozafari, B · 2023
Cited alongside, same era.
The parallelism tradeoff: Limitations of log-precision transformers
Merrill, W. and Sabharwal, A · 2023
Cited alongside, same era.
Giannou, A., Yang, L., Wang, T., Papailiopoulos, D., and Lee, J. D · 2024
Closest in time.
Approximation rate of the transformer architecture for sequence modeling
Jiang, H. and Li, Q · 2024
Closest in time.
Are transformers with one layer self-attention using low-rank weight matrices universal approximators?
Kajitsuka, T. and Sato, I · 2024
Closest in time.
Position: LLMs can’t plan, but can help planning in LLM-modulo frameworks
Kambhampati, S., Valmeekam, K., Guan, L., Verma, M., Stechly, K., Bhambri, S., Saldyt, L. P., and Murthy, A. B · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Team, G · 2024
Closest in time.
Understanding the expressive power and mechanisms of transformer for sequence modeling
Wang, M. and E, W · 2024
Closest in time.
Looped transformers are better at learning learning algorithms
Yang, L., Lee, K., Nowak, R. D., and Papailiopoulos, D · 2024
Closest in time.
Relaxed recursive transformers: Effective parameter sharing with layer-wise loRA
Bae, S., Fisch, A., Harutyunyan, H., Ji, Z., Kim, S., and Schuster, T · 2025
Closest in time.
Reasoning with latent thoughts: On the power of looped transformers
Saunshi, N., Dikkala, N., Li, Z., Kumar, S., and Reddi, S. J · 2025
Closest in time.