Fetching the paper…
Reading the bibliography…
Large language models based on the transformer architecture can solve highly complex tasks, yet their fundamental limitations on simple algorithmic problems remain poorly understood.
Lower bounds on the maximum cross correlation of signals (corresp.)
L. Welch · 1974
Earlier work this paper cites.
Some complexity questions related to distributive computing (preliminary report)
A. C.-C. Yao · 1979
Earlier work this paper cites.
The space complexity of approximating the frequency moments
N. Alon, Y. Matias, and M. Szegedy · 1996
Earlier work this paper cites.
All-distances sketches, revisited: Hip estimators for massive graphs analysis
E. Cohen · 2014
Earlier work this paper cites.
Benefits of depth in neural networks
M. Telgarsky · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Exploring length generalization in large language models
C. Anil, Y. Wu, A. Andreassen, A. Lewkowycz, V. Misra, V. Ramasesh, A. Slone, G. Gur-Ari, E. Dyer, and B. Neyshabur · 2022
Cited alongside, same era.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Cited alongside, same era.
Efficient long-text understanding with short-text models
M. Ivgi, U. Shaham, and J. Berant · 2023
Cited alongside, same era.
The parallelism tradeoff: Limitations of log-precision transformers
W. Merrill and A. Sabharwal · 2023
Cited alongside, same era.
Representational strengths and limitations of transformers
C. Sanford, D. J. Hsu, and M. Telgarsky · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolatingthe capabilities of language models
Transformers need glasses! information over-squashing in language tasks, 2024
F. Barbero, A. Banino, S. Kapturowski, D. Kumaran, J. G. M. Araújo, A. Vitvitskyi, R. Pascanu, and P. Veličković · 2024
Closest in time.
Towards revealing the mystery behind chain of thought: a theoretical perspective
G. Feng, B. Zhang, Y. Gu, H. Ye, D. He, and L. Wang · 2024
Closest in time.
Needle in a haystack - pressure testing llms, 2024
G. Kamradt · 2024
Closest in time.
Same task, more tokens: the impact of input length on the reasoning performance of large language models, 2024
M. Levy, A. Jacoby, and Y. Goldberg · 2024
Closest in time.
How many neurons does it take to approximate the maximum?
I. Safran, D. Reichman, and P. Valiant · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Srivastava, D. Kleyjo, and Z. Wu · 2023
Cited alongside, same era.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
C. Wei, Y. Chen, and T. Ma
Cited in the paper.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al
Cited in the paper.
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2024
Closest in time.