Fetching the paper…
Reading the bibliography…
One way to interpret the reasoning power of transformer-based language models is to describe the types of logical rules they can resolve over some input text.
Parity, circuits, and the polynomial-time hierarchy
Furst, M. L., Saxe, J. B., and Sipser, M · 1984
Earlier work this paper cites.
On uniformity within 𝖭𝖢 1 \mathsf{NC}^{1}
Barrington, D. A. M., Immerman, N., and Straubing, H · 1990
Earlier work this paper cites.
Division is in uniform 𝖳𝖢 0 \mathsf{TC}^{0}
Hesse, W · 2001
Earlier work this paper cites.
Computational Complexity: A Modern Approach
Arora, S. and Barak, B · 2009
Earlier work this paper cites.
Modern computer arithmetic , volume 18
Brent, R. P. and Zimmermann, P · 2010
Earlier work this paper cites.
Computing rational radical sums in uniform TC0
Hunter, P., Bouyer, P., Markey, N., Ouaknine, J., and Worrell, J · 2010
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
On the practical computational power of finite precision RNNs for language recognition
Weiss, G., Goldberg, Y., and Yahav, E · 2018
Earlier work this paper cites.
Reconciling deep learning with symbolic artificial intelligence: representing objects and relations
Garnelo, M. and Shanahan, M · 2019
Earlier work this paper cites.
On the Turing completeness of modern neural network architectures
Pérez, J., Marinković, J., and Barceló, P · 2019
Cited alongside, same era.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Hahn, M · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y · 2020
Cited alongside, same era.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Hao, Y., Angluin, D., and Frank, R · 2022
Closest in time.
Saturated transformers are constant-depth threshold circuits
Merrill, W., Sabharwal, A., and Smith, N. A · 2022
Closest in time.
LaMDA: Language models for dialog applications
Thoppilan, R., Freitas, D. D., Hall, J., Shazeer, N. M., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., Chen, D., Xu, Y., Chen, Z., Roberts, A., Bosma, M., Zhou, Y., Chang, C.-C., Krivokon, I. A., Rusch, W. J., Pickett, M., Meier-Hellstern, K. S., Morris, M. R., Doshi, T., Santos, R. D., Duke, T., Søraker, J. H., Zevenbergen, B., Prabhakaran, V., Díaz, M., Hutchinson, B., Olson, K., Molina, A., Hoffman-John, E., Lee, J., Aroyo, L., Rajakumar, R., Butryna, A., Lamm, M., Kuzmina, V. O., Fenton, J., Cohen, A., Bernstein, R., Kurzweil, R., Aguera-Arcas, B., Cui, C., Croak, M., Chi, E., and Le, Q · 2022
Closest in time.
Tighter bounds on the expressivity of transformer encoders
Chiang, D., Cholak, P., and Pillay, A · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
Merrill, W., Ramanujan, V., Goldberg, Y., Schwartz, R., and Smith, N. A · 2021
Cited alongside, same era.
Thinking like transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Cited alongside, same era.
TC 0 : Constant depth threshold circuits, 2022
Aaronson, S., Kuperberg, G., and Habryka, O · 2022
Cited alongside, same era.
Lindner, D., Kramár, J., Rahtz, M., McGrath, T., and Mikulik, V · 2023
Closest in time.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2023
Closest in time.
The parallelism tradeoff: Limitations of log-precision transformers
Merrill, W. and Sabharwal, A · 2023
Closest in time.
Transformers in uniform TC 0
Chiang, D · 2025
Closest in time.