Fetching the paper…
Reading the bibliography…
Characterizing neural networks in terms of better-understood formal systems has the potential to yield new insights into the power and limitations of these networks.
A logical calculus of ideas immanent in nervous activity
McCulloch, W. and Pitts, W · 1943
Earlier work this paper cites.
Representation of events in nerve nets and finite automata
Kleene, S. C · 1956
Earlier work this paper cites.
Weak second-order arithmetic and finite automata
Büchi, J. R · 1960
Earlier work this paper cites.
Elementary properties of ordered abelian groups
Robinson, A. and Zakon, E · 1960
Earlier work this paper cites.
A decision procedure for the first order theory of real addition with order
Ferrante, J. and Rackoff, C · 1975
Earlier work this paper cites.
Parity, circuits, and the polynomial-time hierarchy
Furst, M., Saxe, J. B., and Sipser, M · 1984
Earlier work this paper cites.
On uniformity within 𝑁𝐶 1 \mathit{NC}^{1}
Barrington, D. A. M., Immerman, N., and Straubing, H · 1990
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. and Sontag, E · 1995
Earlier work this paper cites.
Descriptive Complexity
Immerman, N · 1999
Earlier work this paper cites.
Finite-state computation in analog neural networks: Steps towards biologically plausible models?
Forcada, M. L. and Carrasco, R. C · 2001
Earlier work this paper cites.
On the linguistic capacity of real-time counter automata, 2020
Merrill, W · 2004
Cited alongside, same era.
Computability and Logic
Boolos, G. S., Burgess, J. P., and Jeffrey, R. C · 2007
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Recurrent neural networks as weighted language recognizers, 2017
Chen, Y., Gilroy, S., Knight, K., and May, J · 2018
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Hahn, M · 2020
Later among the works it cites.
Are Transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 2020
Later among the works it cites.
Attention is Turing-complete
Pérez, J., Barceló, P., and Marinkovic, J · 2021
Later among the works it cites.
Thinking like Transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Later among the works it cites.
Overcoming a theoretical limitation of self-attention
Chiang, D. and Cholak, P · 2022
Later among the works it cites.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Hao, Y., Angluin, D., and Frank, R · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schwartz, R., Thomson, S., and Smith, N. A · 2018
Cited alongside, same era.
Transformers without tears: Improving the normalization of self-attention
Nguyen, T. Q. and Salazar, J · 2019
Cited alongside, same era.
Learning deep Transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Cited alongside, same era.
The logical expressiveness of graph neural networks
Barceló, P., Kostylev, E. V., Monet, M., Pérez, J., Reutter, J., and Silva, J.-P · 2020
Cited alongside, same era.
On the ability and limitations of Transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Later among the works it cites.
Transformers can be translated to first-order logic with majority quantifiers, 2022
Merrill, W. and Sabharwal, A · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
Merrill, W., Sabharwal, A., and Smith, N. A · 2022
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers
Merrill, W. and Sabharwal, A · 2023
Closest in time.