Fetching the paper…
Reading the bibliography…
Large language models display remarkable capabilities in logical and mathematical reasoning, allowing them to solve complex tasks.
The perceptron: a probabilistic model for information storage and organization in the brain
Rosenblatt, F · 1958
Earlier work this paper cites.
The realization of symmetric switching functions with linear-input logical elements
Kautz, W. H · 1961
Earlier work this paper cites.
A theory of the learnable
Valiant, L. G · 1984
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. T. and Sontag, E. D · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Kearns, M · 1998
Earlier work this paper cites.
Noise-tolerant learning, the parity problem, and the statistical query model
Blum, A., Kalai, A., and Wasserman, H · 2003
Earlier work this paper cites.
Computational complexity: a modern approach
Arora, S. and Barak, B · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
A linear dynamical system model for text
Belanger, D. and Kakade, S · 2015
Earlier work this paper cites.
Solving general arithmetic word problems
Roy, S. and Roth, D · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Earlier work this paper cites.
Perceptrons, reissue of the 1988 expanded edition with a new foreword by Léon Bottou: an introduction to computational geometry
Minsky, M. and Papert, S. A · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning
Shalev-Shwartz, S., Shamir, O., and Shammah, S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Provable limitations of deep learning
Abbe, E. and Sandon, C · 2018
Earlier work this paper cites.
Fast learning requires good memory: A time-space lower bound for parity learning
Raz, R · 2018
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Learning parities with neural networks
Daniely, A. and Malach, E · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
When hardness of approximation meets hardness of learning
Malach, E. and Shalev-Shwartz, S · 2022
Later among the works it cites.
Limitations of language models in arithmetic and symbolic induction
Qian, J., Wang, H., Li, Z., Li, S., and Yan, X · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Sub-task decomposition enables learning in sequence to sequence tasks
Wies, N., Levine, Y., and Shashua, A · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Cited alongside, same era.
Pay attention to mlps
Liu, H., Dai, Z., So, D., and Le, Q. V · 2021
Cited alongside, same era.
Investigating the limitations of transformers with simple arithmetic tasks
Nogueira, R., Jiang, Z., and Lin, J · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al · 2021
Cited alongside, same era.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Cited alongside, same era.
Zhai, S., Talbott, W., Srivastava, N., Huang, C., Goh, H., Zhang, R., and Susskind, J · 2021
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Edelman, B. L., Goel, S., Kakade, S., and Zhang, C · 2022
Cited alongside, same era.
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Closest in time.
Tinystories: How small can language models be and still speak coherent english?
Eldan, R. and Li, Y · 2023
Closest in time.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Feng, G., Gu, Y., Zhang, B., Ye, H., He, D., and Wang, L · 2023
Closest in time.
Looped transformers as programmable computers
Giannou, A., Rajput, S., Sohn, J.-y., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Closest in time.
Teaching arithmetic to small transformers
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D · 2023
Closest in time.
Symbolic chain-of-thought distillation: Small models can also” think” step-by-step
Li, L. H., Hessel, J., Yu, Y., Ren, X., Chang, K.-W., and Choi, Y · 2023
Closest in time.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Closest in time.
Goat: Fine-tuned llama outperforms gpt-4 on arithmetic tasks
Liu, T. and Low, B. K. H · 2023
Closest in time.
Evaluating transformer language models on arithmetic operations using number decomposition
Muffo, M., Cocco, A., and Bertino, E · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al · 2023
Closest in time.
Languagetool, 2024
LanguageTool · 2024
Closest in time.