Fetching the paper…
Reading the bibliography…
In this paper, we investigate the inherent capabilities of transformer models in learning arithmetic algorithms, such as addition and parity.
The need for biases in learning generalizations
Mitchell, T. M · 1980
Earlier work this paper cites.
Position Information in Transformers: An Overview
Dufter, P., Schmitt, M., and Schütze, H · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
The cognitive roots of regularization in language, 2018
Ferdinand, V., Kirby, S., and Smith, K · 2018
Earlier work this paper cites.
Input combination strategies for multi-source transformer decoder
Libovický, J., Helcl, J., and Mareček, D · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context, 2019
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Bender, E. M. and Koller, A · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Hahn, M · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2020
Earlier work this paper cites.
Rect: A recursive transformer architecture for generalizable mathematical reasoning
Deshpande, R., Chen, J., and Lee, I. G · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Investigating the limitations of transformers with simple arithmetic tasks, 2021
Nogueira, R., Jiang, Z., and Lin, J · 2021
Cited alongside, same era.
Kerple: Kernelized relative positional embedding for length extrapolation, 2022
Chi, T.-C., Fan, T.-H., Ramadge, P. J., and Rudnicky, A. I · 2022
Cited alongside, same era.
Overcoming a theoretical limitation of self-attention
Chiang, D. and Cholak, P · 2022
Cited alongside, same era.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Csordás, R., Irie, K., and Schmidhuber, J · 2022
Cited alongside, same era.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Dissecting transformer length extrapolation via the lens of receptive field analysis
Chi, T.-C., Fan, T.-H., Rudnicky, A., and Ramadge, P · 2023
Closest in time.
Neural networks and the chomsky hierarchy
Deletang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L. K., Catt, E., Cundy, C., Hutter, M., Legg, S., Veness, J., and Ortega, P. A · 2023
Closest in time.
Length generalization in arithmetic transformers, 2023
Jelassi, S., d’Ascoli, S., Domingo-Enrich, C., Wu, Y., Li, Y., and Charton, F · 2023
Closest in time.
The impact of positional encoding on length generalization in transformers, 2023
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Closest in time.
Teaching arithmetic to small transformers, 2023
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D · 2023
Closest in time.
Yarn: Efficient context window extension of large language models, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M · 2022
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2022
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2022
Cited alongside, same era.
Teaching algorithmic reasoning via in-context learning, 2022
Zhou, H., Nova, A., Larochelle, H., Courville, A., Neyshabur, B., and Sedghi, H · 2022
Cited alongside, same era.
Extending context window of large language models via positional interpolation, 2023
Chen, S., Wong, S., Chen, L., and Tian, Y · 2023
Cited alongside, same era.
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Closest in time.
Randomized positional encodings boost length generalization of transformers
Ruoss, A., Delétang, G., Genewein, T., Grau-Moya, J., Csordás, R., Bennani, M., Legg, S., and Veness, J · 2023
Closest in time.
Gpt can solve mathematical problems without a calculator, 2023
Yang, Z., Ding, M., Lv, Q., Jiang, Z., He, Z., Guo, Y., Bai, J., and Tang, J · 2023
Closest in time.
What algorithms can transformers learn? a study in length generalization
Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J., Bengio, S., and Nakkiran, P · 2023
Closest in time.
With greater text comes greater necessity: Inference-time training helps long text generation, 2024
Wang, Y., Ma, D., and Cai, D · 2024
Closest in time.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2074
Closest in time.