Fetching the paper…
Reading the bibliography…
Length generalization, defined as the ability to extrapolate from shorter training sequences to longer test ones, is a significant challenge for language models.
Hybrid computing using a neural network with dynamic external memory
A. Graves, G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwinska, S. G. Colmenarejo, E. Grefenstette, T. Ramalho, J. P. Agapiou, A. P. Badia, K. M. Hermann, Y. Zwols, G. Ostrovski, A. Cain, H. King, C. Summerfield, P. Blunsom, K. Kavukcuoglu, and D. Hassabis · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2018
Earlier work this paper cites.
Self-attention with relative position representations
P. Shaw, J. Uszkoreit, and A. Vaswani · 2018
Earlier work this paper cites.
What algorithms can transformers learn? a study in length generalization
H. Zhou, A. Bradley, E. Littwin, N. Razin, O. Saremi, J. Susskind, S. Bengio, and P. Nakkiran · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Earlier work this paper cites.
Location attention for extrapolation to longer sequences
Y. Dubois, G. Dagan, D. Hupkes, and E. Bruni · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
D. Hupkes, V. Dankers, M. Mul, and E. Bruni · 2020
Earlier work this paper cites.
Understanding the failure modes of out-of-distribution generalization
V. Nagarajan, A. Andreassen, and B. Neyshabur · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
A data-centric approach for training deep neural networks with less data
M. Motamedi, N. Sakharnykh, and T. Kaldewey · 2021
Earlier work this paper cites.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
A. Schwarzschild, E. Borgnia, A. Gupta, F. Huang, U. Vishkin, M. Goldblum, and T. Goldstein · 2021
Earlier work this paper cites.
Primer: Searching for efficient transformers for language modeling
D. R. So, W. Mańke, H. Liu, Z. Dai, N. Shazeer, and Q. V. Le · 2021
Cited alongside, same era.
Exploring length generalization in large language models
C. Anil, Y. Wu, A. Andreassen, A. Lewkowycz, V. Misra, V. Ramasesh, A. Slone, G. Gur-Ari, E. Dyer, and B. Neyshabur · 2022
Cited alongside, same era.
Kerple: Kernelized relative positional embedding for length extrapolation
T.-C. Chi, T.-H. Fan, P. J. Ramadge, and A. Rudnicky · 2022
Cited alongside, same era.
Transformer language models without positional encodings still learn positional information
A. Haviv, O. Ram, O. Press, P. Izsak, and O. Levy · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al · 2022
Cited alongside, same era.
Extending context window of large language models via positional interpolation
S. Chen, S. Wong, L. Chen, and Y. Tian · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Later among the works it cites.
Neural networks and the chomsky hierarchy
G. Deletang, A. Ruoss, J. Grau-Moya, T. Genewein, L. K. Wenliang, E. Catt, C. Cundy, M. Hutter, S. Legg, J. Veness, and P. A. Ortega · 2023
Later among the works it cites.
From interpolation to extrapolation: Complete length generalization for arithmetic transformers
S. Duan and Y. Shi · 2023
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jian, B. Y. Lin, P. West, C. Bhagavatula, R. L. Bras, J. D. Hwang, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Competition-level code generation with alphacode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al · 2022
Cited alongside, same era.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
O. Press, N. Smith, and M. Lewis · 2022
Cited alongside, same era.
Lamda: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Cited alongside, same era.
The clrs algorithmic reasoning benchmark
P. Veličković, A. P. Badia, D. Budden, R. Pascanu, A. Banino, M. Dashevskiy, R. Hadsell, and C. Blundell · 2022
Cited alongside, same era.
Autoformalization with large language models
Y. Wu, A. Q. Jiang, W. Li, M. Rabe, C. Staats, M. Jamnik, and C. Szegedy · 2022
Cited alongside, same era.
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Length generalization in arithmetic transformers
S. Jelassi, S. d’Ascoli, C. Domingo-Enrich, Y. Wu, Y. Li, and F. Charton · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
A. Kazemnejad, I. Padhi, K. N. Ramamurthy, P. Das, and S. Reddy · 2023
Later among the works it cites.
Teaching arithmetic to small transformers
N. Lee, K. Sreenivasan, J. D. Lee, K. Lee, and D. Papailiopoulos · 2023
Later among the works it cites.
Functional interpolation for relative positions improves long context transformers
S. Li, C. You, G. Guruganesh, J. Ainslie, S. Ontanon, M. Zaheer, S. Sanghai, Y. Yang, S. Kumar, and S. Bhojanapalli · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Yarn: Efficient context window extension of large language models
B. Peng, J. Quesnelle, H. Fan, and E. Shippole · 2023
Later among the works it cites.
Randomized positional encodings boost length generalization of transformers
A. Ruoss, G. Delétang, T. Genewein, J. Grau-Moya, R. Csordás, M. Bennani, S. Legg, and J. Veness · 2023
Later among the works it cites.
Positional description matters for transformers arithmetic
R. Shen, S. Bubeck, R. Eldan, Y. T. Lee, Y. Li, and Y. Zhang · 2023
Later among the works it cites.
Rectified rotary position embeddings
J. Su · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Closest in time.