Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are commonly trained on datasets consisting of fixed-length token sequences.
Learning and development in neural networks: The importance of starting small
J. L. Elman · 1993
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
M. Roemmele, C. A. Bejan, and A. S. Gordon · 2011
Earlier work this paper cites.
The winograd schema challenge
H. Levesque, E. Davis, and L. Morgenstern · 2012
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Towards machine comprehension of spoken content: Initial toefl listening comprehension test by machine, 2016
B.-H. Tseng, S.-S. Shen, H.-Y. Lee, and L.-S. Lee · 2016
Earlier work this paper cites.
Cyclical learning rates for training neural networks
L. N. Smith · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Supervised and unsupervised transfer learning for question answering
Y.-A. Chung, H.-Y. Lee, and J. Glass · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
P. Rajpurkar, R. Jia, and P. Liang · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
K. Lee, M.-W. Chang, and K. Toutanova · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
S. Reddy, D. Chen, and C. D. Manning · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al · 2020
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Cited alongside, same era.
Speed up training with variable length inputs by efficient batching strategies
Z. Ge, L. Kaushik, M. Omote, and S. Kumar · 2021
Cited alongside, same era.
M. M. Krell, M. Kosec, S. P. Perez, and A. Fitzgibbon · 2021
Cited alongside, same era.
Pre-training a bert with curriculum learning by increasing block-size of input text
K. Nagatsuka, C. Broni-Bediako, and M. Atsumi · 2021
Cited alongside, same era.
Lm-infinite: Simple on-the-fly length generalization for large language models
C. Han, Q. Wang, W. Xiong, Y. Chen, H. Ji, and S. Wang · 2023
Later among the works it cites.
Growlength: Accelerating llms pretraining by progressively growing training length
H. Jin, X. Han, J. Yang, Z. Jiang, C.-Y. Chang, and X. Hu · 2023
Later among the works it cites.
Scaling laws of rope-based extrapolation
X. Liu, H. Yan, S. Zhang, C. An, X. Qiu, and D. Lin · 2023
Later among the works it cites.
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Cited alongside, same era.
Sequence length is a domain: Length-based overfitting in transformer models
D. Variš and O. Bojar · 2021
Cited alongside, same era.
Exploring length generalization in large language models
C. Anil, Y. Wu, A. Andreassen, A. Lewkowycz, V. Misra, V. Ramasesh, A. Slone, G. Gur-Ari, E. Dyer, and B. Neyshabur · 2022
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang, et al · 2022
Cited alongside, same era.
xformers: A modular and hackable transformer modelling library
B. Lefaudeux, F. Massa, D. Liskovich, W. Xiong, V. Caggiano, S. Naren, M. Xu, J. Hu, M. Tintore, S. Zhang, P. Labatut, D. Haziza, L. Wehrstedt, J. Reizenstein, and G. Sizov · 2022
Cited alongside, same era.
The stability-efficiency dilemma: Investigating sequence length warmup for training gpt models
C. Li, M. Zhang, and Y. He · 2022
Cited alongside, same era.
Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation
B. Peng and J. Quesnelle · 2023
Later among the works it cites.
Yarn: Efficient context window extension of large language models
B. Peng, J. Quesnelle, H. Fan, and E. Shippole · 2023
Later among the works it cites.
Randomized positional encodings boost length generalization of transformers
A. Ruoss, G. Delétang, T. Genewein, J. Grau-Moya, R. Csordás, M. Bennani, S. Legg, and J. Veness · 2023
Later among the works it cites.
In-context pretraining: Language modeling beyond document boundaries
W. Shi, S. Min, M. Lomeli, C. Zhou, M. Li, V. Lin, N. A. Smith, L. Zettlemoyer, S. Yih, and M. Lewis · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
W. Xiong, J. Liu, I. Molybog, H. Zhang, P. Bhargava, R. Hou, L. Martin, R. Rungta, K. A. Sankararaman, B. Oguz, et al · 2023
Later among the works it cites.
Pose: Efficient context window extension of llms via positional skip-wise training
D. Zhu, N. Yang, L. Wang, Y. Song, W. Wu, F. Wei, and S. Li · 2023
Later among the works it cites.
Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
M. Dehghani, B. Mustafa, J. Djolonga, J. Heek, M. Minderer, M. Caron, A. Steiner, J. Puigcerver, R. Geirhos, I. M. Alabdulmohsin, et al · 2024
Closest in time.
Fewer truncations improve language modeling
H. Ding, Z. Wang, G. Paolini, V. Kumar, A. Deoras, D. Roth, and S. Soatto · 2024
Closest in time.
The impact of positional encoding on length generalization in transformers
A. Kazemnejad, I. Padhi, K. Natesan Ramamurthy, P. Das, and S. Reddy · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date, 2024
Meta · 2024
Closest in time.
Random-access infinite context length for transformers
A. Mohtashami and M. Jaggi · 2024
Closest in time.
S. Pawar, S. Tonmoy, S. Zaman, V. Jain, A. Chadha, and A. Das · 2024
Closest in time.
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, V. Hofmann, A. H. Jha, S. Kumar, L. Lucy, X. Lyu, N. Lambert, I. Magnusson, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. E. Peters, A. Ravichander, K. Richardson, Z. Shen, E. Strubell, N. Subramani, O. Tafjord, P. Walsh, L. Zettlemoyer, N. A. Smith, H. Hajishirzi, I. Beltagy, D. Groeneveld, J. Dodge, and K. Lo · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Closest in time.
Transformers can achieve length generalization but not robustly
Y. Zhou, U. Alon, X. Chen, X. Wang, R. Agarwal, and D. Zhou · 2024
Closest in time.