Fetching the paper…
Reading the bibliography…
Large language models often struggle with length generalization and solving complex problem instances beyond their training distribution.
Bagging predictors
Breiman, L · 1996
Earlier work this paper cites.
Deep learning is robust to massive label noise
Rolnick, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Location attention for extrapolation to longer sequences
Dubois, Y., Dagan, G., Hupkes, D., and Bruni, E · 2019
Earlier work this paper cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Zhang, L., Song, J., Gao, A., Chen, J., Bao, C., and Ma, K · 2019
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
Hupkes, D., Dankers, V., Mul, M., and Bruni, E · 2020
Earlier work this paper cites.
The eos decision and length extrapolation
Newman, B., Hewitt, J., Liang, P., and Manning, C. D · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, J. E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Earlier work this paper cites.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Schwarzschild, A., Borgnia, E., Gupta, A., Huang, F., Vishkin, U., Goldblum, M., and Goldstein, T · 2021
Earlier work this paper cites.
From local structures to size generalization in graph neural networks
Yehudai, G., Fetaya, E., Meirom, E., Chechik, G., and Maron, H · 2021
Earlier work this paper cites.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Earlier work this paper cites.
End-to-end algorithm synthesis with recurrent networks: Extrapolation without overthinking
Bansal, A., Schwarzschild, A., Borgnia, E., Emam, Z., Huang, F., Goldblum, M., and Goldstein, T · 2022
Earlier work this paper cites.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., and Goodman, N. D · 2022
Earlier work this paper cites.
Self-consuming generative models go mad
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H. R., LeJeune, D., Siahkoohi, A., and Baraniuk, R · 2023
Earlier work this paper cites.
On the stability of iterative retraining of generative models on their own data
Bertrand, Q., Bose, A. J., Duplessis, A., Jiralerspong, M., and Gidel, G · 2023
Earlier work this paper cites.
Large language models suffer from their own output: An analysis of the self-consuming training loop
Briesch, M., Sobania, D., and Rothlauf, F · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al · 2023
Earlier work this paper cites.
Teaching large language models to self-debug
Chen, X., Lin, M., Schärli, N., and Zhou, D · 2023
Earlier work this paper cites.
de Arcaute, G. M. R., Watson, L., Reviriego, P., Hernández, J. A., Juárez, M., and Sarkar, R · 2023
Earlier work this paper cites.
From interpolation to extrapolation: Complete length generalization for arithmetic transformers
Duan, S., Shi, Y., and Xu, W · 2023
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al · 2023
Earlier work this paper cites.
Will large-scale generative models corrupt future datasets?
Hataya, R., Bao, H., and Arai, H · 2023
Cited alongside, same era.
Length generalization in arithmetic transformers
Jelassi, S., d’Ascoli, S., Domingo-Enrich, C., Wu, Y., Li, Y., and Charton, F · 2023
Cited alongside, same era.
Teaching arithmetic to small transformers
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D · 2023
Cited alongside, same era.
Functional interpolation for relative positions improves long context transformers
Li, S., You, C., Guruganesh, G., Ainslie, J., Ontanon, S., Zaheer, M., Sanghai, S., Yang, Y., Kumar, S., and Bhojanapalli, S · 2023
Cited alongside, same era.
Beyond model collapse: Scaling up with synthesized data requires reinforcement
Feng, Y., Dohmatob, E., Yang, P., Charton, F., and Kempe, J · 2024
Later among the works it cites.
Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Sleight, H., Hughes, J., Korbak, T., Agrawal, R., Pai, D., Gromov, A., et al · 2024
Later among the works it cites.
Self-correcting self-consuming loops for generative model training
Gillman, N., Freeman, M., Aggarwal, D., Hsu, C.-H., Luo, C., Tian, Y., and Sun, C · 2024
Later among the works it cites.
The unreasonable effectiveness of easy training data for hard tasks
Hase, P., Bansal, M., Clark, P., and Wiegreffe, S · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Understanding addition in transformers
Quirke, P. and Barez, F · 2023
Cited alongside, same era.
Randomized positional encodings boost length generalization of transformers
Ruoss, A., Delétang, G., Genewein, T., Grau-Moya, J., Csordás, R., Bennani, M., Legg, S., and Veness, J · 2023
Cited alongside, same era.
Positional description matters for transformers arithmetic
Shen, R., Bubeck, S., Eldan, R., Lee, Y. T., Li, Y., and Zhang, Y · 2023
Cited alongside, same era.
The curse of recursion: Training on generated data makes models forget
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R · 2023
Cited alongside, same era.
Beyond human data: Scaling self-training for problem-solving with language models
Singh, A., Co-Reyes, J. D., Agarwal, R., Anand, A., Patil, P., Garcia, X., Liu, P. J., Harrison, J., Lee, J., Xu, K., et al · 2023
Cited alongside, same era.
Chain-of-thought reasoning is a policy improvement operator
Zhang, H. and Parkes, D. C · 2023
Cited alongside, same era.
What algorithms can transformers learn? a study in length generalization
Zhou, H., Bradley, A., Littwin, E., Razin, N., Saremi, O., Susskind, J., Bengio, S., and Nakkiran, P · 2023
Cited alongside, same era.
Hosseini, A., Yuan, X., Malkin, N., Courville, A., Sordoni, A., and Agarwal, R · 2024
Later among the works it cites.
Self-improvement in language models: The sharpening mechanism, 2024
Huang, A., Block, A., Foster, D. J., Rohatgi, D., Zhang, C., Simchowitz, M., Ash, J. T., and Krishnamurthy, A · 2024
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Natesan Ramamurthy, K., Das, P., and Reddy, S · 2024
Later among the works it cites.
I-sheep: Self-alignment of llm from scratch through an iterative self-enhancement paradigm
Liang, Y., Zhang, G., Qu, X., Zheng, T., Guo, J., Du, X., Yang, Z., Liu, J., Lin, C., Ma, L., et al · 2024
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2024
Later among the works it cites.
Transformers can do arithmetic with the right embeddings
McLeish, S., Bansal, A., Stein, A., Jain, N., Kirchenbauer, J., Bartoldson, B. R., Kailkhura, B., Bhatele, A., Geiping, J., Schwarzschild, A., et al · 2024
Later among the works it cites.
Iterative reasoning preference optimization, 2024
Pang, R. Y., Yuan, W., Cho, K., He, H., Sukhbaatar, S., and Weston, J · 2024
Later among the works it cites.
Regenesis: Llms can grow into reasoning generalists via self-improvement
Peng, X., Xia, C., Yang, X., Xiong, C., Wu, C.-S., and Xing, C · 2024
Later among the works it cites.
Recursive introspection: Teaching language model agents how to self-improve
Qu, Y., Zhang, T., Garg, N., and Kumar, A · 2024
Later among the works it cites.
Explicitly encoding structural symmetry is key to length generalization in arithmetic tasks
Sabbaghi, M., Pappas, G., Hassani, H., and Goel, S · 2024
Later among the works it cites.
Weak-to-strong generalization through the data-centric lens
Shin, C., Cooper, J., and Sala, F · 2024
Later among the works it cites.
Ai models collapse when trained on recursively generated data
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y · 2024
Later among the works it cites.
Mind the gap: Examining the self-improvement capabilities of large language models
Song, Y., Zhang, H., Eisenach, C., Kakade, S., Foster, D., and Ghai, U · 2024
Later among the works it cites.
Easy-to-hard generalization: Scalable alignment beyond human supervision
Sun, Z., Yu, L., Shen, Y., Liu, W., Yang, Y., Welleck, S., and Gan, C · 2024
Later among the works it cites.
Understanding warmup-stable-decay learning rates: A river valley loss landscape perspective
Wen, K., Li, Z., Wang, J., Hall, D., Liang, P., and Ma, T · 2024
Later among the works it cites.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Later among the works it cites.
Transcendence: Generative models can outperform the experts that train them
Zhang, E., Zhu, V., Saphra, N., Kleiman, A., Edelman, B. L., Tambe, M., Kakade, S. M., and Malach, E · 2024
Later among the works it cites.
Transformers can achieve length generalization but not robustly
Zhou, Y., Alon, U., Chen, X., Wang, X., Agarwal, R., and Zhou, D · 2024
Later among the works it cites.