Fetching the paper…
Reading the bibliography…
Large language models (LLMs) exhibit remarkable task generalization, solving tasks they were never explicitly trained on with only a few demonstrations.
Covariate shift adaptation by importance weighted cross validation
Sugiyama, M., Krauledat, M., and Müller, K.-R · 2007
Earlier work this paper cites.
Domain adaptation: Learning bounds and algorithms
Mansour, Y., Mohri, M., and Rostamizadeh, A · 2009
Earlier work this paper cites.
Learning bounds for importance weighting
Cortes, C., Mansour, Y., and Mohri, M · 2010
Earlier work this paper cites.
Domain adaptation–can quantity compensate for quality?
Ben-David, S. and Urner, R · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2017
Earlier work this paper cites.
Hypothesis transfer learning via transformation functions
Du, S. S., Koushik, J., Singh, A., and Póczos, B · 2017
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Lake, B. and Baroni, M · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Measuring compositional generalization: A comprehensive method on realistic data
Keysers, D., Schärli, N., Scales, N., Buisman, H., Furrer, D., Kashubin, S., Momchev, N., Sinopalnikov, D., Stafiniak, L., Tihon, T., Tsarkov, D., Wang, X., van Zee, M., and Bousquet, O · 2020
Earlier work this paper cites.
Understanding and mitigating the tradeoff between robustness and accuracy
Raghunathan, A., Xie, S. M., Yang, F., Duchi, J., and Liang, P · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 2020
Cited alongside, same era.
Marginal singularity and the benefits of labels in covariate-shift
Kpotufe, S. and Martinet, G · 2021
Cited alongside, same era.
Near-optimal linear regression under distribution shift
Lei, Q., Hu, W., and Lee, J · 2021
Cited alongside, same era.
Towards a theoretical framework of out-of-distribution generalization
Ye, H., Xie, C., Cai, T., Li, R., Li, Z., and Wang, L · 2021
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Optimally tackling covariate shift in rkhs-based nonparametric regression
Ma, C., Pathak, R., and Wainwright, M. J · 2023
Later among the works it cites.
Testing the general deductive reasoning capacity of large language models using ood examples
Saparov, A., Pang, R. Y., Padmakumar, V., Joshi, N., Kazemi, M., Kim, N., and He, H · 2023
Later among the works it cites.
How far can transformers reason? the locality barrier and inductive scratchpad
Abbe, E., Bengio, S., Lotfi, A., Sandon, C., and Saremi, O · 2024
Later among the works it cites.
Understanding in-context learning in transformers and LLMs by learning to learn discrete functions
Bhattamishra, S., Patel, A., Blunsom, P., and Kanade, V · 2024
Later among the works it cites.
Discovering modular solutions that generalize compositionally
Schug, S., Kobayashi, S., Akram, Y., Wolczyk, M., Proca, A. M., Von Oswald, J., Pascanu, R., Sacramento, J., and Steger, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Domain generalization: A survey
Zhou, K., Liu, Z., Qiao, Y., Xiang, T., and Loy, C. C · 2022
Cited alongside, same era.
How do in-context examples affect compositional generalization?
An, S., Lin, Z., Fu, Q., Chen, B., Zheng, N., Lou, J.-G., and Zhang, D · 2023
Cited alongside, same era.
A theory for emergence of complex skills in language models
Arora, S. and Goyal, A · 2023
Cited alongside, same era.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Bai, Y., Chen, F., Wang, H., Xiong, C., and Mei, S · 2023
Cited alongside, same era.
Distributionally robust losses for latent covariate mixtures
Duchi, J., Hashimoto, T., and Namkoong, H · 2023
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., Welleck, S., West, P., Bhagavatula, C., Le Bras, R., et al · 2023
Cited alongside, same era.
Compositional generalization from first principles
Wiedemer, T., Mayilvahanan, P., Bethge, M., and Brendel, W · 2024
Later among the works it cites.
Do large language models have compositional ability? an investigation into limitations and scalability
Xu, Z., Shi, Z., and Liang, Y · 2024
Later among the works it cites.
Can models learn skill composition from examples?
Zhao, H., Kaur, S., Yu, D., Goyal, A., and Arora, S · 2024
Later among the works it cites.
Instruct-skillmix: A powerful pipeline for llm instruction tuning
Kaur, S., Park, S., Goyal, A., and Arora, S · 2025
Closest in time.
When does compositional structure yield compositional generalization? a kernel theory
Lippl, S. and Stachenfeld, K · 2025
Closest in time.
Out-of-distribution generalization via composition: a lens through induction heads in transformers
Song, J., Xu, Z., and Zhong, Y · 2025
Closest in time.
From sparse dependence to sparse attention: Unveiling how chain-of-thought enhances transformer sample efficiency
Wen, K., Zhang, H., Lin, H., and Zhang, J · 2025
Closest in time.