Fetching the paper…
Reading the bibliography…
Language models (LMs) have shown impressive performance on tasks within their training distribution, but often struggle with structurally novel tasks even when given a small number of in-context task examples.
On the measure of intelligence, 2019
Chollet, F · 1911
Earlier work this paper cites.
Local learning algorithms
Bottou, L. and Vapnik, V · 1992
Earlier work this paper cites.
Transductive inference for text classification using support vector machines
Joachims, T · 1999
Earlier work this paper cites.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H · 2017
Earlier work this paper cites.
Fixing weight decay regularization in Adam, 2018
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Test-time training with self-supervision for generalization under distribution shifts
Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., and Hardt, M · 2020
Earlier work this paper cites.
Communicating natural programs to humans and machines
Acquaviva, S., Pu, Y., Kryven, M., Sechopoulos, T., Wong, C., Ecanow, G. E., Nye, M. I., Tessler, M. H., and Tenenbaum, J · 2022
Earlier work this paper cites.
Test-time training with masked autoencoders
Gandelsman, Y., Sun, Y., Chen, X., and Efros, A. A · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
MetaICL: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
What learning algorithm is in-context learning? Investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2023
Earlier work this paper cites.
QLoRA: Efficient finetuning of quantized LLMs
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2023
Cited alongside, same era.
Challenging BIG-Bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., and Wei, J · 2023
Cited alongside, same era.
Solving ARC visual analogies with neural embeddings and vector arithmetic: A generalized method
Veldkamp, K., Rosenbusch, H., Thoms, L., and Stevenson, C · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2023
Cited alongside, same era.
Tree of Thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
Do large language models solve ARC visual analogies like people do?, 2024
Opiełka, G., Rosenbusch, H., Vijverberg, V., and Stevenson, C. E · 2024
Closest in time.
Learning to (learn at test time): RNNs with expressive hidden states, 2024
Sun, Y., Li, X., Dalal, K., Xu, J., Vikram, A., Zhang, G., Dubois, Y., Chen, X., Wang, X., Koyejo, S., Hashimoto, T., and Guestrin, C · 2024
Closest in time.
Function vectors in large language models
Todd, E., Li, M., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2024
Closest in time.
torchtune: PyTorch’s finetuning library, 2024
torchtune Maintainers and Contributors · 2024
Closest in time.
Hypothesis Search: Inductive reasoning with language models
Wang, R., Zelikman, E., Poesia, G., Pu, Y., Haber, N., and Goodman, N · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large language monkeys: Scaling inference compute with repeated sampling, 2024
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
CodeIt: Self-improving language models with prioritized hindsight replay
Butt, N., Manczak, B., Wiggers, A., Rainone, C., Zhang, D. W., Defferrard, M., and Cohen, T · 2024
Cited alongside, same era.
Getting 50% (SoTA) on ARC-AGI with GPT-4o, 2024
Greenblatt, R · 2024
Cited alongside, same era.
Test-time training on nearest neighbors for large language models
Hardt, M. and Sun, Y · 2024
Cited alongside, same era.
Addressing the Abstraction and Reasoning Corpus via procedural example generation, 2024
Hodel, M · 2024
Cited alongside, same era.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Cited alongside, same era.
H-ARC: A robust estimate of human performance on the Abstraction and Reasoning Corpus benchmark
LeGris, S., Vong, W. K., Lake, B. M., and Gureckis, T. M · 2024
Cited alongside, same era.
Reasoning or reciting? Exploring the capabilities and limitations of language models through counterfactual tasks
Wu, Z., Qiu, L., Ross, A., Akyürek, E., Chen, B., Wang, B., Kim, N., Andreas, J., and Kim, Y · 2024
Closest in time.
Probing the decision boundaries of in-context learning in large language models
Zhao, S., Nguyen, T., and Grover, A · 2024
Closest in time.
Titans: Learning to memorize at test time, 2025
Behrouz, A., Zhong, P., and Mirrokni, V · 2025
Closest in time.
ARC Prize 2024: Technical report, 2025
Chollet, F., Knoop, M., Kamradt, G., and Landers, B · 2025
Closest in time.
Learning how hard to think: Input-adaptive allocation of LM computation
Damani, M., Shenfeld, I., Peng, A., Bobu, A., and Andreas, J · 2025
Closest in time.
Efficiently learning at test-time: Active fine-tuning of LLMs
Hübotter, J., Bongni, S., Hakimi, I., and Krause, A · 2025
Closest in time.
Combining induction and transduction for abstract reasoning
Li, W.-D., Hu, K., Larsen, C., Wu, Y., Alford, S., Woo, C., Dunn, S. M., Tang, H., Zheng, W.-L., Pu, Y., and Ellis, K · 2025
Closest in time.
Scaling test-time compute optimally can be more effective than scaling LLM parameters
Snell, C. V., Lee, J., Xu, K., and Kumar, A · 2025
Closest in time.
Test-time regression: a unifying framework for designing sequence models with associative memory
Wang, K. A., Shi, J., and Fox, E. B · 2025
Closest in time.
Neural networks for abstraction and reasoning
Bober-Irizar, M. and Banerjee, S · 2045
Closest in time.