Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown impressive capabilities across tasks such as mathematics, coding, and reasoning, yet their learning ability, which is crucial for adapting to dynamic environments and acquiring new knowledge, remains underexplored.
The Canterbury Puzzles and Other Curious Problems
Dudeney, H. E · 1907
Earlier work this paper cites.
International Law: A Treatise, Volume 1
Oppenheim, L · 1920
Earlier work this paper cites.
International Law: A Treatise, Volume 2
Oppenheim, L · 1921
Earlier work this paper cites.
Steps toward artificial intelligence
Minsky, M · 1961
Earlier work this paper cites.
The modularity of mind
Fodor, J. A · 1983
Earlier work this paper cites.
Experiential learning: Experience as the source of learning and development , volume 1
Kolb, D. A · 1984
Earlier work this paper cites.
Fantastic Book of Logic Puzzles
Mandell, M · 1986
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D. X., and Steinhardt, J · 2009
Earlier work this paper cites.
Experiential learning and learning environments: The case of active listening skills
Huerta-Wong, J. E. and Schoech, R · 2010
Earlier work this paper cites.
Putting students on the path to learning: The case for fully guided instruction
Clark, R. E., Kirschner, P. A., and Sweller, J · 2012
Earlier work this paper cites.
Rapid instructed task learning: A new window into the human brain’s unique capacity for flexible cognitive control
Cole, M. W., Laurent, P., and Stocco, A · 2012
Earlier work this paper cites.
The 2014 international planning competition: Progress and trends
Vallati, M., Chrpa, L., Grześ, M., McCluskey, T. L., Roberts, M., Sanner, S., et al · 2015
Earlier work this paper cites.
Mawps: A math word problem repository
Koncel-Kedziorski, R., Roy, S., Amini, A., Kushman, N., and Hajishirzi, H · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W., Yogatama, D., Dyer, C., and Blunsom, P · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Rapid instruction-based task learning (ritl) in schizophrenia
Sheffield, J. M., Ruge, H., Kandala, S., and Barch, D. M · 2018
Earlier work this paper cites.
Babyai: A platform to study the sample efficiency of grounded language learning
Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Are nlp models really able to solve simple math word problems?
Patel, A., Bhattamishra, S., and Goyal, N · 2021
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Shridhar, M., Yuan, X., Côté, M., Bisk, Y., Trischler, A., and Hausknecht, M. J · 2021
Earlier work this paper cites.
Explicit instruction and executive functioning capacity: A new direction in cognitive load theory
Siregar, N. R · 2021
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, e · 2022
Cited alongside, same era.
A survey on in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Liu, T., et al · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Cited alongside, same era.
Numglue: A suite of fundamental yet challenging mathematical reasoning tasks
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2023
Later among the works it cites.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2023
Later among the works it cites.
Abdin, M., Aneja, J., Behl, H., Bubeck, S., Eldan, R., Gunasekar, S., Harrison, M., Hewett, R. J., Javaheripi, M., Kauffmann, P., et al · 2024
Later among the works it cites.
Agarwal, R., Singh, A., Zhang, L. M., Bohnet, B., Chan, S., Zhang, B., Anand, A., Abbas, Z., Nova, A., Co-Reyes, J. D., Chu, E., Behbahani, F. M. P., Faust, A., and Larochelle, H · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mishra, S., Mitra, A., Varshney, N., Sachdeva, B., Clark, P., Baral, C., and Kalyan, A · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Cited alongside, same era.
ScienceWorld: Is your agent smarter than a 5th grader?
Wang, R., Jansen, P., Côté, M.-A., and Ammanabrolu, P · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
An alternative to cognitivism: Computational phenomenology for deep learning
Beckmann, P., Köstner, G., and Hipólito, I · 2023
Cited alongside, same era.
Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Muyan, Z., Zhang, Q., Zhu, X., Lu, L., Li, B., Luo, P., Lu, T., Qiao, Y., and Dai, J · 2023
Cited alongside, same era.
Llm-assisted content analysis: Using large language models to support deductive coding
Chew, R., Bollenbacher, J., Wenger, M., Speer, J., and Kim, A · 2023
Cited alongside, same era.
Learning agent-based modeling with llm companions: Experiences of novices and experts using chatgpt & netlogo chat
Chen, J., Lu, X., Du, Y., Rejtig, M., Bagley, R., Horn, M., and Wilensky, U · 2024
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, C. X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P. S., and Hashimoto, T. B · 2024
Later among the works it cites.
Self-evolving gpt: A lifelong autonomous experiential learner
Gao, J., Ding, X., Cui, Y., Zhao, J., Wang, H., Liu, T., and Qin, B · 2024
Later among the works it cites.
Using an llm to help with code understanding
Nam, D., Macvean, A., Hellendoorn, V., Vasilescu, B., and Myers, B · 2024
Later among the works it cites.
Democratizing large language models via personalized parameter-efficient fine-tuning
Tan, Z., Zeng, Q., Tian, Y., Liu, Z., Yin, B., and Jiang, M · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Multitask-based evaluation of open-source llm on software vulnerability
Yin, X., Ni, C., and Wang, S · 2024
Later among the works it cites.
Mathvc: An llm-simulated multi-character virtual classroom for mathematics education
Yue, M., Lyu, W., Mifdal, W., Suh, J., Zhang, Y., and Yao, Z · 2024
Later among the works it cites.
Expel: Llm agents are experiential learners
Zhao, A., Huang, D., Xu, Q., Lin, M., Liu, Y.-J., and Huang, G · 2024
Later among the works it cites.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Zheng, Y., Zhang, R., Zhang, J., Ye, Y., Luo, Z., Feng, Z., and Ma, Y · 2024
Later among the works it cites.
Current and future use of large language models for knowledge work
Brachman, M., El-Ashry, A., Dugan, C., and Geyer, W · 2025
Closest in time.
Dfpe: A diverse fingerprint ensemble for enhancing llm performance
Cohen, S., Goldshlager, N., Cohen-Inger, N., Shapira, B., and Rokach, L · 2025
Closest in time.
Ai language model rivals expert ethicist in perceived moral expertise
Dillion, D., Mondal, D., Tandon, N., and Gray, K · 2025
Closest in time.
Guertler, L., Cheng, B., Yu, S., Liu, B., Choshen, L., and Tan, C · 2025
Closest in time.
Explaining length bias in llm-based preference evaluations
Hu, Z., Song, L., Zhang, J., Xiao, Z., Chen, Z., and Xiong, H · 2025
Closest in time.
Towards accurate differential diagnosis with large language models
McDuff, D., Schaekermann, M., Tu, T., Palepu, A., Wang, A., Garrison, J., Singhal, K., Sharma, Y., Azizi, S., Kulkarni, K., et al · 2025
Closest in time.
Performance evaluation of large language models: A comprehensive review
Meva, D. and Kukadiya, H · 2025
Closest in time.
Maintaining long-distance relationships with (mediocre) llm-based chatbots: A collaborative ethnographic study
Ploderer, B., Capel, T., Davaakhuu, N.-E., Tung, N. H., Maichal, D. M., Kuzhiparambil, A. K., Mai, Q. H., and Reitberger, W · 2025
Closest in time.
Toward expert-level medical question answering with large language models
Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Amin, M., Hou, L., Clark, K., Pfohl, S. R., Cole-Lewis, H., et al · 2025
Closest in time.