Fetching the paper…
Reading the bibliography…
We study the in-context learning dynamics of large language models (LLMs) using three instrumental learning tasks adapted from cognitive psychology.
The magical number seven, plus or minus two: Some limits on our capacity for processing information
Miller, G. A · 1956
Earlier work this paper cites.
A theory of Pavlovian conditioning: The effectiveness of reinforcement and non-reinforcement
Rescorla, R. and Wagner, A · 1972
Earlier work this paper cites.
Estimating the dimension of a model
Schwarz, G · 1978
Earlier work this paper cites.
Function Optimization Using Connectionist Reinforcement Learning Algorithms
Williams, R. and Peng, J · 1991
Earlier work this paper cites.
Confirmation bias: A ubiquitous phenomenon in many guises
Nickerson, R. S · 1998
Earlier work this paper cites.
Learning the value of information in an uncertain world
Behrens, T. E. J., Woolrich, M. W., Walton, M. E., and Rushworth, M. F. S · 2007
Earlier work this paper cites.
An approximately bayesian delta-rule model explains the dynamics of belief updating in a changing environment
Nassar, M. R., Wilson, R. C., Heasly, B., and Gold, J. I · 2010
Earlier work this paper cites.
Model-based influences on humans’ choices and striatal prediction errors
Daw, N. D., Gershman, S. J., Seymour, B., Dayan, P., and Dolan, R. J · 2011
Earlier work this paper cites.
Bayesian learning theory applied to human cognition
Jacobs, R. A. and Kruschke, J. K · 2011
Earlier work this paper cites.
The optimism bias
Sharot, T · 2011
Earlier work this paper cites.
Adaptive properties of differential learning rates for positive and negative outcomes
Cazé, R. D. and van der Meer, M. A · 2013
Earlier work this paper cites.
Taking Stock of Unrealistic Optimism
Shepperd, J. A., Klein, W. M. P., Waters, E. A., and Weinstein, N. D · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
A unifying probabilistic view of associative learning
Gershman, S. J · 2015
Earlier work this paper cites.
Rl 2 ^ \hat{2} : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Behavioural and neural characterization of optimistic reinforcement learning
Lefebvre, G., Lebreton, M., Meyniel, F., Bourgeois-Gironde, S., and Palminteri, S · 2017
Earlier work this paper cites.
Confirmation bias in human reinforcement learning: Evidence from counterfactual feedback processing
Palminteri, S., Lefebvre, G., Kilford, E. J., and Blakemore, S.-J · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Learning to reinforcement learn, January 2017
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2017
Cited alongside, same era.
Reinforcement Learning, Fast and Slow
Botvinick, M., Ritter, S., Wang, J., Kurth-Nelson, Z., Blundell, C., and Hassabis, D · 2019
Cited alongside, same era.
Meta-learning of sequential strategies
Ortega, P. A., Wang, J. X., Rowland, M., Genewein, T., Kurth-Nelson, Z., Pascanu, R., Heess, N., Veness, J., Pritzel, A., Sprechmann, P., et al · 2019
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Later among the works it cites.
A Normative Account of Confirmation Bias During Reinforcement Learning
Lefebvre, G., Summerfield, C., and Bogacz, R · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Later among the works it cites.
The computational roots of positivity and confirmation biases in reinforcement learning
Palminteri, S. and Lebreton, M · 2022
Later among the works it cites.
In-context learning dynamics with random binary sequences
Bigelow, E. J., Lubana, E. S., Dick, R. P., Tanaka, H., and Ullman, T. D · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Structured, uncertainty-driven exploration in real-world consumer choice
Schulz, E., Bhui, R., Love, B. C., Brier, B., Todd, M. T., and Gershman, S. J · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Information about action outcomes differentially affects learning from self-determined versus imposed choices
Chambon, V., Théro, H., Vidal, M., Vandendriessche, H., Haggard, P., and Palminteri, S · 2020
Cited alongside, same era.
Meta-trained agents implement bayes-optimal agents
Mikulik, V., Delétang, G., McGrath, T., Genewein, T., Martic, M., Legg, S., and Ortega, P · 2020
Cited alongside, same era.
Explainable machine learning for scientific insights and discoveries
Roscher, R., Bohn, B., Duarte, M. F., and Garcke, J · 2020
Cited alongside, same era.
Simple use of bic to assess model selection uncertainty: An illustration using mediation and moderation models
Wu, H., Fai Cheung, S., and On Leung, S · 2020
Cited alongside, same era.
Using cognitive psychology to understand gpt-3
Binz, M. and Schulz, E · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Later among the works it cites.
Meta-in-context learning in large language models
Coda-Forno, J., Binz, M., Akata, Z., Botvinick, M., Wang, J. X., and Schulz, E · 2023
Later among the works it cites.
Gpts are gpts: An early look at the labor market impact potential of large language models
Eloundou, T., Manning, S., Mishkin, P., and Rock, D · 2023
Later among the works it cites.
Zero-shot compositional reinforcement learning in humans, July 2023
Jagadish, A. K., Binz, M., Saanum, T., Wang, J. X., and Schulz, E · 2023
Later among the works it cites.
Chatgpt for good? on opportunities and challenges of large language models for education
Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., et al · 2023
Later among the works it cites.
Large language models are state-of-the-art evaluators of translation quality
Kocmi, T. and Federmann, C · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Larger language models do in-context learning differently
Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al · 2023
Later among the works it cites.
Impaired adaptation of learning to contingency volatility in internalizing psychopathology
Gagne, C., Zika, O., Dayan, P., and Bishop, S. J · 2050
Closest in time.
Ten simple rules for the computational modeling of behavioral data
Wilson, R. C. and Collins, A. G · 2050
Closest in time.