Fetching the paper…
Reading the bibliography…
In-context learning enables large language models (LLMs) to perform a variety of tasks, including learning to make reward-maximizing choices in simple bandit tasks.
Effects of blocked versus interleaved training on relative value learning
William M Hayes and Douglas H Wedell · 1907
Earlier work this paper cites.
A theory of pavlovian conditioning: Variations in the effectiveness of reinforcement and non-reinforcement
Robert A Rescorla · 1972
Earlier work this paper cites.
Contextual modulation of value signals in reward and punishment learning
Stefano Palminteri, Mehdi Khamassi, Mateus Joffily, and Giorgio Coricelli · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Learning relative values in the striatum induces violations of normative decision making
Tilmann A Klein, Markus Ullsperger, and Gerhard Jocham · 2017
Earlier work this paper cites.
Behavioural and neural characterization of optimistic reinforcement learning
Germain Lefebvre, Maël Lebreton, Florent Meyniel, Sacha Bourgeois-Gironde, and Stefano Palminteri · 2017
Earlier work this paper cites.
Confirmation bias in human reinforcement learning: Evidence from counterfactual feedback processing
Stefano Palminteri, Germain Lefebvre, Emma J Kilford, and Sarah-Jayne Blakemore · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Reference-point centering and range-adaptation enhance human reinforcement learning at the cost of irrational preferences
Sophie Bavard, Maël Lebreton, Mehdi Khamassi, Giorgio Coricelli, and Stefano Palminteri · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Two sides of the same coin: Beneficial and detrimental consequences of range adaptation in human reinforcement learning
Sophie Bavard, Aldo Rustichini, and Stefano Palminteri · 2021
Earlier work this paper cites.
Regret in experience-based decisions: The effects of expected value differences and mixed gains and losses
William M Hayes and Douglas H Wedell · 2021
Earlier work this paper cites.
Context-dependent outcome encoding in human reinforcement learning
Stefano Palminteri and Maël Lebreton · 2021
Earlier work this paper cites.
Reinforcement learning in and out of context: The effects of attentional focus
William M Hayes and Douglas H Wedell · 2022
Earlier work this paper cites.
Human value learning and representation reflect rational adaptation to task demands
Keno Juechems, Tugba Altun, Rita Hira, and Andreas Jarvstad · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Cited alongside, same era.
Using large language models to simulate multiple humans and replicate human subject studies
Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai · 2023
Cited alongside, same era.
The functional form of value normalization in human reinforcement learning
Sophie Bavard and Stefano Palminteri · 2023
Cited alongside, same era.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Contextual influence of reinforcement learning performance of depression: evidence for a negativity bias?
Henri Vandendriessche, Amel Demmou, Sophie Bavard, Julien Yadak, Cédric Lemogne, Thomas Mauras, and Stefano Palminteri · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Using cognitive psychology to understand gpt-3
Marcel Binz and Eric Schulz · 2023
Cited alongside, same era.
The emergence of economic rationality of gpt
Yiting Chen, Tracy Xiao Liu, You Shan, and Songfa Zhong · 2023
Cited alongside, same era.
Inducing anxiety in large language models increases exploration and bias
Julian Coda-Forno, Kristin Witte, Akshay K Jagadish, Marcel Binz, Zeynep Akata, and Eric Schulz · 2023
Cited alongside, same era.
Thilo Hagendorff · 2023
Cited alongside, same era.
Large language models as simulated economic agents: What can we learn from homo silicus?
John J Horton · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2023
Cited alongside, same era.
Intrinsic rewards explain context-sensitive valuation in reinforcement learning
Gaia Molinaro and Anne GE Collins · 2023
Cited alongside, same era.
Nicolas Yax, Hernan Anlló, and Stefano Palminteri · 2023
Later among the works it cites.
Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods
Yuji Cao, Huan Zhao, Yuheng Cheng, Ting Shu, Guolong Liu, Gaoqi Liang, Junhua Zhao, and Yun Li · 2024
Closest in time.
Cogbench: a large language model walks into a psychology lab
Julian Coda-Forno, Marcel Binz, Jane X Wang, and Eric Schulz · 2024
Closest in time.
The transformative power of generative ai and large language models, Apr 2024
Gary Fowler · 2024
Closest in time.
Relative value biases in large language models
William M Hayes, Nicolas Yax, and Stefano Palminteri · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
How quickly do large language models learn unexpected skills?, Feb 2024
Stephen Ornes · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
In-context learning agents are asymmetric belief updaters
Johannes A Schubert, Akshay K Jagadish, Marcel Binz, and Eric Schulz · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team and Google DeepMind · 2024
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.