Fetching the paper…
Reading the bibliography…
Recent studies show that instruction tuning (IT) and reinforcement learning from human feedback (RLHF) improve the abilities of large language models (LMs) dramatically.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Hila Gonen and Yoav Goldberg. 2019 · 1903
Earlier work this paper cites.
The probable error of a mean
Student. 1908 · 1908
Earlier work this paper cites.
The utility analysis of choices involving risk
Milton Friedman and Leonard J Savage. 1948 · 1948
Earlier work this paper cites.
Conditional logit analysis of qualitative choice behavior
D McFadden. 1974 · 1974
Earlier work this paper cites.
Prospect theory: An analysis of decisions under risk
Daniel Kahneman. 1979 · 1979
Earlier work this paper cites.
Adding asymmetrically dominated alternatives: Violations of regularity and the similarity hypothesis
Joel Huber, John W Payne, and Christopher Puto. 1982 · 1982
Earlier work this paper cites.
On the conflict between logic and belief in syllogistic reasoning
JSBT Evans, Julie L Barston, and Paul Pollard. 1983 · 1983
Earlier work this paper cites.
Methods for evaluating changes in health care policy: the difference-in-differences approach
Justin B Dimick and Andrew M Ryan. 2014 · 2014
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Do neural language models overcome reporting bias?
Vered Shwartz and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Cognitive biases and decision-making strategies in times of change: a systematic literature review
Chiara Acciarini, Federica Brunetta, and Paolo Boccardelli. 2021 · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Surface form competition: Why the highest probability answer isn’t always right
Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021 · 2021
Machine intuition: Uncovering human-like intuitive decision-making in gpt-3.5
Thilo Hagendorff, Sarah Fabi, and Michal Kosinski. 2022 · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022 · 2022
Later among the works it cites.
Aristotle’s Logic
Robin Smith. 2022 · 2022
Later among the works it cites.
Fewer errors, but more stereotypes? the effect of model size on gender bias
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Cited alongside, same era.
The impact of cognitive biases on professionals’ decision-making: A review of four occupational areas
Vincent Berthet. 2022 · 2022
Cited alongside, same era.
Using cognitive psychology to understand gpt-3
Marcel Binz and Eric Schulz. 2022 · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022 · 2022
Cited alongside, same era.
Language models show human-like content effects on reasoning
Ishita Dasgupta, Andrew K Lampinen, Stephanie CY Chan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. 2022 · 2022
Cited alongside, same era.
Yarden Tal, Inbal Magar, and Roy Schwartz. 2022 · 2022
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Closest in time.
The unlocking spell on base llms: Rethinking alignment via in-context learning
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2023 · 2023
Closest in time.
Gpt-4 technical report
OpenAI. 2023 · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al. 2023 · 2023
Closest in time.
Are large language models rational investors?
Yuhang Zhou, Yuchen Ni, Xiang Liu, Jian Zhang, Sen Liu, Guangnan Ye, and Hongfeng Chai. 2024 · 2024
Closest in time.