Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown impressive emergent abilities in a wide range of tasks, but the associated expensive API cost greatly limits the real application.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
Representativeness revisited: Attribute substitution in intuitive judgment
Kahneman, D.; Frederick, S.; et al. 2002 · 2002
Earlier work this paper cites.
The heuristic-analytic theory of reasoning: Extension and evaluation
Evans, J. S. B. 2006 · 2006
Earlier work this paper cites.
Clinical cognition and diagnostic error: applications of a dual process model of reasoning
Croskerry, P. 2009 · 2009
Earlier work this paper cites.
Intuition and reasoning: A dual-process perspective
Evans, J. S. B. 2010 · 2010
Earlier work this paper cites.
Thinking, fast and slow
Kahneman, D. 2011 · 2011
Earlier work this paper cites.
Rationality and the reflective mind
Stanovich, K. 2011 · 2011
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A. M.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2022 · 2022
Earlier work this paper cites.
Compression of generative pre-trained language models via quantization
Tao, C.; Hou, L.; Zhang, W.; Shang, L.; Jiang, X.; Liu, Q.; Luo, P.; and Wong, N. 2022 · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2022 · 2022
Cited alongside, same era.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Zhou, D.; Schärli, N.; Hou, L.; Wei, J.; Scales, N.; Wang, X.; Schuurmans, D.; Cui, C.; Bousquet, O.; Le, Q. V.; et al. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Anil, R.; Dai, A. M.; Firat, O.; Johnson, M.; Lepikhin, D.; Passos, A.; Shakeri, S.; Taropa, E.; Bailey, P.; Chen, Z.; et al. 2023 · 2023
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Self-refine: Iterative refinement with self-feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. 2023 · 2023
Later among the works it cites.
Does Writing with Language Models Reduce Content Diversity?
Padmakumar, V.; and He, H. 2023 · 2023
Later among the works it cites.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Šakota, M.; Peyrard, M.; and West, R. 2023 · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K. R.; and Yao, S. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chan, C.-M.; Chen, W.; Su, Y.; Yu, J.; Xue, W.; Zhang, S.; Fu, J.; and Liu, Z. 2023 · 2023
Cited alongside, same era.
AutoAgents: A Framework for Automatic Agent Generation
Chen, G.; Dong, S.; Shu, Y.; Zhang, G.; Sesay, J.; Karlsson, B. F.; Fu, J.; and Shi, Y. 2023 · 2023
Cited alongside, same era.
Reconcile: Round-table conference improves reasoning via consensus among diverse llms
Chen, J. C.-Y.; Saha, S.; and Bansal, M. 2023 · 2023
Cited alongside, same era.
FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Chen, L.; Zaharia, M.; and Zou, J. 2023 · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2023 · 2023
Cited alongside, same era.
Improving Factuality and Reasoning in Language Models through Multiagent Debate
Du, Y.; Li, S.; Torralba, A.; Tenenbaum, J. B.; and Mordatch, I. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
Understanding the Effects of RLHF on LLM Generalisation and Diversity
Kirk, R.; Mediratta, I.; Nalmpantis, C.; Luketina, J.; Hambro, E.; Grefenstette, E.; and Raileanu, R. 2023 · 2023
Cited alongside, same era.
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; and Narasimhan, K. 2023 · 2023
Later among the works it cites.
Exchange-of-thought: Enhancing large language model capabilities through cross-model communication
Yin, Z.; Sun, Q.; Chang, C.; Guo, Q.; Dai, J.; Huang, X.-J.; and Qiu, X. 2023 · 2023
Later among the works it cites.
Large language model cascades with mixture of thoughts representations for cost-efficient reasoning
Yue, M.; Zhao, J.; Zhang, M.; Du, L.; and Yao, Z. 2023 · 2023
Later among the works it cites.
Cumulative reasoning with large language models
Zhang, Y.; Yang, J.; Yuan, Y.; and Yao, A. C.-C. 2023 · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023 · 2023
Later among the works it cites.
Yi: Open foundation models by 01. ai
Young, A.; Chen, B.; Li, C.; Huang, C.; Zhang, G.; Zhang, G.; Li, H.; Zhu, J.; Chen, J.; Chang, J.; et al. 2024 · 2024
Closest in time.