Fetching the paper…
Reading the bibliography…
Trained on vast corpora of human language, language models demonstrate emergent human-like reasoning abilities.
Human problem solving , volume 104
A. Newell, H. A. Simon, et al · 1972
Earlier work this paper cites.
Development of expertise in mathematical problem solving
J. Sweller, R. F. Mawer, and M. R. Ward · 1983
Earlier work this paper cites.
Abstract planning and perceptual chunks: Elements of expertise in geometry
K. R. Koedinger and J. R. Anderson · 1990
Earlier work this paper cites.
How people learn to skip steps
S. Blessing and J. R. Anderson · 1996
Earlier work this paper cites.
Cognitive load theory and complex learning: Recent developments and future directions
J. J. Van Merrienboer and J. Sweller · 2005
Earlier work this paper cites.
Sympy: symbolic computing in python
A. Meurer, C. P. Smith, M. Paprocki, O. Čertík, S. B. Kirpichev, M. Rocklin, A. Kumar, S. Ivanov, J. K. Moore, S. Singh, T. Rathnayake, S. Vig, B. E. Granger, R. P. Muller, F. Bonazzi, H. Gupta, S. Vats, F. Johansson, F. Pedregosa, M. J. Curry, A. R. Terrel, v. Roučka, A. Saboo, I. Fernando, S. Kulal, R. Cimrman, and A. Scopatz · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Why machine reading comprehension models learn shortcuts?
Y. Lai, C. Zhang, Y. Feng, Q. Huang, and D. Zhao · 2021
Earlier work this paper cites.
Exploring generalization ability of pretrained language models on arithmetic and logical reasoning
C. Wang, B. Zheng, Y. Niu, and Y. Zhang · 2021
Earlier work this paper cites.
Using cognitive psychology to understand GPT-3
M. Binz and E. Schulz · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
S. R. Bowman, J. Hyun, E. Perez, E. Chen, C. Pettit, S. Heiner, K. Lukosiute, A. Askell, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Olah, D. Amodei, D. Amodei, D. Drain, D. Li, E. Tran-Johnson, J. Kernion, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, L. Lovitt, N. Elhage, N. Schiefer, N. Joseph, N. Mercado, N. DasSarma, R. Larson, S. McCandlish, S. Kundu, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Telleen-Lawton, T. Brown, T. Henighan, T. Hume, Y. Bai, Z. Hatfield-Dodds, B. Mann, and J. Kaplan · 2022
Earlier work this paper cites.
K. M. Collins, C. Wong, J. Feng, M. Wei, and J. B. Tenenbaum · 2022
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
A. Creswell, M. Shanahan, and I. Higgins · 2022
Earlier work this paper cites.
Language models show human-like content effects on reasoning
I. Dasgupta, A. K. Lampinen, S. C. Y. Chan, A. Creswell, D. Kumaran, J. L. McClelland, and F. Hill · 2022
Earlier work this paper cites.
223C11The Cognitive Unconscious and Dual Process Theories of Reasoning
W. De Neys · 2022
Cited alongside, same era.
Shortcut learning of large language models in natural language understanding: A survey
M. Du, F. He, N. Zou, D. Tao, and X. Hu · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou · 2022
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, O. Bousquet, Q. Le, and E. H. Chi · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. M. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang · 2023
Chameleon: Plug-and-play compositional reasoning with large language models
P. Lu, B. Peng, H. Cheng, M. Galley, K. Chang, Y. N. Wu, S. Zhu, and J. Gao · 2023
Later among the works it cites.
Towards A holistic landscape of situated theory of mind in large language models
Z. Ma, J. Sansom, R. Peng, and J. Chai · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Welleck, B. P. Majumder, S. Gupta, A. Yazdanbakhsh, and P. Clark · 2023
Later among the works it cites.
Levels of AGI: operationalizing progress on the path to AGI
M. R. Morris, J. Sohl-Dickstein, N. Fiedel, T. Warkentin, A. Dafoe, A. Faust, C. Farabet, and S. Legg · 2023
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, I. Sutskever, and J. Wu · 2023
Cited alongside, same era.
Implicit chain of thought reasoning via knowledge distillation
Y. Deng, K. Prasad, R. Fernandez, P. Smolensky, V. Chaudhary, and S. M. Shieber · 2023
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, S. Welleck, P. West, C. Bhagavatula, R. L. Bras, J. D. Hwang, S. Sanyal, X. Ren, A. Ettinger, Z. Harchaoui, and Y. Choi · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu · 2023
Cited alongside, same era.
L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu · 2023
Cited alongside, same era.
Unveiling theory of mind in large language models: A parallel to single neurons in the human brain
M. Jamali, Z. M. Williams, and J. Cai · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. Bang, A. Madotto, and P. Fung · 2023
Cited alongside, same era.
A. Saparov and H. He · 2023
Later among the works it cites.
Positional description matters for transformers arithmetic
R. Shen, S. Bubeck, R. Eldan, Y. T. Lee, Y. Li, and Y. Zhang · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. Canton-Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
Survey on factuality in large language models: Knowledge, retrieval and domain-specificity
C. Wang, X. Liu, Y. Yue, X. Tang, T. Zhang, C. Jiayang, Y. Yao, W. Gao, X. Hu, Z. Qi, Y. Wang, L. Yang, J. Wang, X. Xie, Z. Zhang, and Y. Zhang · 2023
Later among the works it cites.
Cogbench: a large language model walks into a psychology lab
J. Coda-Forno, M. Binz, J. X. Wang, and E. Schulz · 2024
Closest in time.
From explicit cot to implicit cot: Learning to internalize cot step by step
Y. Deng, Y. Choi, and S. M. Shieber · 2024
Closest in time.
The unreasonable effectiveness of easy training data for hard tasks
P. Hase, M. Bansal, P. Clark, and S. Wiegreffe · 2024
Closest in time.
Beyond accuracy: Evaluating the reasoning behavior of large language models - A survey
P. Mondorf and B. Plank · 2024
Closest in time.
Easy-to-hard generalization: Scalable alignment beyond human supervision
Z. Sun, L. Yu, Y. Shen, W. Liu, Y. Yang, S. Welleck, and C. Gan · 2024
Closest in time.
Y. Yang, Y. Ma, and P. Liu · 2024
Closest in time.
Distilling system 2 into system 1
P. Yu, J. Xu, J. Weston, and I. Kulikov · 2024
Closest in time.