Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have accomplished remarkable reasoning performance in various domains.
Estimates of the regression coefficient based on kendall’s tau
P. K. Sen · 1968
Earlier work this paper cites.
The effect of premise order in conditional reasoning: A test of the mental model theory
V. Girotto, A. Mazzocco, and A. Tasso · 1997
Earlier work this paper cites.
Preferred premise order in propositional reasoning: Semantic informativeness and co-reference
M. Dekeyser, W. Schroyens, W. Schaeken, O. Spitaels, and G. d’Ydewalle · 2000
Earlier work this paper cites.
Deepmath-deep sequence models for premise selection
G. Irving, C. Szegedy, A. A. Alemi, N. Eén, F. Chollet, and J. Urban · 2016
Earlier work this paper cites.
Premise selection for theorem proving by deep graph embedding
M. Wang, Y. Tang, J. Wang, and J. Deng · 2017
Earlier work this paper cites.
Kendall tau sequence distance: Extending kendall tau from ranks to sequences
V. A. Cicirello · 2019
Earlier work this paper cites.
Premise selection in natural language mathematical texts
D. Ferreira and A. Freitas · 2020
Earlier work this paper cites.
K. Sinha, P. Parthasarathi, J. Pineau, and A. Williams · 2020
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Word order does matter and shuffled language models know it
M. Abdou, V. Ravishankar, A. Kulmizev, and A. Søgaard · 2022
Earlier work this paper cites.
Folio: Natural language reasoning with first-order logic
S. Han, H. Schoelkopf, Y. Zhao, Z. Qi, M. Riddell, L. Benson, L. Sun, E. Zubova, Y. Qiao, M. Burtell, et al · 2022
Cited alongside, same era.
Capturing failures of large language models via human cognitive biases
E. Jones and J. Steinhardt · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al · 2022
Cited alongside, same era.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
A. Saparov and H. He · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
R. T. McCoy, S. Yao, D. Friedman, M. Hardy, and T. L. Griffiths · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Testing the general deductive reasoning capacity of large language models using ood examples
A. Saparov, R. Y. Pang, V. Padmakumar, N. Joshi, S. M. Kazemi, N. Kim, and H. He · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
F. Shi, X. Chen, K. Misra, N. Scales, D. Dohan, E. H. Chi, N. Schärli, and D. Zhou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the paradox of learning to reason from data
H. Zhang, L. H. Li, T. Meng, K.-W. Chang, and G. V. d. Broeck · 2022
Cited alongside, same era.
The reversal curse: Llms trained on" a is b" fail to learn" b is a"
L. Berglund, M. Tong, M. Kaufmann, M. Balesni, A. C. Stickland, T. Korbak, and O. Evans · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Cited alongside, same era.
Unnatural error correction: Gpt-4 can almost perfectly handle unnatural scrambled text
Q. Cao, T. Kojima, Y. Matsuo, and Y. Iwasawa · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini · 2023
Cited alongside, same era.
Google · 2023
Cited alongside, same era.
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt
T. Hagendorff, S. Fabi, and M. Kosinski · 2023
Cited alongside, same era.
P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, Q. Liu, T. Liu, and Z. Sui · 2023
Later among the works it cites.
F. Xu, Q. Lin, J. Han, T. Zhao, J. Liu, and E. Cambria · 2023
Later among the works it cites.
Concise and organized perception facilitates large language models for deductive reasoning
S. Yan, C. Shen, J. Liu, and J. Ye · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
H. Zhou, A. Bradley, E. Littwin, N. Razin, O. Saremi, J. Susskind, S. Bengio, and P. Nakkiran · 2023
Later among the works it cites.
Large language models can learn rules
Z. Zhu, Y. Xue, X. Chen, D. Zhou, J. Tang, D. Schuurmans, and H. Dai · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2024
Closest in time.
A & b== b & a: Triggering logical reasoning failures in large language models
Y. Wan, W. Wang, Y. Yang, Y. Yuan, J.-t. Huang, P. He, W. Jiao, and M. R. Lyu · 2024
Closest in time.
Transformers can achieve length generalization but not robustly
Y. Zhou, U. Alon, X. Chen, X. Wang, R. Agarwal, and D. Zhou · 2024
Closest in time.