Fetching the paper…
Reading the bibliography…
We investigate the logical reasoning capabilities of large language models (LLMs) and their scalability in complex non-monotonic reasoning.
The complexity of completing partial latin squares
Colbourn, C. J · 1984
Earlier work this paper cites.
Hybrid algorithms for the constraint satisfaction problem
Prosser, P · 1993
Earlier work this paper cites.
Completing quasigroups or latin squares: A structured graph coloring problem
Gomes, C. P. and Shmoys, D. B · 2002
Earlier work this paper cites.
Constraint Processing
Dechter, R · 2003
Earlier work this paper cites.
Z3: An efficient smt solver
de Moura, L. M. and Bjørner, N. S · 2008
Earlier work this paper cites.
Automatic solutions of logic puzzles
Sempolinski, P · 2009
Earlier work this paper cites.
Learning to automatically solve logic grid puzzles
Mitra, A. and Baral, C · 2015
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Transformers as soft reasoners over language
Clark, P., Tafjord, O., and Richardson, K · 2020
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N. M., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., García, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Díaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K. S., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Cited alongside, same era.
Pushing the limits of rule reasoning in transformers through natural language satisfiability
Richardson, K. and Sabharwal, A · 2022
Cited alongside, same era.
Can transformers reason in fragments of natural language?
Schlegel, V., Pavlov, K. V., and Pratt-Hartmann, I · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E. H., Xia, F., Le, Q., and Zhou, D · 2022
Skywork-reward: Bag of tricks for reward modeling in llms
Liu, C. Y., Zeng, L., Liu, J., Yan, R., He, J., Wang, C., Yan, S., Liu, Y., and Zhou, Y · 2024
Later among the works it cites.
Natural language satisfiability: Exploring the problem distribution and evaluating transformer-based language models
Madusanka, T., Pratt-Hartmann, I., and Batista-Navarro, R · 2024
Later among the works it cites.
OpenAI · 2024
Later among the works it cites.
Can transformers reason logically? a study in sat solving
Pan, L., Ganesh, V., Abernethy, J., Esposo, C., and Lee, W · 2024
Later among the works it cites.
Logicbench: Towards systematic evaluation of logical reasoning ability of large language models
Parmar, M., Patel, N., Varshney, N., Nakamura, M., Luo, M., Mashetty, S., Mitra, A., and Baral, C · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jian, L., Lin, B. Y., West, P., Bhagavatula, C., Le Bras, R., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., and Choi, Y · 2023
Cited alongside, same era.
Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding
Liu, H., Liu, J., Cui, L., Teng, Z., Duan, N., Zhou, M., and Zhang, Y · 2023
Cited alongside, same era.
Identifying the limits of transformers when performing model-checking with natural language
Madusanka, T., Batista-navarro, R., and Pratt-hartmann, I · 2023
Cited alongside, same era.
AI@Meta · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I · 2024
Cited alongside, same era.
A closer look at logical reasoning with llms: The choice of tool matters
Lam, L. H. M., Thatikonda, R. K., and Shareghi, E · 2024
Cited alongside, same era.
Tülu 3: Pushing frontiers in open language model post-training
Lambert, N., Morrison, J. D., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, X., Gu, Y., Malik, S., Graf, V., Hwang, J. D., Yang, J., Le Bras, R., Tafjord, O., Wilhelm, C., Soldaini, L., Smith, N. A., Wang, Y., Dasigi, P., and Hajishirzi, H
Cited in the paper.
Divide and translate: Compositional first-order logic translation and verification for complex logical reasoning
Ryu, H., Kim, G., Lee, H. S., and Yang, E · 2024
Later among the works it cites.
Step-by-step reasoning to solve grid puzzles: Where do llms falter?
Tyagi, N., Parmar, M., Kulkarni, M., Rrv, A., Patel, N., Nakamura, M., Mitra, A., and Baral, C · 2024
Later among the works it cites.
On memorization of large language models in logical reasoning
Xie, C., Huang, Y., Zhang, C., Yu, D., Chen, X., Lin, B. Y., Li, B., Ghazi, B., and Kumar, R · 2024
Later among the works it cites.
Do large language models understand logic or just mimick context?
Yan, J., Wang, C., Huang, J., and Zhang, W · 2024
Later among the works it cites.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
DeepSeek-AI · 2025
Closest in time.