Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown significant general language understanding abilities.
Newell, A., Simon, H.: The logic theory machine–a complex information processing system. IRE Transactions on Information Theory (1956)
1956
Earlier work this paper cites.
McCarthy, J., Hayes, P.J.: Some philosophical problems from the standpoint of artificial intelligence. In: Machine Intelligence 4 (1969)
1969
Earlier work this paper cites.
Cresswell, M.J.: Logics and languages (1st ed.). Routledge (1973)
1973
Earlier work this paper cites.
Kowalski, R.: Logic for problem solving. Ediciones Díaz de Santos (1979)
1979
Earlier work this paper cites.
Poole, D., Goebel, R., Aleliunas, R.: Theorist: A Logical Reasoning System for Defaults and Diagnosis, pp. 331–352 (1987)
1987
Earlier work this paper cites.
Iwańska, L.: Logical reasoning in natural language: It is all about knowledge. Minds and Machines pp. 475–510 (1993)
1993
Earlier work this paper cites.
Pulman, S.G.: Using the framework (1996)
1996
Earlier work this paper cites.
Mccarthy, J.: Programs with common sense (2002)
2002
Earlier work this paper cites.
Dagan, I., Glickman, O., Magnini, B.: The pascal recognising textual entailment challenge. In: MLCW (2005)
2005
Earlier work this paper cites.
MacCartney, B., Manning, C.D.: Natural logic for textual inference. In: Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing. pp. 193–200 (2007)
2007
Earlier work this paper cites.
MacCartney, B., Manning, C.D.: Natural logic for textual inference. In: Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing (2007)
2007
Earlier work this paper cites.
Bowman, S.R., Angeli, G., Potts, C., Manning, C.D.: A large annotated corpus for learning natural language inference. In: Proc. of EMNLP. pp. 632–642 (2015)
2015
Earlier work this paper cites.
Chen, D., Bolton, J., Manning, C.D.: A thorough examination of the CNN/daily mail reading comprehension task. In: ACL (2016)
2016
Earlier work this paper cites.
Lai, G., Xie, Q., Liu, H., Yang, Y., Hovy, E.: RACE: Large-scale Reading Comprehension dataset from Examinations. In: EMNLP (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. CoRR (2017)
2017
Earlier work this paper cites.
Demszky, D., Guu, K., Liang, P.: Transforming question answering datasets into natural language inference datasets (2018)
2018
Earlier work this paper cites.
Poliak, A., Haldar, A., Rudinger, R., Hu, J.E., Pavlick, E., White, A.S., Van Durme, B.: Collecting diverse natural language inference problems for sentence representation evaluation. In: Proc. of EMNLP (2018)
2018
Earlier work this paper cites.
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S.: GLUE: A multi-task benchmark and analysis platform for natural language understanding. In: Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP (2018)
2018
Earlier work this paper cites.
Williams, A., Nangia, N., Bowman, S.: A broad-coverage challenge corpus for sentence understanding through inference. In: Proc. of AACL (2018)
2018
Earlier work this paper cites.
Williams, A., Nangia, N., Bowman, S.: A broad-coverage challenge corpus for sentence understanding through inference. In: Proc. of NAACL (2018)
2018
Earlier work this paper cites.
Clark, C., Lee, K., Chang, M.W., Kwiatkowski, T., Collins, M., Toutanova, K.: Boolq: Exploring the surprising difficulty of natural yes/no questions (2019)
2019
Earlier work this paper cites.
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., Gardner, M.: DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In: Proc. of AACL (2019)
2019
Earlier work this paper cites.
Li, T., Srikumar, V.: Augmenting neural networks with first-order logic. In: Proc. of ACL. pp. 292–302 (2019)
2019
Earlier work this paper cites.
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V.: Roberta: A robustly optimized bert pretraining approach. arXiv (2019)
2019
Earlier work this paper cites.
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S.R.: SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems (2019)
2019
Earlier work this paper cites.
Yanaka, H., Mineshima, K., Bekki, D., Inui, K., Sekine, S., , Bos, J.: Help: A dataset for identifying shortcomings of neural models in monotonicity reasoning. In: Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM2019) (2019)
2019
Earlier work this paper cites.
Clark, P., Tafjord, O., Richardson, K.: Transformers as soft reasoners over language. In: Proc. of IJCAI (2020)
2020
Earlier work this paper cites.
Joshi, P., Aditya, S., Sathe, A., Choudhury, M.: Taxinli: Taking a ride up the NLU hill. CoRR (2020)
2020
Cited alongside, same era.
Liu, H., Cui, L., Liu, J., Zhang, Y.: Natural language inference in context - investigating contextual reasoning over long texts. CoRR (2020)
2020
Cited alongside, same era.
Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., Zhang, Y.: Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. CoRR (2020)
2020
Cited alongside, same era.
Saha, S., Ghosh, S., Srivastava, S., Bansal, M.: Prover: Proof generation for interpretable reasoning over rules (2020)
2020
Cited alongside, same era.
Yu, W., Jiang, Z., Dong, Y., Feng, J.: Reclor: A reading comprehension dataset requiring logical reasoning. In: Proc. of ICLR (2020)
2020
Cited alongside, same era.
2023
Closest in time.
Choi, J.H., Hickman, K.E., Monahan, A., Schwarcz, D.: Chatgpt goes to law school. Available at SSRN (2023)
2023
Closest in time.
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., Li, L., Sui, Z.: A survey on in-context learning (2023)
2023
Closest in time.
Frieder, S., Pinchetti, L., Griffiths, R.R., Salvatori, T., Lukasiewicz, T., Petersen, P.C., Chevalier, A., Berner, J.: Mathematical capabilities of chatgpt (2023)
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, M., Tworek, J., Jun, H., Yuan, Q., Others: Evaluating large language models trained on code (2021)
2021
Cited alongside, same era.
Chen, M., Tworek, J., Jun, H., Yuan, Q., Others: Evaluating large language models trained on code (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., Steinhardt, J.: Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR) (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Koreeda, Y., Manning, C.: ContractNLI: A dataset for document-level natural language inference for contracts. In: Proc. of EMNLP Findings (2021)
2021
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
Huang, J., Chang, K.C.C.: Towards reasoning in large language models: A survey (2023)
2023
Closest in time.
Kung, T.H., Cheatham, M., Medenilla, A., Sillos, C., De Leon, L., Elepaño, C., Madriaga, M., Aggabao, R., Diaz-Candido, G., Maningo, J., et al.: Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models. PLoS digital health p. e0000198 (2023)
2023
Closest in time.
Liu, H., Liu, J., Cui, L., Teng, Z., Duan, N., Zhou, M., Zhang, Y.: Logiqa 2.0—an improved dataset for logical reasoning in natural language understanding. IEEE/ACM Transactions on Audio, Speech, and Language Processing pp. 2947–2962 (2023)
2023
Closest in time.
Liu, H., Teng, Z., Cui, L., Zhang, C., Zhou, Q., Zhang, Y.: Logicot: Logical chain-of-thought instruction tuning. In: Proc. of EMNLP Findings. pp. 2908–2921 (2023)
2023
Closest in time.
OpenAI: Gpt-4 technical report (2023)
2023
Closest in time.
Orrù, G., Piarulli, A., Conversano, C., Gemignani, A.: Human-like problem-solving abilities in large language models using chatgpt. Frontiers in Artificial Intelligence p. 1199350 (2023)
2023
Closest in time.
2023
Closest in time.
Qin, C., Zhang, A., Zhang, Z., Chen, J., Yasunaga, M., Yang, D.: Is chatgpt a general-purpose natural language processing task solver? (2023)
2023
Closest in time.
2023
Closest in time.
Sawada, T., Paleka, D., Havrilla, A., Tadepalli, P., Vidas, P., Kranias, A., Nay, J.J., Gupta, K., Komatsuzaki, A.: Arb: Advanced reasoning benchmark for large language models (2023)
2023
Closest in time.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., Lample, G.: Llama: Open and efficient foundation language models (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2024
Closest in time.
Jiang, A.Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Others.: Mixtral of experts (2024)
2024
Closest in time.
OpenAI: Openai o1 system card (2024), https://arxiv.org/abs/2412.16720
2024
Closest in time.
Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., et al.: Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In: Proc. of NeurIPS (2024)
2024
Closest in time.
2025
Closest in time.
2025
Closest in time.
Team, Q.: Qwq-32b: Embracing the power of reinforcement learning (2025)
2025
Closest in time.