Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly deployed in everyday applications, demanding robust general reasoning capabilities and diverse reasoning skillset.
Defaults and normality in causal structures
J. Y. Halpern · 2008
Earlier work this paper cites.
Graded causation and defaults
J. Y. Halpern and C. Hitchcock · 2015
Earlier work this paper cites.
Causal superseding
J. F. Kominsky, J. Phillips, T. Gerstenberg, D. Lagnado, and J. Knobe · 2015
Earlier work this paper cites.
Unifying morality’s influence on non-moral judgments: The relevance of alternative possibilities
J. Phillips, J. B. Luguri, and J. Knobe · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
J. Weston, A. Bordes, S. Chopra, A. M. Rush, B. Van Merriënboer, A. Joulin, and T. Mikolov · 2015
Earlier work this paper cites.
Actual Causality
J. Y. Halpern · 2016
Earlier work this paper cites.
Normality and actual causal strength
T. F. Icard, J. F. Kominsky, and J. Knobe · 2017
Earlier work this paper cites.
A large self-annotated corpus for sarcasm
M. Khodak, N. Saunshi, and K. Vodrahalli · 2017
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang · 2018
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Proofwriter: Generating implications, proofs, and abductive statements over natural language
O. Tafjord, B. D. Mishra, and P. Clark · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
A comprehensive survey of the actual causality literature
K. R. Kueffner · 2021
Earlier work this paper cites.
Spartqa:: A textual question answering benchmark for spatial reasoning
R. Mirzaee, H. R. Faghihi, Q. Ning, and P. Kordjmashidi · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Earlier work this paper cites.
Time-aware language models as temporal knowledge bases
B. Dhingra, J. R. Cole, J. M. Eisenschlos, D. Gillick, J. Eisenstein, and W. W. Cohen · 2022
Earlier work this paper cites.
J. Hessel, A. Marasović, J. D. Hwang, L. Lee, J. Da, R. Zellers, R. Mankoff, and Y. Choi · 2022
Earlier work this paper cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
A. Saparov and H. He · 2022
Cited alongside, same era.
Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts
Z. Shi, Q. Zhang, and A. Lipani · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Cited alongside, same era.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024
E. Glazer, E. Erdil, T. Besiroglu, D. Chicharro, E. Chen, A. Gunning, C. F. Olsson, J.-S. Denain, A. Ho, E. de Oliveira Santos, O. Järviniemi, M. Barnett, R. Sandler, M. Vrzala, J. Sevilla, Q. Ren, E. Pratt, L. Levine, G. Barkley, N. Stewart, B. Grechuk, T. Grechuk, S. V. Enugandla, and M. Wildon · 2024
Later among the works it cites.
Not all llm reasoners are created equal
A. Hosseini, A. Sordoni, D. Toyama, A. Courville, and R. Agarwal · 2024
Later among the works it cites.
Remi: A dataset for reasoning with multiple images
M. Kazemi, N. Dikkala, A. Anand, P. Devic, I. Dasgupta, F. Liu, B. Fatemi, P. Awasthi, D. Guo, S. Gollapudi, et al · 2024
Later among the works it cites.
Scheherazade: Evaluating chain-of-thought math reasoning in llms with chain-of-problems
S. Miner, Y. Takashima, S. Han, F. Erata, T. Antonopoulos, R. Piskac, and S. J. Shapiro · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. Abrego, J. Ahn, J. Austin, P. Barham, J. Botha, J. Bradbury, S. Brahma, K. Brooks, M. Catasta, Y. Cheng, C. Cherry, C. A. Choquette-Choo, A. Chowdhery, C. Crepy, S. Dave, M. Dehghani, S. Dev, J. Devlin, M. Díaz, N. Du, E. Dyer, V. Feinberg, F. Feng, V. Fienber, M. Freitag, X. Garcia, S. Gehrmann, L. Gonzalez, G. Gur-Ari, S. Hand, H. Hashemi, L. Hou, J. Howland, A. Hu, J. Hui, J. Hurwitz, M. Isard, A. Ittycheriah, M. Jagielski, W. Jia, K. Kenealy, M. Krikun, S. Kudugunta, C. Lan, K. Lee, B. Lee, E. Li, M. Li, W. Li, Y. Li, J. Li, H. Lim, H. Lin, Z. Liu, F. Liu, M. Maggioni, A. Mahendru, J. Maynez, V. Misra, M. Moussalem, Z. Nado, J. Nham, E. Ni, A. Nystrom, A. Parrish, M. Pellat, M. Polacek, A. Polozov, R. Pope, S. Qiao, E. Reif, B. Richter, P. Riley, A. C. Ros, A. Roy, B. Saeta, R. Samuel, R. Shelby, A. Slone, D. Smilkov, D. R. So, D. Sohn, S. Tokumine, D. Valter, V. Vasudevan, K. Vodrahalli, X. Wang, P. Wang, Z. Wang, T. Wang, J. Wieting, Y. Wu, K. Xu, Y. Xu, L. Xue, P. Yin, J. Yu, Q. Zhang, S. Zheng, C. Zheng, W. Zhou, D. Zhou, S. Petrov, and Y. Wu · 2023
Cited alongside, same era.
Causal reasoning and large language models: Opening a new frontier for causality
E. Kıcıman, R. Ness, A. Sharma, and C. Tan · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Cited alongside, same era.
Inverse scaling: When bigger isn’t better
I. R. McKenzie, A. Lyzhov, M. Pieler, A. Parrish, A. Mueller, A. Prabhu, E. McLean, A. Kirtland, A. Ross, A. Liu, et al · 2023
Cited alongside, same era.
MoCa: Measuring human-language model alignment on causal and moral judgment tasks
A. Nie, Y. Zhang, A. S. Amdekar, C. Piech, T. B. Hashimoto, and T. Gerstenberg · 2023
Cited alongside, same era.
CRAB: Assessing the strength of causal relationships between real-world events
A. Romanou, S. Montariol, D. Paul, L. Laugier, K. Aberer, and A. Bosselut · 2023
Cited alongside, same era.
Testing the general deductive reasoning capacity of large language models using ood examples
A. Saparov, R. Y. Pang, V. Padmakumar, N. Joshi, M. Kazemi, N. Kim, and H. He · 2023
Cited alongside, same era.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar · 2024
Later among the works it cites.
Logicbench: Towards systematic evaluation of logical reasoning ability of large language models
M. Parmar, N. Patel, N. Varshney, M. Nakamura, M. Luo, S. Mashetty, A. Mitra, and C. Baral · 2024
Later among the works it cites.
Linguini: A benchmark for language-agnostic linguistic reasoning
E. Sánchez, B. Alastruey, C. Ropers, P. Stenetorp, M. Artetxe, and M. R. Costa-jussà · 2024
Later among the works it cites.
Transformers struggle to learn to search
A. Saparov, S. Pawar, S. Pimpalgaonkar, N. Joshi, R. Y. Pang, V. Padmakumar, S. M. Kazemi, N. Kim, and H. He · 2024
Later among the works it cites.
Causal language modeling can elicit search and reasoning capabilities on logic puzzles
K. Shah, N. Dikkala, X. Wang, and R. Panigrahy · 2024
Later among the works it cites.
LLMs cannot find reasoning errors, but can correct them given the error location
G. Tyen, H. Mansoor, V. Carbune, P. Chen, and T. Mak · 2024
Later among the works it cites.
Michelangelo: Long context evaluations beyond haystacks via latent structure queries
K. Vodrahalli, S. Ontanon, N. Tripuraneni, K. Xu, S. Jain, R. Shivanna, J. Hui, N. Dikkala, M. Kazemi, B. Fatemi, et al · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al · 2024
Later among the works it cites.
Livebench: A challenging, contamination-free llm benchmark
C. White, S. Dooley, M. Roberts, A. Pal, B. Feuer, S. Jain, R. Shwartz-Ziv, N. Jain, K. Saifullah, S. Naidu, et al · 2024
Later among the works it cites.
Sportqa: A benchmark for sports understanding in large language models
H. Xia, Z. Yang, Y. Wang, R. Tracy, Y. Zhao, D. Huang, Z. Chen, Y. Zhu, Y.-f. Wang, and W. Shen · 2024
Later among the works it cites.
Large language models can learn temporal reasoning
S. Xiong, A. Payani, R. Kompella, and F. Fekri · 2024
Later among the works it cites.
Humor in ai: Massive scale crowd-sourced preferences and benchmarks for cartoon captioning, 2024
J. Zhang, L. Jain, Y. Guo, J. Chen, K. L. Zhou, S. Suresh, A. Wagenmaker, S. Sievert, T. Rogers, K. Jamieson, R. Mankoff, and R. Nowak · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
L. Phan, A. Gatti, Z. Han, and N. L. et. al · 2025
Closest in time.
G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, et al · 2025
Closest in time.