Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated notable capabilities across various tasks, showcasing complex problem-solving abilities.
Can recursive neural tensor networks learn logical reasoning?
S. R. Bowman · 2013
Earlier work this paper cites.
Deep learning for symbolic mathematics
G. Lample and F. Charton · 2019
Earlier work this paper cites.
Clutrr: A diagnostic benchmark for inductive reasoning from text
K. Sinha, S. Sodhani, J. Dong, J. Pineau, and W. L. Hamilton · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Transformers as soft reasoners over language
P. Clark, O. Tafjord, and K. Richardson · 2020
Earlier work this paper cites.
Isarstep: a benchmark for high-level mathematical reasoning
W. Li, L. Yu, Y. Wu, and L. C. Paulson · 2020
Earlier work this paper cites.
Reclor: A reading comprehension dataset requiring logical reasoning
W. Yu, Z. Jiang, Y. Dong, and J. Feng · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
M. Geva, D. Khashabi, E. Segal, T. Khot, D. Roth, and J. Berant · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Creak: A dataset for commonsense reasoning over entity knowledge
Y. Onoe, M. J. Zhang, E. Choi, and G. Durrett · 2021
Earlier work this paper cites.
Folio: Natural language reasoning with first-order logic
S. Han, H. Schoelkopf, Y. Zhao, Z. Qi, M. Riddell, L. Benson, L. Sun, E. Zubova, Y. Qiao, M. Burtell, et al · 2022
Earlier work this paper cites.
Lila: A unified benchmark for mathematical reasoning
S. Mishra, M. Finlayson, P. Lu, L. Tang, S. Welleck, C. Baral, T. Rajpurohit, O. Tafjord, A. Sabharwal, P. Clark, et al · 2022
Earlier work this paper cites.
Numglue: A suite of fundamental yet challenging mathematical reasoning tasks
S. Mishra, A. Mitra, N. Varshney, B. Sachdeva, P. Clark, C. Baral, and A. Kalyan · 2022
Cited alongside, same era.
Introducing chatgpt, 2022
OpenAI · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
A. Saparov and H. He · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models
Y. He, Y. Wu, Y. Jia, R. Mihalcea, Y. Chen, and N. Deng · 2023
Later among the works it cites.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Later among the works it cites.
Agentbench: Evaluating llms as agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, et al · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of gpt-4
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, and A. Awadallah · 2023
Later among the works it cites.
Gpt-4 technical report
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Glm-130b: An open bilingual pre-trained model
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia, et al · 2022
Cited alongside, same era.
Introducing claude, 2023
Anthropic · 2023
Cited alongside, same era.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Cited alongside, same era.
Black-box prompt optimization: Aligning large language models without model training
J. Cheng, X. Liu, K. Zheng, P. Ke, H. Wang, Y. Dong, J. Tang, and M. Huang · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Cited alongside, same era.
R. OpenAI · 2023
Later among the works it cites.
Cognitive architectures for language agents
T. R. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths · 2023
Later among the works it cites.
Internlm: A multilingual language model with progressively enhanced capabilities
I. Team · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
Autodetect: Towards a unified framework for automated weakness detection in large language models
J. Cheng, Y. Lu, X. Gu, P. Ke, X. Liu, Y. Dong, H. Wang, J. Tang, and M. Huang · 2024
Closest in time.
Language is primarily a tool for communication rather than thought
E. Fedorenko, S. T. Piantadosi, and E. A. Gibson · 2024
Closest in time.
Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
T. GLM, :, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Rojas, G. Feng, H. Zhao, H. Lai, H. Yu, H. Wang, J. Sun, J. Zhang, J. Cheng, J. Gui, J. Tang, J. Zhang, J. Li, L. Zhao, L. Wu, L. Zhong, M. Liu, M. Huang, P. Zhang, Q. Zheng, R. Lu, S. Duan, S. Zhang, S. Cao, S. Yang, W. L. Tam, W. Zhao, X. Liu, X. Xia, X. Zhang, X. Gu, X. Lv, X. Liu, X. Liu, X. Yang, X. Song, X. Zhang, Y. An, Y. Xu, Y. Niu, Y. Yang, Y. Li, Y. Bai, Y. Dong, Z. Qi, Z. Wang, Z. Yang, Z. Du, Z. Hou, and Z. Wang · 2024
Closest in time.
Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent ai
Z. Huang, Z. Wang, S. Xia, X. Li, H. Zou, R. Xu, R.-Z. Fan, L. Ye, E. Chern, Y. Ye, et al · 2024
Closest in time.