Fetching the paper…
Reading the bibliography…
We introduce a comprehensive Linguistic Benchmark designed to evaluate the limitations of Large Language Models (LLMs) in domains such as logical reasoning, spatial intelligence, and linguistic understanding, among others.
Collins, Harry M. “Embedded or embodied? A review of Hubert Dreyfus’ what computers still can’t do.” Artificial Intelligence 80.1: 99-118
1996
Earlier work this paper cites.
Asher, Nicholas, et al. “Limits for Learning with Language Models,” arXiv preprint arXiv:2306.12213
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Davis, E., et al. “Mathematics, word problems, common sense, and artificial intelligence,” Bulletin of the American Mathematical Society
2024
Earlier work this paper cites.
Miyu Sasaki, Natsumi Watanabe, Tsukihito Komanaka et al. “Enhancing Contextual Understanding of Mistral LLM with External Knowledge Bases,” Research Square rs.3.rs-4215447/v1
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Wachter, S., Mittelstadt, B., et al. “Do large language models have a legal duty to tell the truth?,” Available at SSRN 4771884
2024
Cited alongside, same era.
Li, Zhiming, et al. “LLMs for Relational Reasoning: How Far are We?” arXiv preprint arXiv:2401.09042
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Closest in time.