Fetching the paper…
Reading the bibliography…
Error prediction in large language models often relies on domain-specific information.
Introspective Perception: Learning to Predict Failures in Vision Systems
Daftry, S.; Zeng, S.; Bagnell, J. A.; and Hebert, M. 2016 · 2016
Earlier work this paper cites.
Annotating Derivations: A New Evaluation Strategy and Dataset for Algebra Word Problems
Upadhyay, S.; and Chang, M.-W. 2017 · 2017
Earlier work this paper cites.
Hierarchical Neural Story Generation
Fan, A.; Lewis, M.; and Dauphin, Y. 2018 · 2018
Earlier work this paper cites.
Failing to Learn: Autonomously Identifying Perception Failures for Self-driving Cars
Ramanagopal, M. S.; Anderson, C.; Vasudevan, R.; and Johnson-Roberson, M. 2018 · 2018
Earlier work this paper cites.
Complex Sequential Question Answering: Towards Learning to Converse over Linked Question Answer Pairs with a Knowledge Graph
Saha, A.; Pahuja, V.; Khapra, M. M.; Sankaranarayanan, K.; and Chandar, S. 2018 · 2018
Earlier work this paper cites.
The Curious Case of Neural Text Degeneration
Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; and Choi, Y. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Earlier work this paper cites.
Meta-Learning in Neural Networks: A Survey
Hospedales, T.; Antoniou, A.; Micaelli, P.; and Storkey, A. 2022 · 2022
Earlier work this paper cites.
Lang2LTL: Translating Natural Language Commands to Temporal Specification with Large Language Models
Liu, J. X.; Yang, Z.; Schornstein, B.; Liang, S.; Idrees, I.; Tellex, S.; and Shah, A. 2022 · 2022
Earlier work this paper cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2022 · 2022
Earlier work this paper cites.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2022 · 2022
Cited alongside, same era.
Domain Generalization: A Survey
Zhou, K.; Liu, Z.; Qiao, Y.; Xiang, T.; and Loy, C. C. 2022 · 2022
Cited alongside, same era.
PyReason: Software for Open World Temporal Logic
Aditya, D.; Mukherji, K.; Balasubramanian, S.; Chaudhary, A.; and Shakarian, P. 2023 · 2023
Cited alongside, same era.
Do Language Models Know When They’re Hallucinating References?
Agrawal, A.; Mackey, L.; and Kalai, A. T. 2023 · 2023
Cited alongside, same era.
Active Prompting with Chain-of-Thought for Large Language Models
Diao, S.; Wang, P.; Lin, Y.; and Zhang, T. 2023 · 2023
Cited alongside, same era.
Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
Liu, H.; Ning, R.; Teng, Z.; Liu, J.; Zhou, Q.; and Zhang, Y. 2023 · 2023
Closest in time.
GPT-4 Technical Report
OpenAI. 2023 · 2023
Closest in time.
Learning gain differences between ChatGPT and human tutor generated algebra hints
Pardos, Z. A.; and Bhandari, S. 2023 · 2023
Closest in time.
ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope
Ray, P. P. 2023 · 2023
Closest in time.
An Independent Evaluation of ChatGPT on Mathematical Word Problems (MWP)
Shakarian, P.; Koyyalamudi, A.; Ngu, N.; and Mareedu, L. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Controlling Linguistic Style Aspects in Neural Language Generation
Ficler, J.; and Goldberg, Y. 2023 · 2023
Cited alongside, same era.
Mathematical Capabilities of ChatGPT
Frieder, S.; Pinchetti, L.; Griffiths, R.-R.; Salvatori, T.; Lukasiewicz, T.; Petersen, P. C.; Chevalier, A.; and Berner, J. 2023 · 2023
Cited alongside, same era.
ChatGPT: Jack of all trades, master of none
Kocoń, J.; Cichecki, I.; Kaszyca, O.; Kochanek, M.; Szydło, D.; Baran, J.; Bielaniewicz, J.; Gruza, M.; Janz, A.; Kanclerz, K.; Kocoń, A.; Koptyra, B.; Mieleszczenko-Kowszewicz, W.; Miłkowski, P.; Oleksy, M.; Piasecki, M.; RadliÅski, Å.; Wojtasik, K.; Woźniak, S.; and Kazienko, P. 2023 · 2023
Cited alongside, same era.
Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
Kuhn, L.; Gal, Y.; and Farquhar, S. 2023 · 2023
Cited alongside, same era.
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
Lin, Z.; Trivedi, S.; and Sun, J. 2023 · 2023
Cited alongside, same era.
Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models
Vaithilingam, P.; Zhang, T.; and Glassman, E. L. ????
Cited in the paper.
Tian, K.; Mitchell, E.; Zhou, A.; Sharma, A.; Rafailov, R.; Yao, H.; Finn, C.; and Manning, C. D. 2023 · 2023
Closest in time.
ChatLog: Recording and Analyzing ChatGPT Across Time
Tu, S.; Li, C.; Yu, J.; Wang, X.; Hou, L.; and Li, J. 2023 · 2023
Closest in time.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; and Narasimhan, K. 2023 · 2023
Closest in time.
How well do Large Language Models perform in Arithmetic tasks?
Yuan, Z.; Yuan, H.; Tan, C.; Wang, W.; and Huang, S. 2023 · 2023
Closest in time.
Navigating the Grey Area: Expressions of Overconfidence and Uncertainty in Language Models
Zhou, K.; Jurafsky, D.; and Hashimoto, T. 2023 · 2023
Closest in time.