Fetching the paper…
Reading the bibliography…
Hallucination continues to be one of the most critical challenges in the institutional adoption journey of Large Language Models (LLMs).
Nogueira, R.; and Cho, K. 2020 · 1901
Earlier work this paper cites.
The Curious Case of Neural Text Degeneration
Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; and Choi, Y. 2020 · 1904
Earlier work this paper cites.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Wang, A.; Pruksachatkun, Y.; Nangia, N.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 1905
Earlier work this paper cites.
UNIFIEDQA: Crossing Format Boundaries with a Single QA System
Khashabi, D.; Min, S.; Khot, T.; Sabharwal, A.; Tafjord, O.; Clark, P.; and Hajishirzi, H. 2020 · 1907
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
The theory of the estimation of test reliability
Kuder, G.; and Richardson, M. 1937 · 1937
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E. 1948 · 1948
Earlier work this paper cites.
Coefficient alpha and the internal structure of tests
Cronbach, L. J. 1951 · 1951
Earlier work this paper cites.
The relation of the reliability of multiple-choice tests to the distribution of item difficulties
Lord, F. M. 1952 · 1952
Earlier work this paper cites.
A Coefficient of Agreement for Nominal Scales
Cohen, J. 1960 · 1960
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Fleiss, J. L. 1971 · 1971
Earlier work this paper cites.
Indices of Qualitative Variation and Political Measurement
Wilcox, A. R. 1973 · 1973
Earlier work this paper cites.
The Division of Labor: Conceptualization and Related Measures*
Gibbs, J. P.; and Poston, J., Dudley L. 1975 · 1975
Earlier work this paper cites.
Regression Models with Ordinal Variables
Winship, C.; and Mare, R. D. 1984 · 1984
Earlier work this paper cites.
Probabilistic Outputs for Support vector Machines and Comparisons to Regularized Likelihood Methods
Platt, J. 1999 · 1999
Earlier work this paper cites.
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
Wang, W.; Wei, F.; Dong, L.; Bao, H.; Yang, N.; and Zhou, M. 2020 · 2002
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Niculescu-Mizil, A.; and Caruana, R. 2005 · 2005
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Russell, S.; and Norvig, P. 2009 · 2009
Earlier work this paper cites.
Extracting Training Data from Large Language Models
Carlini, N.; Tramer, F.; Wallace, E.; Jagielski, M.; Herbert-Voss, A.; Lee, K.; Roberts, A.; Brown, T.; Song, D.; Erlingsson, U.; Oprea, A.; and Raffel, C. 2021 · 2012
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Murphy, K. P. 2012 · 2012
Earlier work this paper cites.
Improving deep neural networks using softplus units
Zheng, H.; Yang, Z.; Liu, W.; Liang, J.; and Li, Y. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
On Calibration of Modern Neural Networks
Guo, C.; Pleiss, G.; Sun, Y.; and Weinberger, K. Q. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing Multiple Choice Science Questions
Johannes Welbl, M. G., Nelson F. Liu. 2017 · 2017
Earlier work this paper cites.
triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Joshi, M.; Choi, E.; Weld, D.; and Zettlemoyer, L. 2017 · 2017
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2017 · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Hierarchical Neural Story Generation
Fan, A.; Lewis, M.; and Dauphin, Y. 2018 · 2018
Earlier work this paper cites.
Identifying Well-formed Natural Language Questions
Faruqui, M.; and Das, D. 2018 · 2018
Cited alongside, same era.
Learning to Write with Cooperative Discriminators
Holtzman, A.; Buys, J.; Forbes, M.; Bosselut, A.; Golub, D.; and Choi, Y. 2018 · 2018
Cited alongside, same era.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Mihaylov, T.; Clark, P.; Khot, T.; and Sabharwal, A. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Cited alongside, same era.
WikiQA: A Challenge Dataset for Open-Domain Question Answering
Finetuned Language Models Are Zero-Shot Learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022 · 2022
Later among the works it cites.
PromptChainer: Chaining Large Language Model Prompts through Visual Programming
Wu, T.; Jiang, E.; Donsbach, A.; Gray, J.; Molina, A.; Terry, M.; and Cai, C. J. 2022 · 2022
Later among the works it cites.
Selectively Answering Ambiguous Questions
Cole, J.; Zhang, M.; Gillick, D.; Eisenschlos, J.; Dhingra, B.; and Eisenstein, J. 2023 · 2023
Later among the works it cites.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Abbasi, B.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; Le Noac’h, A.; Li, H.; McDonell, K.; Muennighoff, N.; Ociepa, C.; Phang, J.; Reynolds, L.; Schoelkopf, H.; Skowron, A.; Sutawika, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; and Zou, A. 2023 · 2023
Later among the works it cites.
Introducing Gemini: our largest and most capable AI model
Google. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, Y.; Yih, W.-t.; and Meek, C. 2015 · 2018
Cited alongside, same era.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Cited alongside, same era.
MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Amini, A.; Gabriel, S.; Lin, S.; Koncel-Kedziorski, R.; Choi, Y.; and Hajishirzi, H. 2019 · 2019
Cited alongside, same era.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Billion-scale similarity search with GPUs
Johnson, J.; Douze, M.; and Jégou, H. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Cited alongside, same era.
PIQA: Reasoning about Physical Commonsense in Natural Language
Bisk, Y.; Zellers, R.; Bras, R. L.; Gao, J.; and Choi, Y. 2020 · 2020
Cited alongside, same era.
Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; and Liu, T. 2023 · 2023
Later among the works it cites.
Survey of Hallucination in Natural Language Generation
Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y. J.; Madotto, A.; and Fung, P. 2023 · 2023
Later among the works it cites.
Large Language Models are Zero-Shot Reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2023 · 2023
Later among the works it cites.
Making Large Language Models Better Reasoners with Step-Aware Verifier
Li, Y.; Lin, Z.; Zhang, S.; Fu, Q.; Chen, B.; Lou, J.-G.; and Chen, W. 2023 · 2023
Later among the works it cites.
Query Rewriting for Retrieval-Augmented Large Language Models
Ma, X.; Gong, Y.; He, P.; Zhao, H.; and Duan, N. 2023 · 2023
Later among the works it cites.
Self-Refine: Iterative Refinement with Self-Feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; Gupta, S.; Majumder, B. P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P. 2023 · 2023
Later among the works it cites.
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Manakul, P.; Liusie, A.; and Gales, M. J. F. 2023 · 2023
Later among the works it cites.
Your Everyday AI Companion: Microsoft Bing
Microsoft. 2023 · 2023
Later among the works it cites.
Peng, B.; Galley, M.; He, P.; Cheng, H.; Xie, Y.; Hu, Y.; Huang, Q.; Liden, L.; Yu, Z.; Chen, W.; and Gao, J. 2023 · 2023
Later among the works it cites.
On Early Detection of Hallucinations in Factual Question Answering
Snyder, B.; Moisescu, M.; and Zafar, M. B. 2023 · 2023
Later among the works it cites.
Varshney, N.; Yao, W.; Zhang, H.; Chen, J.; and Yu, D. 2023 · 2023
Later among the works it cites.
Calibration of Neural Networks
Vasilev, R.; and D’yakonov, A. 2023 · 2023
Later among the works it cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023 · 2023
Later among the works it cites.
HiddenTables and PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of Taxonomies
Watson, W.; Cho, N.; Balch, T.; and Veloso, M. 2023 · 2023
Later among the works it cites.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2023 · 2023
Later among the works it cites.
Generate rather than retrieve: Large language models are strong context generators
Yu, W.; Iter, D.; Wang, S.; Xu, Y.; Ju, M.; Sanyal, S.; Zhu, C.; Zeng, M.; and Jiang, M. 2023 · 2023
Later among the works it cites.
Noisy Exemplars Make Large Language Models More Robust: A Domain-Agnostic Behavioral Analysis
Zheng, H.; and Saparov, A. 2023 · 2023
Later among the works it cites.
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Alzahrani, N.; Alyahya, H. A.; Alnumay, Y.; Alrashed, S.; Alsubaie, S.; Almushaykeh, Y.; Mirza, F.; Alotaibi, N.; Altwairesh, N.; Alowisheq, A.; Bari, M. S.; and Khan, H. 2024 · 2024
Closest in time.
FISHNET: Financial Intelligence from Sub-querying, Harmonizing, Neural-Conditioning, Expert Swarms, and Task Planning
Cho, N.; Srishankar, N.; Cecchi, L.; and Watson, W. 2024 · 2024
Closest in time.
Hallucinations Leaderboard
Minervini, P.; Nie, P.; Fourrier, C.; Saxena, R.; Gema, A. P.; He, X.; et al. 2024 · 2024
Closest in time.
A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
Tonmoy, S. M. T. I.; Zaman, S. M. M.; Jain, V.; Rani, A.; Rawte, V.; Chadha, A.; and Das, A. 2024 · 2024
Closest in time.
Corrective Retrieval Augmented Generation
Yan, S.-Q.; Gu, J.-C.; Zhu, Y.; and Ling, Z.-H. 2024 · 2024
Closest in time.
FlowMind: Automatic Workflow Generation with LLMs
Zeng, Z.; Watson, W.; Cho, N.; Rahimi, S.; Reynolds, S.; Balch, T.; and Veloso, M. 2024 · 2024
Closest in time.