Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) demonstrate remarkable performance in semantic understanding and generation, yet accurately assessing their output reliability remains a significant challenge.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Learning to count objects with few exemplar annotations
Jianfeng Wang, Rong Xiao, Yandong Guo, and Lei Zhang. 2019 · 1905
Earlier work this paper cites.
Eli5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 1907
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N Reimers. 2019 · 1908
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryściński, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 1910
Earlier work this paper cites.
Multicalibration: Calibration for the (computationally-identifiable) masses
Ursula Hébert-Johnson, Michael Kim, Omer Reingold, and Guy Rothblum. 2018 · 1948
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Glenn W Brier. 1950 · 1950
Earlier work this paper cites.
A bayesian approach to calibration
David N DeJong, Beth Fisher Ingram, and Charles H Whiteman. 1996 · 1996
Earlier work this paper cites.
Axiomatic characterization of the quadratic scoring rule
Reinhard Selten. 1998 · 1998
Earlier work this paper cites.
Learning and making decisions when costs and probabilities are both unknown
Bianca Zadrozny and Charles Elkan. 2001 · 2001
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown. 2020 · 2005
Earlier work this paper cites.
Some comparisons among quadratic, spherical, and logarithmic scoring rules
J Eric Bickel. 2007 · 2007
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery. 2007 · 2007
Earlier work this paper cites.
Sampling uncertainty and confidence intervals for the brier score and brier skill score
A Allen Bradley, Stuart S Schwartz, and Tempei Hashino. 2008 · 2008
Earlier work this paper cites.
Pearson correlation coefficient
Israel Cohen, Yiteng Huang, Jingdong Chen, Jacob Benesty, Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. 2009 · 2009
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2009
Earlier work this paper cites.
Accuracy-rejection curves (arcs) for comparing classification methods with a reject option
Malik Sajjad Ahmed Nadeem, Jean-Daniel Zucker, and Blaise Hanczar. 2009 · 2009
Earlier work this paper cites.
Smooth isotonic regression: a new method to calibrate predictive models
Xiaoqian Jiang, Melanie Osl, Jihoon Kim, and Lucila Ohno-Machado. 2011 · 2011
Earlier work this paper cites.
Area under the precision-recall curve: point estimates and confidence intervals
Kendrick Boyd, Kevin H Eng, and C David Page. 2013 · 2013
Earlier work this paper cites.
Bigbench: Towards an industry standard benchmark for big data analytics
Ahmad Ghazal, Tilmann Rabl, Minqing Hu, Francois Raab, Meikel Poess, Alain Crolotte, and Hans-Arno Jacobsen. 2013 · 2013
Earlier work this paper cites.
Question answering with subgraph embeddings
Antoine Bordes, Sumit Chopra, and Jason Weston. 2014 · 2014
Earlier work this paper cites.
Area under precision-recall curves for weighted and unweighted data
Jens Keilwagen, Ivo Grosse, and Jan Grau. 2014 · 2014
Earlier work this paper cites.
Spearman’s rank correlation coefficient
Philip Sedgwick. 2014 · 2014
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
Beta calibration: a well-founded and easily implemented improvement on logistic calibration for binary classifiers
Meelis Kull, Telmo Silva Filho, and Peter Flach. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. 2018 · 2018
Earlier work this paper cites.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. 2018 · 2018
Earlier work this paper cites.
To trust or not to trust a classifier
Heinrich Jiang, Been Kim, Melody Guan, and Maya Gupta. 2018 · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Retrieval augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021 · 2021
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021 · 2021
Earlier work this paper cites.
Top-label calibration and multiclass-to-binary reductions
Chirag Gupta and Aaditya Ramdas. 2021 · 2021
Earlier work this paper cites.
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021 · 2021
Cited alongside, same era.
Bayesian confidence calibration for epistemic uncertainty modelling
Fabian Küppers, Jan Kronenberger, Jonas Schneider, and Anselm Haselhoff. 2021 · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Cited alongside, same era.
Are nlp models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2021
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023 · 2023
Later among the works it cites.
A comprehensive capability analysis of gpt-3 and gpt-3.5 series models
Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, et al. 2023 · 2023
Later among the works it cites.
Jiaxin Zhang, Zhuohang Li, Kamalika Das, Bradley A Malin, and Sricharan Kumar. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic optimization of areas under precision-recall curves with provable convergence
Qi Qi, Youzhi Luo, Zhao Xu, Shuiwang Ji, and Tianbao Yang. 2021 · 2021
Cited alongside, same era.
Can explanations be useful for calibrating black box models?
Xi Ye and Greg Durrett. 2021 · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Chatgpt: Fundamentals, applications and social impacts
Malak Abdullah, Alia Madain, and Yaser Jararweh. 2022 · 2022
Cited alongside, same era.
Samuel Joseph Amouyal, Tomer Wolfson, Ohad Rubin, Ori Yoran, Jonathan Herzig, and Jonathan Berant. 2022 · 2022
Cited alongside, same era.
Metrics of calibration for probabilistic predictions
Imanol Arrieta-Ibarra, Paman Gujral, Jonathan Tannen, Mark Tygert, and Cherie Xu. 2022 · 2022
Cited alongside, same era.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
Chiwei Zhu, Benfeng Xu, Quan Wang, Yongdong Zhang, and Zhendong Mao. 2023 · 2023
Later among the works it cites.
Linguistic calibration of long-form generations
Neil Band, Xuechen Li, Tengyu Ma, and Tatsunori Hashimoto. 2024 · 2024
Closest in time.
Cycles of thought: Measuring llm confidence through stable explanations
Evan Becker and Stefano Soatto. 2024 · 2024
Closest in time.
Decoding by contrasting knowledge: Enhancing llms’ confidence on edited facts
Baolong Bi, Shenghua Liu, Lingrui Mei, Yiwei Wang, Pengliang Ji, and Xueqi Cheng. 2024 · 2024
Closest in time.
Claude 2.0 large language model: Tackling a real-world classification problem with a new iterative prompt engineering approach
Loredana Caruccio, Stefano Cirillo, Giuseppe Polese, Giandomenico Solimando, Shanmugam Sundaramurthy, and Genoveffa Tortora. 2024 · 2024
Closest in time.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024 · 2024
Closest in time.
Quantifying uncertainty in answers from any language model and enhancing their trustworthiness
Jiuhai Chen and Jonas Mueller. 2024 · 2024
Closest in time.
Multicalibration for confidence scoring in llms
Gianluca Detommaso, Martin Bertran, Riccardo Fogliato, and Aaron Roth. 2024 · 2024
Closest in time.
Facilitating human-llm collaboration through factuality scores and source attributions
Hyo Jin Do, Rachel Ostrand, Justin D Weisz, Casey Dugan, Prasanna Sattigeri, Dennis Wei, Keerthiram Murugesan, and Werner Geyer. 2024 · 2024
Closest in time.
Counterfactual debating with preset stances for hallucination elimination of llms
Yi Fang, Moxin Li, Wenjie Wang, Hui Lin, and Fuli Feng. 2024 · 2024
Closest in time.
Don’t hallucinate, abstain: Identifying llm knowledge gaps via multi-llm collaboration
Shangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Vidhisha Balachandran, and Yulia Tsvetkov. 2024 · 2024
Closest in time.
A survey of confidence estimation and calibration in large language models
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2024 · 2024
Closest in time.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, et al. 2024 · 2024
Closest in time.
Llm-rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts
Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, and Chris Kedzie. 2024 · 2024
Closest in time.
Calibration-tuning: Teaching large language models to know what they don’t know
Sanyam Kapoor, Nate Gruver, Manley Roberts, Arka Pal, Samuel Dooley, Micah Goldblum, and Andrew Wilson. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2024 · 2024
Closest in time.
Shiyu Ni, Keping Bi, Jiafeng Guo, and Xueqi Cheng. 2024 · 2024
Closest in time.
Kernel language entropy: Fine-grained uncertainty quantification for llms from semantic similarities
Alexander Nikitin, Jannik Kossen, Yarin Gal, and Pekka Marttinen. 2024 · 2024
Closest in time.
Large language model confidence estimation via black-box access
Tejaswini Pedapati, Amit Dhurandhar, Soumya Ghosh, Soham Dan, and Prasanna Sattigeri. 2024 · 2024
Closest in time.
Enhancing healthcare llm trust with atypical presentations recalibration
Jeremy Qin, Bang Liu, and Quoc Dinh Nguyen. 2024 · 2024
Closest in time.
Large language model uncertainty measurement and calibration for medical diagnosis and treatment
Thomas Savage, John Wang, Robert Gallo, Abdessalem Boukil, Vishwesh Patel, Seyed Amir Ahmad Safavi-Naini, Ali Soroush, and Jonathan H Chen. 2024 · 2024
Closest in time.
Thermometer: Towards universal calibration for large language models
Maohao Shen, Subhro Das, Kristjan Greenewald, Prasanna Sattigeri, Gregory Wornell, and Soumya Ghosh. 2024 · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024 · 2024
Closest in time.
Towards robust evaluation: A comprehensive taxonomy of datasets and metrics for open domain question answering in the era of large language models
Akchay Srivastava and Atif Memon. 2024 · 2024
Closest in time.
The calibration gap between model and human confidence in large language models
Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Sheer Karny, Xinyue Hu, Lukas Mayer, and Padhraic Smyth. 2024 · 2024
Closest in time.
Bayesian prompt ensembles: Model uncertainty estimation for black-box large language models
Francesco Tonolini, Nikolaos Aletras, Jordan Massiah, and Gabriella Kazai. 2024 · 2024
Closest in time.
Yao-Hung Hubert Tsai, Walter Talbott, and Jian Zhang. 2024 · 2024
Closest in time.
Calibrating large language models using their generations only
Dennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun, and Seong Joon Oh. 2024 · 2024
Closest in time.
Black-box uncertainty quantification method for llm-as-a-judge
Nico Wagner, Michael Desmond, Rahul Nair, Zahra Ashktorab, Elizabeth M Daly, Qian Pan, Martín Santillán Cooper, James M Johnson, and Werner Geyer. 2024 · 2024
Closest in time.
Sportqa: A benchmark for sports understanding in large language models
Haotian Xia, Zhengbang Yang, Yuqing Wang, Rhys Tracy, Yun Zhao, Dongdong Huang, Zezhi Chen, Yan Zhu, Yuan-fang Wang, and Weining Shen. 2024 · 2024
Closest in time.
Mirror: A multiple-perspective self-reflection method for knowledge-rich reasoning
Hanqi Yan, Qinglin Zhu, Xinyu Wang, Lin Gui, and Yulan He. 2024 · 2024
Closest in time.
Can we trust llms? mitigate overconfidence bias in llms through knowledge transfer
Haoyan Yang, Yixuan Wang, Xingyin Xu, Hanyuan Zhang, and Yirong Bian. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Efficiently deploying llms with controlled risk
Michael J Zellinger and Matt Thomson. 2024 · 2024
Closest in time.
Hallucinations in llms: Understanding and addressing challenges
Gabrijela Perković, Antun Drobnjak, and Ivica Botički. 2024 · 2088
Closest in time.