Fetching the paper…
Reading the bibliography…
Recently, Large Language Models (LLMs) have been increasingly used to support various decision-making tasks, assisting humans in making informed decisions.
“Verification of forecasts expressed in terms of probability”
Glenn Brier · 1950
Earlier work this paper cites.
“Elicitation of personal probabilities and expectations”
Leonard Savage · 1971
Earlier work this paper cites.
“Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval”
Stephen Robertson and Steve Walker · 1994
Earlier work this paper cites.
“Strictly proper scoring rules, prediction, and estimation”
Tilmann Gneiting and Adrian Raftery · 2007
Earlier work this paper cites.
“Obtaining well calibrated probabilities using bayesian binning”
Mahdi Naeini, Gregory Cooper and Milos Hauskrecht · 2015
Earlier work this paper cites.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“On calibration of modern neural networks”
Chuan Guo, Geoff Pleiss, Yu Sun and Kilian Weinberger · 2017
Earlier work this paper cites.
“Proximal policy optimization algorithms”
John Schulman et al · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension”
Mandar Joshi, Eunsol Choi, Daniel Weld and Luke Zettlemoyer · 2017
Earlier work this paper cites.
“Accurate uncertainties for deep learning using calibrated regression”
Volodymyr Kuleshov, Nathan Fenner and Stefano Ermon · 2018
Earlier work this paper cites.
“Know what you don’t know: Unanswerable questions for SQuAD”
Pranav Rajpurkar, Robin Jia and Percy Liang · 2018
Earlier work this paper cites.
“WikiQA: A challenge dataset for open-domain question answering”
Yi Yang, Wen-tau Yih and Christopher Meek · 2018
Earlier work this paper cites.
“HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering”
Zhilin Yang et al · 2018
Earlier work this paper cites.
“Don’t say that! making inconsistent dialogue unlikely with unlikelihood training”
Margaret Li et al · 2019
Earlier work this paper cites.
“Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration”
Meelis Kull et al · 2019
Earlier work this paper cites.
“Evaluating model calibration in classification”
Juozas Vaicenavicius et al · 2019
Earlier work this paper cites.
“Natural Questions: A Benchmark for Question Answering Research”
Tom Kwiatkowski et al · 2019
Earlier work this paper cites.
“Pytorch: An imperative style, high-performance deep learning library”
Adam Paszke et al · 2019
Earlier work this paper cites.
“HuggingFace’s Transformers: State-of-the-art natural language processing”
Thomas Wolf et al · 2019
Earlier work this paper cites.
“Retrieval-augmented generation for knowledge-intensive nlp tasks”
Patrick Lewis et al · 2020
Cited alongside, same era.
“Well-calibrated regression uncertainty in medical imaging with deep learning”
Max-Heinrich Laves et al · 2020
Cited alongside, same era.
“Dense passage retrieval for open-domain question answering”
Vladimir Karpukhin et al · 2020
Cited alongside, same era.
“Distilling knowledge from reader to retriever for question answering”
Gautier Izacard and Edouard Grave · 2020
Cited alongside, same era.
“On the opportunities and risks of foundation models”
Rishi Bommasani et al · 2021
Cited alongside, same era.
“Ra-dit: Retrieval-augmented dual instruction tuning”
Xi Lin et al · 2023
Later among the works it cites.
“Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback”
Katherine Tian et al · 2023
Later among the works it cites.
“Is ChatGPT good at search? investigating large language models as re-ranking agents”
Weiwei Sun et al · 2023
Later among the works it cites.
“Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection”
Akari Asai et al · 2023
Later among the works it cites.
“MedCPT: Contrastive Pre-trained Transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval”
Qiao Jin et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kurt Shuster et al · 2021
Cited alongside, same era.
“Revisiting the calibration of modern neural networks”
Matthias Minderer et al · 2021
Cited alongside, same era.
“Unsupervised dense information retrieval with contrastive learning”
Gautier Izacard et al · 2021
Cited alongside, same era.
“LoRA: Low-Rank Adaptation of Large Language Models”
Edward Hu et al · 2021
Cited alongside, same era.
Wei Li et al · 2022
Cited alongside, same era.
“A survey on retrieval-augmented text generation”
Huayang Li et al · 2022
Cited alongside, same era.
“WebQA: Multihop and multimodal QA”
Yingshan Chang et al · 2022
Cited alongside, same era.
Miao Xiong et al · 2023
Later among the works it cites.
“BioASQ-QA: A manually curated corpus for Biomedical Question Answering”
Anastasia Krithara, Anastasios Nentidis, Konstantinos Bougiatiotis and Georgios Paliouras · 2023
Later among the works it cites.
Abhimanyu Dubey et al · 2024
Closest in time.
“Linguistic Calibration of Long-Form Generations”
Neil Band, Xuechen Li, Tengyu Ma and Tatsunori Hashimoto · 2024
Closest in time.
“Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty”
Kaitlyn Zhou, Jena Hwang, Xiang Ren and Maarten Sap · 2024
Closest in time.
“Position paper: Bayesian deep learning in the age of large-scale ai”
Theodore Papamarkou et al · 2024
Closest in time.
“Towards safe large language models for medicine”
Tessa Han, Aounon Kumar, Chirag Agarwal and Himabindu Lakkaraju · 2024
Closest in time.
“Potential for GPT technology to optimize future clinical decision-making using retrieval-augmented generation”
Calvin Wang et al · 2024
Closest in time.
Jiarui Li, Ye Yuan and Zehua Zhang · 2024
Closest in time.
“Calibration-Tuning: Teaching Large Language Models to Know What They Don’t Know”
Sanyam Kapoor et al · 2024
Closest in time.
“Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs”
Miao Xiong et al · 2024
Closest in time.
“Benchmarking Retrieval-Augmented Generation for Medicine”
Guangzhi Xiong, Qiao Jin, Zhiyong Lu and Aidong Zhang · 2024
Closest in time.
Marah Abdin et al · 2024
Closest in time.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”
Daya Guo et al · 2025
Closest in time.