Fetching the paper…
Reading the bibliography…
Large language models (LLMs) specializing in natural language generation (NLG) have recently started exhibiting promising capabilities across a variety of domains.
On Optimum Recognition Error and Reject Tradeoff
C. K. Chow · 1970
Earlier work this paper cites.
Reject option with multiple thresholds
Giorgio Fumera, Fabio Roli, and Giorgio Giacinto · 2000
Earlier work this paper cites.
On spectral clustering: Analysis and an algorithm
Andrew Ng, Michael Jordan, and Yair Weiss · 2001
Earlier work this paper cites.
Learning and making decisions when costs and probabilities are both unknown
Bianca Zadrozny and Charles Elkan · 2001
Earlier work this paper cites.
A tutorial on spectral clustering
Ulrike Von Luxburg · 2007
Earlier work this paper cites.
Accuracy-rejection curves (arcs) for comparing classification methods with a reject option
Malik Sajjad Ahmed Nadeem, Jean-Daniel Zucker, and Blaise Hanczar · 2009
Earlier work this paper cites.
Sabrina J. Mielke, Arthur Szlam, Y-Lan Boureau, and Emily Dinan · 2012
Earlier work this paper cites.
Uncertainties have a meaning: Information entropy as a quality measure for 3-d geological models
J Florian Wellmann and Klaus Regenauer-Lieb · 2012
Earlier work this paper cites.
Align, disambiguate and walk: A unified approach for measuring semantic similarity
Mohammad Taher Pilehvar, David Jurgens, and Roberto Navigli · 2013
Earlier work this paper cites.
Reliable classification: Learning classifiers that distinguish aleatoric and epistemic uncertainty
Robin Senge, Stefan Bösner, Krzysztof Dembczyński, Jörg Haasenritter, Oliver Hirsch, Norbert Donner-Banzhoff, and Eyke Hüllermeier · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Probabilistic backpropagation for scalable learning of bayesian neural networks
José Miguel Hernández-Lobato and Ryan P. Adams · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
A comparison of rule-based and machine learning approaches for classifying patient portal messages
Robert M Cronin, Daniel Fabbri, Joshua C Denny, S Trent Rosenbloom, and Gretchen Purcell Jackson · 2017
Earlier work this paper cites.
Selective classification for deep neural networks
Yonatan Geifman and Ran El-Yaniv · 2017
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
To trust or not to trust a classifier
Heinrich Jiang, Been Kim, Maya Gupta, and Melody Y. Guan · 2018
Earlier work this paper cites.
Predictive uncertainty estimation via prior networks
Andrey Malinin and Mark Gales · 2018
Earlier work this paper cites.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
L. McInnes, J. Healy, and J. Melville · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
Addressing failure prediction by learning model confidence
Charles Corbière, Nicolas THOME, Avner Bar-Hen, Matthieu Cord, and Patrick Pérez · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Cited alongside, same era.
Measuring calibration in deep learning
Jeremy Nixon, Michael W. Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran · 2019
Cited alongside, same era.
Reducing conversational agents’ overconfidence through linguistic calibration
Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau · 2022
Later among the works it cites.
Confident adaptive language modeling
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Donald Metzler · 2022
Later among the works it cites.
Re-examining calibration: The case of question answering
Chenglei Si, Chen Zhao, Sewon Min, and Jordan Boyd-Graber · 2022
Later among the works it cites.
Investigating selective prediction approaches across several tasks in IID, OOD, and adversarial settings
Neeraj Varshney, Swaroop Mishra, and Chitta Baral · 2022
Later among the works it cites.
Towards improving selective prediction ability of NLP systems
Neeraj Varshney, Swaroop Mishra, and Chitta Baral · 2022
Later among the works it cites.
Uncertainty estimation and reduction of pre-trained models for text regression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CoQA: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning · 2019
Cited alongside, same era.
Feature selection using neighborhood entropy-based uncertainty measures for gene expression data classification
Lin Sun, Xiaoyu Zhang, Yuhua Qian, Jiucheng Xu, and Shiguang Zhang · 2019
Cited alongside, same era.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett · 2020
Cited alongside, same era.
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang · 2020
Cited alongside, same era.
A survey on recognizing textual entailment as an NLP evaluation
Adam Poliak · 2020
Cited alongside, same era.
Document processing: Methods for semantic text similarity analysis
Abdul Wahab Qurashi, Violeta Holmes, and Anju P. Johnson · 2020
Cited alongside, same era.
A review of uncertainty quantification in deep learning: Techniques, applications and challenges
Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Xiaochun Cao, Abbas Khosravi, U. Rajendra Acharya, Vladimir Makarenkov, and Saeid Nahavandi · 2021
Cited alongside, same era.
Yuxia Wang, Daniel Beck, Timothy Baldwin, and Karin Verspoor · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Jiuhai Chen and Jonas Mueller · 2023
Closest in time.
What comes next? evaluating uncertainty in neural text generators against human production variability, 2023
Mario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández, and Barbara Plank · 2023
Closest in time.
Gpt-4 passes the bar exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2023
Closest in time.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Closest in time.
DEUP: Direct epistemic uncertainty prediction
Salem Lahlou, Moksh Jain, Hadi Nekoei, Victor I Butoi, Paul Bertin, Jarrid Rector-Brooks, Maksym Korablyov, and Yoshua Bengio · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Out-of-distribution detection and selective generation for conditional language models
Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu · 2023
Closest in time.
Prompting GPT-3 to be reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Lee Boyd-Graber, and Lijuan Wang · 2023
Closest in time.
Post-abstention: Towards reliably re-attempting the abstained instances in QA
Neeraj Varshney and Chitta Baral · 2023
Closest in time.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms, 2023
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2023
Closest in time.
Navigating the grey area: Expressions of overconfidence and uncertainty in language models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto · 2023
Closest in time.