Fetching the paper…
Reading the bibliography…
Trustworthy language models should abstain from answering questions when they do not know the answer.
Calibration of encoder decoder models for neural machine translation
Aviral Kumar and Sunita Sarawagi. 2019 · 1903
Earlier work this paper cites.
An optimum character recognition system using decision functions
Chi-Keung Chow. 1957 · 1957
Earlier work this paper cites.
Knowing more about questions can help: Improving calibration in question answering
Shujian Zhang, Chengyue Gong, and Eunsol Choi. 2021b · 1970
Earlier work this paper cites.
Indices of qualitative variation and political measurement
Allen R. Wilcox. 1973 · 1973
Earlier work this paper cites.
Learning to parse database queries using inductive logic programming
John M Zelle and Raymond J Mooney. 1996 · 1996
Earlier work this paper cites.
Investigating selective prediction approaches across several tasks in IID, OOD, and adversarial settings
Neeraj Varshney, Swaroop Mishra, and Chitta Baral. 2022 · 2002
Earlier work this paper cites.
On the foundations of noise-free selective classification
Ran El-Yaniv et al. 2010 · 2010
Earlier work this paper cites.
Calibrated structured prediction
Volodymyr Kuleshov and Percy S Liang. 2015 · 2015
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Confidence modeling for neural semantic parsing
Li Dong, Chris Quirk, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang. 2020 · 2020
Cited alongside, same era.
AmbigQA: Answering ambiguous open-domain questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Conditional Poisson stochastic beams
Clara Meister, Afra Amini, Tim Vieira, and Ryan Cotterell. 2021 · 2021
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022 · 2022
Later among the works it cites.
ASQA: Factoid questions meet long-form answers
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022 · 2022
Later among the works it cites.
Model cascading: Towards jointly improving efficiency and accuracy of NLP systems
Neeraj Varshney and Chitta Baral. 2022 · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy. 2021 · 2021
Cited alongside, same era.
SituatedQA: Incorporating extra-linguistic contexts into QA
Michael Zhang and Eunsol Choi. 2021 · 2021
Cited alongside, same era.
Estimating the entropy of linguistic distributions
Aryaman Arora, Clara Meister, and Ryan Cotterell. 2022 · 2022
Cited alongside, same era.
Tomayto, tomahto. beyond token-level answer equivalence for question answering evaluation
Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022 · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022 · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022 · 2022
Cited alongside, same era.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022 · 2022
Cited alongside, same era.
Do language models know when they’re hallucinating references?
Ayush Agrawal, Lester Mackey, and Adam Tauman Kalai. 2023 · 2023
Closest in time.
Sources of uncertainty in machine learning–a statisticians’ view
Cornelia Gruber, Patrick Oliver Schenk, Malte Schierholz, Frauke Kreuter, and Göran Kauermann. 2023 · 2023
Closest in time.
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Closest in time.
Generating with confidence: Uncertainty quantification for black-box large language models
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023 · 2023
Closest in time.
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023 · 2023
Closest in time.
Locally typical sampling
Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2023 · 2023
Closest in time.
Out-of-distribution detection and selective generation for conditional language models
Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu. 2023 · 2023
Closest in time.
Navigating the grey area: Expressions of overconfidence and uncertainty in language models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Calibrating structured output predictors for natural language processing
Abhyuday Jagannatha and Hong Yu. 2020 · 2092
Closest in time.