Fetching the paper…
Reading the bibliography…
We propose a novel approach to conformal prediction for generative language models (LMs).
A simple sequentially rejective multiple test procedure
Sture Holm · 1979
Earlier work this paper cites.
Inductive confidence machines for regression
Harris Papadopoulos, Kostas Proedrou, Volodya Vovk, and Alex Gammerman · 2002
Earlier work this paper cites.
On-line confidence machines are well-calibrated
Vladimir Vovk · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Algorithmic Learning in a Random World
Vladimir Vovk, Alex Gammerman, and Glenn Shafer · 2005
Earlier work this paper cites.
Inductive conformal prediction: Theory and application to neural networks
Harris Papadopoulos · 2008
Earlier work this paper cites.
Distribution-free prediction sets
Jing Lei, James Robins, and Larry Wasserman · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Large-scale probabilistic predictors with and without guarantees of validity
Vladimir Vovk, Ivan Petej, and Valentina Fedorova · 2015
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension, 2017
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Earlier work this paper cites.
Nonparametric predictive distributions based on conformal prediction
Vladimir Vovk, Jieli Shen, Valery Manokhin, and Min-ge Xie · 2017
Earlier work this paper cites.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark · 2018
Earlier work this paper cites.
Distribution-free predictive inference for regression
Jing Lei, Max G’Sell, Alessandro Rinaldo, Ryan J. Tibshirani, and Larry Wasserman · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost, 2018
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs
Alistair EW Johnson, Tom J Pollard, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Yifan Peng, Zhiyong Lu, Roger G Mark, Seth J Berkowitz, and Steven Horng · 2019
Earlier work this paper cites.
Clinically accurate chest x-ray report generation
Guanxiong Liu, Tzu-Ming Harry Hsu, Matthew McDermott, Willie Boag, Wei-Hung Weng, Peter Szolovits, and Marzyeh Ghassemi · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Conformalized quantile regression
Yaniv Romano, Evan Patterson, and Emmanuel Candès · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew · 2019
Earlier work this paper cites.
PAWS: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He · 2019
Earlier work this paper cites.
Distribution free, risk controlling prediction sets
Stephen Bates, Anastasios Nikolas Angelopoulos, Lihua Lei, Jitendra Malik, and Michael I. Jordan · 2020
Cited alongside, same era.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Cited alongside, same era.
Distribution-free binary classification: prediction sets, confidence intervals and calibration
Chirag Gupta, Aleksandr Podkopaev, and Aaditya Ramdas · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan · 2022
Later among the works it cites.
SummaC: Re-visiting NLI-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst · 2022
Later among the works it cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Later among the works it cites.
When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi · 2022
Later among the works it cites.
Reducing conversational agents’ overconfidence through linguistic calibration
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chexbert: Combining automatic labelers and expert annotations for accurate radiology report labeling using bert, 2020
Akshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek, Andrew Y. Ng, and Matthew P. Lungren · 2020
Cited alongside, same era.
Predictive inference with the jackknife+
Rina Foygel Barber, Emmanuel J Candes, Aaditya Ramdas, and Ryan J Tibshirani · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig · 2021
Cited alongside, same era.
Hurdles to progress in long-form question answering
Kalpesh Krishna, Aurko Roy, and Mohit Iyyer · 2021
Cited alongside, same era.
Prevent the language model from being overconfident in neural machine translation
Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, and Jie Zhou · 2021
Cited alongside, same era.
Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau · 2022
Later among the works it cites.
Enhancing self-consistency and performance of pre-trained language models through natural language inference
Eric Mitchell, Joseph Noh, Siyan Li, Will Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, and Christopher Manning · 2022
Later among the works it cites.
Improving chest x-ray report generation by leveraging warm-starting, 2022
Aaron Nicolson, Jason Dowling, and Bevan Koopman · 2022
Later among the works it cites.
Characteristics of harmful text: Towards rigorous benchmarking of language models, 2022
Maribeth Rauh, John Mellor, Jonathan Uesato, Po-Sen Huang, Johannes Welbl, Laura Weidinger, Sumanth Dathathri, Amelia Glaese, Geoffrey Irving, Iason Gabriel, William Isaac, and Lisa Anne Hendricks · 2022
Later among the works it cites.
Stretching sentence-pair NLI models to reason over long documents and clusters
Tal Schuster, Sihao Chen, Senaka Buthpitiya, Alex Fabrikant, and Donald Metzler · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2022
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, and et al · 2022
Later among the works it cites.
Conformal risk control, 2023
Anastasios N. Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei, and Tal Schuster · 2023
Closest in time.
Attributed question answering: Evaluation and modeling for attributed large language models, 2023
Bernd Bohnet, Vinh Q. Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, Kai Hui, Tom Kwiatkowski, Ji Ma, Jianmo Ni, Lierni Sestorain Saralegui, Tal Schuster, William W. Cohen, Michael Collins, Dipanjan Das, Donald Metzler, Slav Petrov, and Kellie Webster · 2023
Closest in time.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Closest in time.
Efficiently controlling multiple risks with pareto testing
Bracha Laufer-Goldshtein, Adam Fisch, Regina Barzilay, and Tommi S. Jaakkola · 2023
Closest in time.
Evaluating verifiability in generative search engines, 2023
Nelson F. Liu, Tianyi Zhang, and Percy Liang · 2023
Closest in time.
Gpt-4 technical report
OpenAI · 2023
Closest in time.
Conformal nucleus sampling
Shauli Ravfogel, Yoav Goldberg, and Jacob Goldberger · 2023
Closest in time.
How to trust your diffusion model: A convex optimization approach to conformal risk control
Jacopo Teneggi, Matthew Tivnan, J. Webster Stayman, and Jeremias Sulam · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Generation probabilities are not enough: Exploring the effectiveness of uncertainty highlighting in ai-powered code completions
Helena Vasconcelos, Gagan Bansal, Adam Fourney, Q. Vera Liao, and Jennifer Wortman Vaughan · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Closest in time.
Automatic evaluation of attribution by large language models
Xiang Yue, Boshi Wang, Kai Zhang, Ziru Chen, Yu Su, and Huan Sun · 2023
Closest in time.
On uncertainty calibration and selective generation in probabilistic neural summarization: A benchmark study
Polina Zablotskaia, Du Phan, Joshua Maynez, Shashi Narayan, Jie Ren, and Jeremiah Liu · 2023
Closest in time.
Navigating the grey area: Expressions of overconfidence and uncertainty in language models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto · 2023
Closest in time.