Fetching the paper…
Reading the bibliography…
Accurate uncertainty quantification is crucial for the safe deployment of machine learning models, and prior research has demonstrated improvements in the calibration of modern language models (LMs).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
The well-calibrated bayesian
A Philip Dawid. 1982 · 1982
Earlier work this paper cites.
The comparison and evaluation of forecasters
Morris H DeGroot and Stephen E Fienberg. 1983 · 1983
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt et al. 1999 · 1999
Earlier work this paper cites.
Building a question answering test collection
Ellen M Voorhees and Dawn M Tice. 2000 · 2000
Earlier work this paper cites.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett. 2020 · 2003
Earlier work this paper cites.
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang. 2020 · 2006
Earlier work this paper cites.
Efficient conformal prediction via cascaded inference with expanded admission
Adam Fisch, Tal Schuster, Tommi Jaakkola, and Regina Barzilay. 2020 · 2007
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Novel decompositions of proper scoring rules for classification: Score adjustment as precursor to calibration
Meelis Kull and Peter Flach. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Peter A Jansen, Elizabeth Wainwright, Steven Marmorstein, and Clayton T Morrison. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
Verified uncertainty calibration
Ananya Kumar, Percy S Liang, and Tengyu Ma. 2019 · 2019
Earlier work this paper cites.
Measuring calibration in deep learning
Jeremy Nixon, Michael W Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. 2019 · 2019
Earlier work this paper cites.
Calibration, entropy rates, and memory in language models
Mark Braverman, Xinyi Chen, Sham Kakade, Karthik Narasimhan, Cyril Zhang, and Yi Zhang. 2020 · 2020
Earlier work this paper cites.
Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making
Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. 2020 · 2020
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021 · 2021
Earlier work this paper cites.
Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty
Umang Bhatt, Javier Antorán, Yunfeng Zhang, Q Vera Liao, Prasanna Sattigeri, Riccardo Fogliato, Gabrielle Melançon, Ranganath Krishnan, Jason Stanley, Omesh Tickoo, et al. 2021 · 2021
Cited alongside, same era.
To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021 · 2021
Cited alongside, same era.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021 · 2021
Cited alongside, same era.
Do explanations help users detect errors in open-domain qa? an evaluation of spoken vs. visual explanations
Ana Valeria González, Gagan Bansal, Angela Fan, Yashar Mehdad, Robin Jia, and Srinivasan Iyer. 2021 · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei. 2023 · 2023
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023 · 2023
Closest in time.
Selectively answering ambiguous questions
Jeremy R Cole, Michael JQ Zhang, Daniel Gillick, Julian Martin Eisenschlos, Bhuwan Dhingra, and Jacob Eisenstein. 2023 · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. 2021 · 2021
Cited alongside, same era.
Exploiting cloze-questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
Consistent accelerated inference via confident adaptive transformers
Tal Schuster, Adam Fisch, Tommi Jaakkola, and Regina Barzilay. 2021 · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2021 · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Anastasios N Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei, and Tal Schuster. 2022 · 2022
Cited alongside, same era.
Yu Fei, Yifan Hou, Zeming Chen, and Antoine Bosselut. 2023 · 2023
Closest in time.
Prototypical calibration for few-shot learning of language models
Zhixiong Han, Yaru Hao, Li Dong, Yutao Sun, and Furu Wei. 2023 · 2023
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Closest in time.
Calibrated language models must hallucinate
Adam Tauman Kalai and Santosh S. Vempala. 2023 · 2023
Closest in time.
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Closest in time.
Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas. 2023 · 2023
Closest in time.
Gpt-4 technical report
OpenAI. 2023 · 2023
Closest in time.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks. 2023 · 2023
Closest in time.
What In-Context Learning “Learns” In-Context: Disentangling Task Recognition and Task Learning
Jane Pan. 2023 · 2023
Closest in time.
Simple token-level confidence improves caption correctness
Suzanne Petryk, Spencer Whitehead, Joseph E Gonzalez, Trevor Darrell, Anna Rohrbach, and Marcus Rohrbach. 2023 · 2023
Closest in time.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Allan Raventós, Mansheej Paul, Feng Chen, and Surya Ganguli. 2023 · 2023
Closest in time.
Long horizon temperature scaling
Andy Shih, Dorsa Sadigh, and Stefano Ermon. 2023 · 2023
Closest in time.
Large language models help humans verify truthfulness–except when they are convincingly wrong
Chenglei Si, Navita Goyal, Sherry Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé III, and Jordan Boyd-Graber. 2023 · 2023
Closest in time.
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. 2023 · 2023
Closest in time.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2023 · 2023
Closest in time.
Theodore Zhao, Mu Wei, J Samuel Preston, and Hoifung Poon. 2023 · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Closest in time.