Fetching the paper…
Reading the bibliography…
While large language models (LLMs) achieve strong performance on text-to-SQL parsing, they sometimes exhibit unexpected failures in which they are confidently incorrect.
Calibration of encoder decoder models for neural machine translation
Aviral Kumar and Sunita Sarawagi. 2019 · 1903
Earlier work this paper cites.
Multicalibration: Calibration for the (computationally-identifiable) masses
Ursula Hébert-Johnson, Michael Kim, Omer Reingold, and Guy Rothblum. 2018 · 1948
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt et al. 1999 · 1999
Earlier work this paper cites.
Towards a theory of natural language interfaces to databases
Ana-Maria Popescu, Oren Etzioni, and Henry Kautz. 2003 · 2003
Earlier work this paper cites.
Confidence estimation for machine translation
John Blatz, Erin Fitzgerald, George Foster, Simona Gandrabur, Cyril Goutte, Alex Kulesza, Alberto Sanchis, and Nicola Ueffing. 2004 · 2004
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana. 2005 · 2005
Earlier work this paper cites.
Word-level confidence estimation for machine translation using phrase-based translation models
Nicola Ueffing and Hermann Ney. 2005 · 2005
Earlier work this paper cites.
A tutorial on conformal prediction
Glenn Shafer and Vladimir Vovk. 2008 · 2008
Earlier work this paper cites.
Reliability, sufficiency, and the decomposition of proper scores
Jochen Bröcker. 2009 · 2009
Earlier work this paper cites.
Translating questions to SQL queries with generative parsers discriminatively reranked
Alessandra Giordani and Alessandro Moschitti. 2012 · 2012
Earlier work this paper cites.
A framework for merging and ranking of answers in deepqa
D. C. Gondek, A. Lally, A. Kalyanpur, J. W. Murdock, P. A. Duboue, L. Zhang, Y. Pan, Z. M. Qiu, and C. Welty. 2012 · 2012
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
Beam search strategies for neural machine translation
Markus Freitag and Yaser Al-Onaizan. 2017 · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
Learning a neural semantic parser from user feedback
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Victor Zhong, Caiming Xiong, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Confidence modeling for neural semantic parsing
Li Dong, Chris Quirk, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018 · 2018
Earlier work this paper cites.
Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration
Meelis Kull, Miquel Perello Nieto, Markus Kängsepp, Telmo Silva Filho, Hao Song, and Peter Flach. 2019 · 2019
Earlier work this paper cites.
Measuring calibration in deep learning
Jeremy Nixon, Michael W. Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. 2019 · 2019
Cited alongside, same era.
Model-based interactive semantic parsing: A unified framework and a text-to-SQL case study
Ziyu Yao, Yu Su, Huan Sun, and Wen-tau Yih. 2019 · 2019
Cited alongside, same era.
Distribution-free binary classification: prediction sets, confidence intervals and calibration
Chirag Gupta, Aleksandr Podkopaev, and Aaditya Ramdas. 2020 · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Multivariate confidence calibration for object detection
Fabian Kuppers, Jan Kronenberger, Amirhossein Shantia, and Anselm Haselhoff. 2020 · 2020
Cited alongside, same era.
On the inference calibration of neural machine translation
Online Platt scaling with calibeating
Chirag Gupta and Aaditya Ramdas. 2023 · 2023
Later among the works it cites.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023 · 2023
Later among the works it cites.
A multilingual translator to sql with database schema pruning to improve self-attention
Marcelo Archanjo Jose and Fabio Gagliardi Cozman. 2023 · 2023
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023 · 2023
Later among the works it cites.
Improving generalization in language model-based text-to-SQL semantic parsing: Two simple semantic boundary-based techniques
Daking Rai, Bailin Wang, Yilun Zhou, and Ziyu Yao. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuo Wang, Zhaopeng Tu, Shuming Shi, and Yang Liu. 2020 · 2020
Cited alongside, same era.
How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Picard: Parsing incrementally for constrained auto-regressive decoding from language models
Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. 2021 · 2021
Cited alongside, same era.
Low-degree multicalibration
Parikshit Gopalan, Michael P Kim, Mihir A Singhal, and Shengjia Zhao. 2022 · 2022
Cited alongside, same era.
Post-hoc calibration without distributional assumptions
Chirag Gupta. 2022 · 2022
Cited alongside, same era.
Top-label Calibration and Multiclass-to-binary Reductions
Chirag Gupta and Aaditya Ramdas. 2022 · 2022
Cited alongside, same era.
Batch multivalid conformal prediction
Christopher Jung, Georgy Noarov, Ramya Ramalingam, and Aaron Roth. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Calibrated interpretation: Confidence estimation in semantic parsing
Elias Stengel-Eskin and Benjamin Van Durme. 2023 · 2023
Later among the works it cites.
Exploring chain of thought style prompting for text-to-sql
Chang-Yu Tai, Ziru Chen, Tianshu Zhang, Xiang Deng, and Huan Sun. 2023 · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. 2023 · 2023
Later among the works it cites.
Know what I don’t know: Handling ambiguous and unknown questions for text-to-SQL
Bing Wang, Yan Gao, Zhoujun Li, and Jian-Guang Lou. 2023 · 2023
Later among the works it cites.
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2023 · 2023
Later among the works it cites.
Multicalibration for confidence scoring in LLMs
Gianluca Detommaso, Martin Bertran, Riccardo Fogliato, and Aaron Roth. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Later among the works it cites.
On the sample complexity of parameter estimation in logistic regression with normal design
Daniel Hsu and Arya Mazumdar. 2024 · 2024
Later among the works it cites.
Can LLM already Serve as a Database Interface? A Big Bench for Large-scale Database Grounded Text-to-SQLs
Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et al. 2024 · 2024
Later among the works it cites.
Multi-group uncertainty quantification for long-form text generation
Terrance Liu and Zhiwei Steven Wu. 2024 · 2024
Later among the works it cites.
Language models with conformal factuality guarantees
Christopher Mohri and Tatsunori Hashimoto. 2024 · 2024
Later among the works it cites.
Din-sql: Decomposed in-context learning of text-to-sql with self-correction
Mohammadreza Pourreza and Davood Rafiei. 2024 · 2024
Later among the works it cites.
Relying on the unreliable: The impact of language models’ reluctance to express uncertainty
Kaitlyn Zhou, Jena Hwang, Xiang Ren, and Maarten Sap. 2024 · 2024
Later among the works it cites.