Fetching the paper…
Reading the bibliography…
Current question-answering benchmarks predominantly focus on accuracy in realizable prediction tasks.
Verification of forecasts expressed in terms of probability
Glenn W. Brier · 1950
Earlier work this paper cites.
The theory of signal detectability
WWTG Peterson, T Birdsall, and We Fox · 1954
Earlier work this paper cites.
Test bias: Prediction of grades of negro and white students in integrated colleges
T Anne Cleary · 1968
Earlier work this paper cites.
Reliability of subjective probability forecasts of precipitation and temperature
Allan H Murphy and Robert L Winkler · 1977
Earlier work this paper cites.
The comparison and evaluation of forecasters
Morris H DeGroot and Stephen E Fienberg · 1983
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt et al · 1999
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers
Bianca Zadrozny and Charles Elkan · 2001
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
Survey Methodology
R.M. Groves, F.J. Fowler, M.P. Couper, J.M. Lepkowski, E. Singer, and R. Tourangeau · 2009
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
Armen Der Kiureghian and Ove Ditlevsen · 2009
Earlier work this paper cites.
The black box society: The secret algorithms that control money and information
Frank Pasquale · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
XGBoost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Alexandra Chouldechova · 2017
Earlier work this paper cites.
On fairness and calibration
Geoff Pleiss, Manish Raghavan, Felix Wu, Jon Kleinberg, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Integrated public use microdata series, Current Population Survey: Version 6.0
Sarah Flood, Miriam King, Renae Rodgers, Steven Ruggles, and J Robert Warren · 2018
Earlier work this paper cites.
Automating inequality: How high-tech tools profile, police, and punish the poor
Virginia Eubanks · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Analyzing uncertainty in neural machine translation
Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato · 2018
Earlier work this paper cites.
Multicalibration: Calibration for the (Computationally-identifiable) masses
Ursula Hebert-Johnson, Michael Kim, Omer Reingold, and Guy Rothblum · 2018
Earlier work this paper cites.
Race after technology: Abolitionist tools for the new Jim code
Ruha Benjamin · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
50 years of test (un) fairness: Lessons for machine learning
Ben Hutchinson and Margaret Mitchell · 2019
Earlier work this paper cites.
TableGPT: Few-shot table-to-text generation with table structure reconstruction and content matching
Heng Gong, Yawei Sun, Xiaocheng Feng, Bing Qin, Wei Bi, Xiaojiang Liu, and Ting Liu · 2020
Earlier work this paper cites.
Retiring adult: New datasets for fair machine learning
Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt · 2021
Cited alongside, same era.
Fairness, equality, and power in algorithmic decision-making
Maximilian Kasy and Rediet Abebe · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization, 2021
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Can ai language models replace human participants?
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray · 2023
Later among the works it cites.
A close look into the calibration of pre-trained language models
Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji · 2023
Later among the works it cites.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Later among the works it cites.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al · 2023
Later among the works it cites.
Evaluating the moral beliefs encoded in LLMs
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Cited alongside, same era.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2022
Cited alongside, same era.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
Reducing conversational agents’ overconfidence through linguistic calibration
Sabrina J. Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau · 2022
Cited alongside, same era.
Deep neural networks and tabular data: A survey
Vadim Borisov, Tobias Leemann, Kathrin Seßler, Johannes Haug, Martin Pawelczyk, and Gjergji Kasneci · 2022
Cited alongside, same era.
Evaluating and mitigating discrimination in language model decisions
Alex Tamkin, Amanda Askell, Liane Lovitt, Esin Durmus, Nicholas Joseph, Shauna Kravec, Karina Nguyen, Jared Kaplan, and Deep Ganguli · 2023
Cited alongside, same era.
Unified language representation for question answering over text, tables, and images
Bowen Yu, Cheng Fu, Haiyang Yu, Fei Huang, and Yongbin Li · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning · 2023
Later among the works it cites.
Leveraging large language models for multiple choice question answering
Joshua Robinson and David Wingate · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Auditing the use of language models to guide hiring decisions
Johann D Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe · 2024
Closest in time.
The relative value of prediction in algorithmic decision making
Juan Carlos Perdomo · 2024
Closest in time.
Against predictive optimization: On the legitimacy of decision-making algorithms that optimize predictive accuracy
Angelina Wang, Sayash Kapoor, Solon Barocas, and Arvind Narayanan · 2024
Closest in time.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi · 2024
Closest in time.
A study on the calibration of in-context learning
Hanlin Zhang, Yi-Fan Zhang, Yaodong Yu, Dhruv Madeka, Dean Foster, Eric Xing, Himabindu Lakkaraju, and Sham Kakade · 2024
Closest in time.
Questioning the survey responses of large language models
Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner · 2024
Closest in time.
Do LLMs exhibit human-like response biases? a case study in survey design
Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkar, and Graham Neubig · 2024
Closest in time.
Calibrated language models must hallucinate
Adam Tauman Kalai and Santosh S. Vempala · 2024
Closest in time.
Linguistic calibration of language models
Neil Band, Xuechen Li, Tengyu Ma, and Tatsunori Hashimoto · 2024
Closest in time.
Generating with confidence: Uncertainty quantification for black-box large language models, 2024
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun · 2024
Closest in time.
Large language models(LLMs) on tabular data: Prediction, generation, and understanding – a survey
Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos · 2024
Closest in time.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al · 2024
Closest in time.