Fetching the paper…
Reading the bibliography…
Before deploying outputs from foundation models in high-stakes tasks, it is imperative to ensure that they align with human values.
Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator
Aryeh Dvoretzky, Jack Kiefer, and Jacob Wolfowitz · 1956
Earlier work this paper cites.
An optimum character recognition system using decision functions
Chi-Keung Chow · 1957
Earlier work this paper cites.
The tight constant in the dvoretzky-kiefer-wolfowitz inequality
Pascal Massart · 1990
Earlier work this paper cites.
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg · 1995
Earlier work this paper cites.
The positive false discovery rate: a bayesian interpretation and the q-value
John D Storey · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Chexbert: Combining automatic labelers and expert annotations for accurate radiology report labeling using bert. arxiv [cscl]. published online april 20, 2020, 2004
A Smit, S Jain, P Rajpurkar, A Pareek, AY Ng, and MP Lungren · 2004
Earlier work this paper cites.
Algorithmic learning in a random world
Vladimir Vovk, Alexander Gammerman, and Glenn Shafer · 2005
Earlier work this paper cites.
On the foundations of noise-free selective classification
Ran El-Yaniv et al · 2010
Earlier work this paper cites.
Selective classification for deep neural networks
Yonatan Geifman and Ran El-Yaniv · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
The ethics of artificial intelligence
Nick Bostrom and Eliezer Yudkowsky · 2018
Earlier work this paper cites.
Distribution-free predictive inference for regression
Jing Lei, Max G’Sell, Alessandro Rinaldo, Ryan J Tibshirani, and Larry Wasserman · 2018
Earlier work this paper cites.
Artificial intelligence, bias and clinical safety
Robert Challen, Joshua Denny, Martin Pitt, Luke Gompels, Tom Edwards, and Krasimira Tsaneva-Atanasova · 2019
Earlier work this paper cites.
Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports
Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D Manning · 2019
Earlier work this paper cites.
Conformalized quantile regression
Yaniv Romano, Evan Patterson, and Emmanuel Candes · 2019
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Earlier work this paper cites.
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang · 2020
Earlier work this paper cites.
Consistent estimators for learning to defer to an expert
Hussein Mozannar and David Sontag · 2020
Cited alongside, same era.
Learn then test: Calibrating predictive algorithms to achieve risk control
Anastasios N Angelopoulos, Stephen Bates, Emmanuel J Candès, Michael I Jordan, and Lihua Lei · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Victor Quach, Adam Fisch, Tal Schuster, Adam Yala, Jae Ho Sohn, Tommi S Jaakkola, and Regina Barzilay · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model. arxiv 2023
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Selective classification via neural network training dynamics
Stephan Rabanser, Anvith Thudi, Kimia Hamidieh, Adam Dziedzic, and Nicolas Papernot · 2022
Cited alongside, same era.
Kurt Shuster, Mojtaba Komeili, Leonard Adolphs, Stephen Roller, Arthur Szlam, and Jason Weston · 2022
Cited alongside, same era.
Neeraj Varshney, Swaroop Mishra, and Chitta Baral · 2022
Cited alongside, same era.
Reliable visual question answering: Abstain rather than answer incorrectly
Spencer Whitehead, Suzanne Petryk, Vedaad Shakib, Joseph Gonzalez, Trevor Darrell, Anna Rohrbach, and Marcus Rohrbach · 2022
Cited alongside, same era.
Fairness and machine learning: Limitations and opportunities
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2023
Cited alongside, same era.
Adaptation with self-evaluation to improve selective prediction in llms
Jiefeng Chen, Jinsung Yoon, Sayna Ebrahimi, Sercan O Arik, Tomas Pfister, and Somesh Jha · 2023
Cited alongside, same era.
Allen Z Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Uncertainty-aware language modeling for selective question answering
Qi Yang, Shreya Ravikumar, Fynn Schmitt-Ulms, Satvik Lolla, Ege Demir, Iaroslav Elistratov, Alex Lavaee, Sadhana Lolla, Elaheh Ahmadi, Daniela Rus, et al · 2023
Later among the works it cites.
Selective-lama: Selective prediction for confidence-aware evaluation of language models
Hiyori Yoshikawa and Naoaki Okazaki · 2023
Later among the works it cites.
Evaluating progress in automatic chest x-ray radiology report generation
Feiyang Yu, Mark Endo, Rayan Krishnan, Ian Pan, Andy Tsai, Eduardo Pontes Reis, Eduardo Kaiser Ururahy Nunes Fonseca, Henrique Min Ho Lee, Zahra Shakeri Hossein Abad, Andrew Y Ng, et al · 2023
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2023
Later among the works it cites.
Conformal triage for medical imaging ai deployment
Anastasios Nikolas Angelopoulos, Stuart R Pomerantz, Synho Do, Stephen Bates, Christopher P Bridge, Daniel C Elton, Michael H Lev, R Gilberto Gonzalez, Michael I Jordan, and Jitendra Malik · 2024
Closest in time.
Uncertainty quantification over graph with conformalized graph neural networks
Kexin Huang, Ying Jin, Emmanuel Candes, and Jure Leskovec · 2024
Closest in time.
Uncertainty in language models: Assessment through rank-calibration
Xinmeng Huang, Shuo Li, Mengxin Yu, Matteo Sesia, Hamed Hassani, Insup Lee, Osbert Bastani, and Edgar Dobriban · 2024
Closest in time.
Confidence on the focal: Conformal prediction with selection-conditional coverage
Ying Jin and Zhimei Ren · 2024
Closest in time.
Language models with conformal factuality guarantees
Christopher Mohri and Tatsunori Hashimoto · 2024
Closest in time.
Unintended impacts of llm alignment on global representation
Michael J Ryan, William Held, and Diyi Yang · 2024
Closest in time.
Api is enough: Conformal prediction for large language models without logit-access
Jiayuan Su, Jing Luo, Hongwei Wang, and Lu Cheng · 2024
Closest in time.
Non-exchangeable conformal language generation with nearest neighbors
Dennis Ulmer, Chrysoula Zerva, and André FT Martins · 2024
Closest in time.
Benchmarking llms via uncertainty quantification
Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.