Fetching the paper…
Reading the bibliography…
As machine learning models continue to swiftly advance, calibrating their performance has become a major concern prior to practical and widespread implementation.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant. 1984 · 1984
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt et al. 1999 · 1999
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers
Bianca Zadrozny and Charles Elkan. 2001 · 2001
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
Deep entity matching with pre-trained language models
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan. 2020 · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
Use of brier score to assess binary predictions
Kaspar Rufibach. 2010 · 2010
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Algorithmic accountability: A primer
Joan Donovan, Robyn Caplan, Jeanna Matthews, and Lauren Hanson. 2018 · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018 · 2018
Cited alongside, same era.
Pretrained language models for sequential sentence classification
Arman Cohan, Iz Beltagy, Daniel King, Bhavana Dalvi, and Dan Weld. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Measuring calibration in deep learning
Jeremy Nixon, Michael W Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. 2019 · 2019
Cited alongside, same era.
Automatically identifying complaints in social media
Daniel Preoţiuc-Pietro, Mihaela Gaman, and Nikolaos Aletras. 2019 · 2019
Cited alongside, same era.
On mixup training: Improved calibration and predictive uncertainty for deep neural networks
Accountability in an algorithmic society: relationality, responsibility, and robustness in machine learning
A Feder Cooper, Emanuel Moss, Benjamin Laufer, and Helen Nissenbaum. 2022 · 2022
Later among the works it cites.
Ddxplus: A new dataset for automatic medical diagnosis
Arsene Fansi Tchango, Rishab Goel, Zhi Wen, Julien Martel, and Joumana Ghosn. 2022 · 2022
Later among the works it cites.
Protoformer: Embedding prototypes for transformers
Ashkan Farhangi, Ning Sui, Nan Hua, Haiyan Bai, Arthur Huang, and Zhishan Guo. 2022 · 2022
Later among the works it cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022 · 2022
Later among the works it cites.
Net benefit, calibration, threshold selection, and training objectives for algorithmic fairness in healthcare
Stephen Pfohl, Yizhe Xu, Agata Foryciarz, Nikolaos Ignatiadis, Julian Genkins, and Nigam Shah. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sunil Thulasidasan, Gopinath Chennupati, Jeff A Bilmes, Tanmoy Bhattacharya, and Sarah Michalak. 2019 · 2019
Cited alongside, same era.
Efficient intent detection with dual sentence encoders
Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić. 2020 · 2020
Cited alongside, same era.
Revisiting the calibration of modern neural networks
Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, and Mario Lucic. 2021 · 2021
Cited alongside, same era.
Faircal: Fairness calibration for face verification
Tiago Salvador, Stephanie Cairns, Vikram Voleti, Noah Marshall, and Adam Oberman. 2021 · 2021
Cited alongside, same era.
Rethinking calibration of deep neural networks: Do not be afraid of overconfidence
Deng-Bao Wang, Lei Feng, and Min-Ling Zhang. 2021 · 2021
Cited alongside, same era.
Combining ensembles and data augmentation can harm your calibration
Yeming Wen, Ghassen Jerfel, Rafael Muller, Michael W Dusenberry, Jasper Snoek, Balaji Lakshminarayanan, and Dustin Tran. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Synthetic data generation with large language models for text classification: Potential and limitations
Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, and Ming Yin. 2023 · 2023
Later among the works it cites.
PromptMix: A class boundary augmentation method for large language model distillation
Gaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, and Issam Laradji. 2023 · 2023
Later among the works it cites.
Synthetic data for model selection
Alon Shoshan, Nadav Bhonker, Igor Kviatkovsky, Matan Fintz, and Gérard Medioni. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Synthetic data, real errors: how (not) to publish and use synthetic data
Boris Van Breugel, Zhaozhi Qian, and Mihaela Van Der Schaar. 2023 · 2023
Later among the works it cites.
Can you rely on your model evaluation? Improving model evaluation with synthetic test data
Boris van Breugel, Nabeel Seedat, Fergus Imrie, and Mihaela van der Schaar. 2023 · 2023
Later among the works it cites.