Fetching the paper…
Reading the bibliography…
Pre-trained Transformers are now ubiquitous in natural language processing, but despite their high end-task performance, little is known empirically about whether they are calibrated.
How to Fine-Tune BERT for Text Classification?
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang. 2019 · 1905
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Verification of Forecasts Expressed in Terms of Probability
Glenn W. Brier. 1950 · 1950
Earlier work this paper cites.
A Global Optimization Technique for Statistical Classifier Design
David J. Miller, Ajit V. Rao, Kenneth Rose, and Allen Gersho. 1996 · 1996
Earlier work this paper cites.
Are Artificial Neural Networks Black Boxes?
Jose M. Benítez, Juan Luis Castro, and Ignacio Requena. 1997 · 1997
Earlier work this paper cites.
Artificial Neural Networks: Opening the Black Box
Judith E. Dayhoff and James M. DeLeo. 2001 · 2001
Earlier work this paper cites.
Using Bayesian Model Averaging to Calibrate Forecast Ensembles
Adrian E. Raftery, Tilmann Gneiting, Fadoua Balabdaoui, and Michael Polakowski. 2005 · 2005
Earlier work this paper cites.
Probabilistic Forecasts, Calibration and Sharpness
Tilmann Gneiting, Fadoua Balabdaoui, and Adrian E. Raftery. 2007 · 2007
Earlier work this paper cites.
Toward Seamless Prediction: Calibration of Climate Change Projections using Seasonal Forecasts
Tim Palmer, Francisco Doblas-Reyes, Antje Weisheimer, and Mark Rodwell. 2008 · 2008
Earlier work this paper cites.
Nurses’ Risk Assessment Judgements: A Confidence Calibration Study
Huiqin Yang and Carl Thompson. 2010 · 2010
Earlier work this paper cites.
Calibrating Predictive Model Estimates to Support Personalized Medicine
Xiaoqian Jiang, Melanie Osl, Jihoon Kim, and Lucila Ohno-Machado. 2012 · 2012
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
A Large Annotated Corpus for Learning Natural Language Inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Posterior Calibration and Exploratory Analysis for Natural Language Processing Models
Khanh Nguyen and Brendan O’Connor. 2015 · 2015
Earlier work this paper cites.
Can We Open the Black Box of AI?
Davide Castelvecchi. 2016 · 2016
Cited alongside, same era.
A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Cited alongside, same era.
A Decomposable Attention Model for Natural Language Inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Cited alongside, same era.
Enhanced LSTM for Natural Language Inference
Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, Si Wei, Hui Jiang, and Diana Inkpen. 2017 · 2017
Cited alongside, same era.
On Calibration of Modern Neural Networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Cited alongside, same era.
Quora Question Pairs
Shankar Iyer, Nikhil Dandekar, and Kornél Csernai. 2017 · 2017
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Later among the works it cites.
What Does BERT Look at? An Analysis of BERT’s Attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Evaluating Lottery Tickets Under Distributional Shifts
Shrey Desai, Hongyuan Zhan, and Ahmed Aly. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Revealing the Dark Secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
Alex Kendall and Yarin Gal. 2017 · 2017
Cited alongside, same era.
A Continuously Growing Dataset of Sentential Paraphrases
Wuwei Lan, Siyu Qiu, Hua He, and Wei Xu. 2017 · 2017
Cited alongside, same era.
Regularizing Neural Networks by Penalizing Confident Output Distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton. 2017 · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Adversarial Deep Averaging Networks for Cross-Lingual Sentiment Classification
Xilun Chen, Yu Sun, Ben Athiwaratkun, Claire Cardie, and Kilian Weinberger. 2018 · 2018
Cited alongside, same era.
AllenNLP: A Deep Semantic Natural Language Processing Platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Calibration of Encoder Decoder Models for Neural Machine Translation
Aviral Kumar and Sunita Sarawagi. 2019 · 2019
Later among the works it cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Later among the works it cites.
Simplified Neural Unsupervised Domain Adaptation
Timothy Miller. 2019 · 2019
Later among the works it cites.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R. Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
HellaSWAG: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Closest in time.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.