Fetching the paper…
Reading the bibliography…
Pre-trained language models (PLMs) may fail in giving reliable estimates of their predictive uncertainty.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Scaling out-of-distribution detection for real-world settings
Dan Hendrycks, Steven Basart, Mantas Mazeika, Mohammadreza Mostajabi, Jacob Steinhardt, and Dawn Song. 2019a · 1911
Earlier work this paper cites.
Distributional structure
Zellig S Harris. 1954 · 1954
Earlier work this paper cites.
A statistical approach to mechanized encoding and searching of literary information
Hans Peter Luhn. 1957 · 1957
Earlier work this paper cites.
Knowing more about questions can help: Improving calibration in question answering
Shujian Zhang, Chengyue Gong, and Eunsol Choi. 2021b · 1970
Earlier work this paper cites.
The comparison and evaluation of forecasters
Morris H DeGroot and Stephen E Fienberg. 1983 · 1983
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt et al. 1999 · 1999
Earlier work this paper cites.
From amateurs to connoisseurs: modeling the evolution of user expertise through online reviews
Julian J. McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
SemEval-2013 task 2: Sentiment analysis in Twitter
Preslav Nakov, Sara Rosenthal, Zornitsa Kozareva, Veselin Stoyanov, Alan Ritter, and Theresa Wilson. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013a · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013b · 2013
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017 · 2017
Earlier work this paper cites.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey E. Hinton. 2017 · 2017
Earlier work this paper cites.
Hate speech dataset from a white supremacy forum
Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018a · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018b · 2018
Cited alongside, same era.
Improved zero-shot neural machine translation via ignoring spurious correlations
Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor O. K. Li. 2019 · 2019
Cited alongside, same era.
Using pre-training can improve model robustness and uncertainty
Dan Hendrycks, Kimin Lee, and Mantas Mazeika. 2019b · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Measuring calibration in deep learning
DynaSent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah. 2021 · 2021
Later among the works it cites.
Augmax: Adversarial composition of random augmentations for robust training
Haotao Wang, Chaowei Xiao, Jean Kossaifi, Zhiding Yu, Anima Anandkumar, and Zhangyang Wang. 2021 · 2021
Later among the works it cites.
Bridge the gap between CV and nlp! A gradient-based textual adversarial attack framework
Lifan Yuan, Yichi Zhang, Yangyi Chen, and Wei Wei. 2021 · 2021
Later among the works it cites.
Why should adversarial perturbations be imperceptible? rethink the research paradigm in adversarial NLP
Yangyi Chen, Hongcheng Gao, Ganqu Cui, Fanchao Qi, Longtao Huang, Zhiyuan Liu, and Maosong Sun. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeremy Nixon, Michael W. Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. 2019 · 2019
Cited alongside, same era.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Jasper Snoek, Yaniv Ovadia, Emily Fertig, Balaji Lakshminarayanan, Sebastian Nowozin, D. Sculley, Joshua V. Dillon, Jie Ren, and Zachary Nado. 2019 · 2019
Cited alongside, same era.
Evaluating model calibration in classification
Juozas Vaicenavicius, David Widmann, Carl R. Andersson, Fredrik Lindsten, Jacob Roll, and Thomas B. Schön. 2019 · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
EDA: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Closest in time.
A unified evaluation of textual backdoor learning: Frameworks and benchmarks
Ganqu Cui, Lifan Yuan, Bingxiang He, Yangyi Chen, Zhiyuan Liu, and Maosong Sun. 2022 · 2022
Closest in time.
Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al. 2022 · 2022
Closest in time.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022 · 2022
Closest in time.
Calibrated ensembles can mitigate accuracy tradeoffs under distribution shift
Ananya Kumar, Tengyu Ma, Percy Liang, and Aditi Raghunathan. 2022 · 2022
Closest in time.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Closest in time.
A survey on backdoor attack and defense in natural language processing
Xuan Sheng, Zhaoyang Han, Piji Li, and Xiangmao Chang. 2022 · 2022
Closest in time.
Revisiting calibration for question answering
Chenglei Si, Chen Zhao, Sewon Min, and Jordan Boyd-Graber. 2022 · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022 · 2022
Closest in time.
Exploring predictive uncertainty and calibration in NLP: A study on the impact of method & data scarcity
Dennis Ulmer, Jes Frellsen, and Christian Hardmeier. 2022 · 2022
Closest in time.
Neeraj Varshney, Swaroop Mishra, and Chitta Baral. 2022 · 2022
Closest in time.
Identifying and mitigating spurious correlations for improving robustness in NLP models
Tianlu Wang, Rohit Sridhar, Diyi Yang, and Xuezhi Wang. 2022 · 2022
Closest in time.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022 · 2022
Closest in time.
Can explanations be useful for calibrating black box models?
Xi Ye and Greg Durrett. 2022 · 2022
Closest in time.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. 2022 · 2022
Closest in time.
Interpreting the robustness of neural NLP models to textual perturbations
Yunxiang Zhang, Liangming Pan, Samson Tan, and Min-Yen Kan. 2022 · 2022
Closest in time.
On the effects of transformer size on in- and out-of-domain calibration
Soham Dan and Dan Roth. 2021 · 2096
Closest in time.