Fetching the paper…
Reading the bibliography…
Fine-tuned Large Language Models (LLMs) often suffer from overconfidence and poor calibration, particularly when fine-tuned on small datasets.
A practical Bayesian framework for backpropagation networks
David J. C. MacKay · 1992
Earlier work this paper cites.
Bayesian learning via stochastic dynamics
Radford M Neal · 1993
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
MCMC using Hamiltonian dynamics
Radford M Neal et al · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Diederik P Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Structured and efficient variational deep learning with matrix Gaussian posteriors
Christos Louizos and Max Welling · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Fast and scalable Bayesian deep learning by weight-perturbation in Adam
Mohammad Khan, Didrik Nielsen, Voot Tangkaratt, Wu Lin, Yarin Gal, and Akash Srivastava · 2018
Earlier work this paper cites.
Enhancing the reliability of out-of-distribution image detection in neural networks
Shiyu Liang, Yixuan Li, and R. Srikant · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi · 2018
Earlier work this paper cites.
Deep ensembles: A loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Approximate inference turns deep networks into Gaussian processes
Mohammad Emtiyaz Khan, Alexander Immer, Ehsan Abedi, and Maciej Korzepa · 2019
Earlier work this paper cites.
A simple baseline for bayesian uncertainty in deep learning
Wesley J Maddox, Pavel Izmailov, Timur Garipov, Dmitry P Vetrov, and Andrew Gordon Wilson · 2019
Earlier work this paper cites.
Practical deep learning with Bayesian principles
Kazuki Osawa, Siddharth Swaroop, Mohammad Emtiyaz E Khan, Anirudh Jain, Runa Eschenhagen, Richard E Turner, and Rio Yokota · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Bayesian layers: A module for neural network uncertainty
Dustin Tran, Mike Dusenberry, Mark Van Der Wilk, and Danijar Hafner · 2019
Earlier work this paper cites.
Function space particle optimization for bayesian neural networks
Ziyu Wang, Tongzheng Ren, Jun Zhu, and Bo Zhang · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Conservative uncertainty estimation by fitting prior networks
Kamil Ciosek, Vincent Fortuin, Ryota Tomioka, Katja Hofmann, and Richard Turner · 2020
Cited alongside, same era.
Bayesian deep ensembles via the neural tangent kernel
Bobby He, Balaji Lakshminarayanan, and Yee Whye Teh · 2020
Cited alongside, same era.
Subspace inference for bayesian deep learning
Pavel Izmailov, Wesley J Maddox, Polina Kirichenko, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2020
Cited alongside, same era.
Probing as quantifying inductive bias
Alexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, and Ryan Cotterell · 2022
Later among the works it cites.
Hands-on bayesian neural networks - a tutorial for deep learning users
Laurent Valentin Jospin, Hamid Laga, Farid Boussaid, Wray Buntine, and Mohammed Bennamoun · 2022
Later among the works it cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Later among the works it cites.
Calibrated ensembles can mitigate accuracy tradeoffs under distribution shift
Ananya Kumar, Tengyu Ma, Percy Liang, and Aditi Raghunathan · 2022
Later among the works it cites.
Peft: State-of-the-art parameter-efficient fine-tuning methods
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ensemble distribution distillation
Andrey Malinin, Bruno Mlodozeniec, and Mark J. F. Gales · 2020
Cited alongside, same era.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew Gordon Wilson and Pavel Izmailov · 2020
Cited alongside, same era.
Cyclical stochastic gradient mcmc for bayesian deep learning
Ruqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen, and Andrew Gordon Wilson · 2020
Cited alongside, same era.
Repulsive deep ensembles are Bayesian
Francesco D’Angelo and Vincent Fortuin · 2021
Cited alongside, same era.
On stein variational neural network ensembles
Francesco D’Angelo, Vincent Fortuin, and Florian Wenzel · 2021
Cited alongside, same era.
Laplace Redux - Effortless Bayesian Deep Learning
Erik Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Matthias Bauer, and Philipp Hennig · 2021
Cited alongside, same era.
Bnnpriors: A library for bayesian neural network inference with different prior distributions
Vincent Fortuin, Adrià Garriga-Alonso, Mark van der Wilk, and Laurence Aitchison · 2021
Cited alongside, same era.
Data augmentation in bayesian neural networks and the cold posterior effect
Seth Nabarro, Stoil Ganev, Adrià Garriga-Alonso, Vincent Fortuin, Mark van der Wilk, and Laurence Aitchison · 2022
Later among the works it cites.
Pac-bayesian meta-learning: From theory to practice
Jonas Rothfuss, Martin Josifoski, Vincent Fortuin, and Andreas Krause · 2022
Later among the works it cites.
Last layer marginal likelihood for invariance learning
Pola Schwöbel, Martin Jørgensen, Sebastian W Ober, and Mark Van Der Wilk · 2022
Later among the works it cites.
Quantifying uncertainty in foundation models via ensembles
Meiqi Sun, Wilson Yan, Pieter Abbeel, and Igor Mordatch · 2022
Later among the works it cites.
Learning invariant weights in neural networks
Tycho FA van der Ouderaa and Mark van der Wilk · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le · 2022
Later among the works it cites.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2022
Later among the works it cites.
One-for-all: Generalized lora for parameter-efficient fine-tuning
Arnav Chavan, Zhuang Liu, Deepak Gupta, Eric Xing, and Zhiqiang Shen · 2023
Later among the works it cites.
Calibrating transformers via sparse gaussian processes
Wenlong Chen and Yingzhen Li · 2023
Later among the works it cites.
Longlora: Efficient fine-tuning of long-context large language models
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Later among the works it cites.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al · 2023
Later among the works it cites.
Preserving pre-trained features helps calibrate fine-tuned language models
Guande He, Jianfei Chen, and Jun Zhu · 2023
Later among the works it cites.
Gpt-4 technical report. arxiv 2303.08774
R OpenAI · 2023
Later among the works it cites.
Beyond deep ensembles: A large-scale evaluation of bayesian deep learning under distribution shift
Florian Seligmann, Philipp Becker, Michael Volpp, and Gerhard Neumann · 2023
Later among the works it cites.
Incorporating unlabelled data into bayesian neural networks
Mrinank Sharma, Tom Rainforth, Yee Whye Teh, and Vincent Fortuin · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Later among the works it cites.
Lora ensembles for large language model fine-tuning
Xi Wang, Laurence Aitchison, and Maja Rudolph · 2023
Later among the works it cites.
Bayesian low-rank adaptation for large language models
Adam Yang, Maxime Robeyns, Xi Wang, and Laurence Aitchison · 2023
Later among the works it cites.
Uncertainty quantification in fine-tuned llms using lora ensembles, 2024
Oleksandr Balabanov and Hampus Linander · 2024
Closest in time.
Position paper: Bayesian deep learning in the age of large-scale AI
Theodore Papamarkou, Maria Skoularidou, Konstantina Palla, Laurence Aitchison, Julyan Arbel, David Dunson, Maurizio Filippone, Vincent Fortuin, Philipp Hennig, Aliaksandr Hubin, Alexander Immer, Theofanis Karaletsos, Mohammad Emtiyaz Khan, Agustinus Kristiadi, Yingzhen Li, Stephan Mandt, Christopher Nemeth, Michael A Osborne, Tim GJ Rudner, David Rügamer, Yee Whye Teh, Max Welling, Andrew Gordon Wilson, and Ruqi Zhang · 2024
Closest in time.
Benchmarking llms via uncertainty quantification, 2024
Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.