Fetching the paper…
Reading the bibliography…
Models trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weighted-averaging to mismatches in the gradients and propose a new uncertainty-based scheme to improve the performance by reducing the mismatch.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 1903
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
The infinitesimal jackknife
Louis A Jaeckel · 1972
Earlier work this paper cites.
Detection of influential observation in linear regression
R Dennis Cook · 1977
Earlier work this paper cites.
Accurate approximations for posterior moments and marginal densities
Luke Tierney and Joseph B Kadane · 1986
Earlier work this paper cites.
A practical Bayesian framework for backpropagation networks
David JC MacKay · 1992
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
Yann LeCun · 1998
Earlier work this paper cites.
Decentralized estimation and control for multisensor systems
Arthur G. O. Mutambara · 1998
Earlier work this paper cites.
A Bayesian committee machine
Volker Tresp · 2000
Earlier work this paper cites.
Data fusion in decentralised sensing networks
Hugh Durrant-Whyte · 2001
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
The variational Gaussian approximation revisited
Manfred Opper and Cédric Archambeau · 2009
Earlier work this paper cites.
SUN database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer Yuval · 2011
Earlier work this paper cites.
Detection of traffic signs in real-world images: The German Traffic Sign Detection Benchmark
Sebastian Houben, Johannes Stallkamp, Jan Salmen, Marc Schlipsing, and Christian Igel · 2013
Earlier work this paper cites.
3D object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Revisiting natural gradient for deep networks, 2013
Razvan Pascanu and Yoshua Bengio · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Describing Textures in the Wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
Distributed gaussian processes
Marc Deisenroth and Jun Wei Ng · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Character-level Convolutional Networks for Text Classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Cited alongside, same era.
Remote Sensing Image Scene Classification: Benchmark and State of the Art
Gong Cheng, Junwei Han, and Xiaoqiang Lu · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
What is being transferred in transfer learning?
Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Introducing EuroSAT: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth · 2018
Cited alongside, same era.
Note on the quadratic penalties in elastic weight consolidation
Ferenc Huszár · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend · 2021
Later among the works it cites.
OpenCLIP, 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Later among the works it cites.
Knowledge-adaptation priors
Mohammad Emtiyaz Khan and Siddharth Swaroop · 2021
Later among the works it cites.
Composable sparse fine-tuning for cross-lingual transfer
Alan Ansell, Edoardo Ponti, Anna Korhonen, and Ivan Vulić · 2022
Later among the works it cites.
Scaling instruction-finetuned language models, 2022
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei · 2022
Later among the works it cites.
FaithDial: A faithful benchmark for information-seeking dialogue
Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar Zaiane, Mo Yu, Edoardo M. Ponti, and Siva Reddy · 2022
Later among the works it cites.
Revisiting checkpoint averaging for neural machine translation
Yingbo Gao, Christian Herold, Zijian Yang, and Hermann Ney · 2022
Later among the works it cites.
Merging models with Fisher-weighted averaging
Michael S Matena and Colin Raffel · 2022
Later among the works it cites.
Bayesian data fusion with shared priors, 2022
Peng Wu, Tales Imbiriba, Victor Elvira, and Pau Closas · 2022
Later among the works it cites.
Git Re-Basin: Merging models modulo permutation symmetries
Samuel Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2023
Closest in time.
Elastic weight removal for faithful and abstractive dialogue generation, 2023
Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, and Edoardo M. Ponti · 2023
Closest in time.
Detoxify
Laura Hanu and Unitary team · 2023
Closest in time.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Closest in time.
Dataless knowledge fusion by merging weights of language models
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng · 2023
Closest in time.
The Flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, and Adam Roberts · 2023
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models, 2023
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard · 2023
Closest in time.
GPT-J-6B: a 6 billion parameter autoregressive language model, 2021
Ben Wang and Aran Komatsuzaki · 2023
Closest in time.
Resolving interference when merging models, 2023
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal · 2023
Closest in time.