Fetching the paper…
Reading the bibliography…
Solving multi-objective optimization problems for large deep neural networks is a challenging task due to the complexity of the loss landscape and the expensive computational cost of training and evaluating models.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing, July 2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 1910
Earlier work this paper cites.
Linear Mode Connectivity and the Lottery Ticket Hypothesis, July 2020
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin · 1912
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 1912
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Le’on Bottou, Yoshua Bengio, and Parick Haffner · 1998
Earlier work this paper cites.
A fast and elitist multiobjective genetic algorithm: NSGA-II
K. Deb, A. Pratap, S. Agarwal, and T. Meyarivan · 2002
Earlier work this paper cites.
Stochastic Method for the Solution of Unconstrained Vector Optimization Problems
S. Schäffler, R. Schultz, and K. Weinzierl · 2002
Earlier work this paper cites.
Convex Optimization
Stephen P. Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
SMS-EMOA: Multiobjective selection based on dominated hypervolume
Nicola Beume, Boris Naujoks, and Michael Emmerich · 2006
Earlier work this paper cites.
MOEA/D: A Multiobjective Evolutionary Algorithm Based on Decomposition
Qingfu Zhang and Hui Li · 2007
Earlier work this paper cites.
Learning the Pareto Front with Hypernetworks, April 2021
Aviv Navon, Aviv Shamsian, Gal Chechik, and Ethan Fetaya · 2010
Earlier work this paper cites.
SUN database: Large-scale scene recognition from abbey to zoo
Jianxiong Xiao, James Hays, Krista A. Ehinger, Aude Oliva, and Antonio Torralba · 2010
Earlier work this paper cites.
Multiple-gradient descent algorithm (MGDA) for multiobjective optimization
Jean-Antoine Désidéri · 2012
Earlier work this paper cites.
Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition
J. Stallkamp, M. Schlipsing, J. Salmen, and C. Igel · 2012
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Visualizing and Understanding Convolutional Networks, November 2013
Matthew D. Zeiler and Rob Fergus · 2013
Earlier work this paper cites.
Search Methodologies: Introductory Tutorials in Optimization and Decision Support Techniques
Edmund K. Burke and Graham Kendall, editors · 2014
Earlier work this paper cites.
Describing Textures in the Wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
Remote Sensing Image Scene Classification: Benchmark and State of the Art
Gong Cheng, Junwei Han, and Xiaoqiang Lu · 2017
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization: 5th International Conference on Learning Representations, ICLR 2017
C. Daniel Freeman and Joan Bruna · 2017
Earlier work this paper cites.
An Overview of Multi-Task Learning in Deep Neural Networks, June 2017
Sebastian Ruder · 2017
Earlier work this paper cites.
Introducing eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth · 2018
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
Essentially No Barriers in Neural Network Energy Landscape, February 2019
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A. Hamprecht · 2019
Cited alongside, same era.
Pareto Multi-Task Learning
Xi Lin, Hui-Ling Zhen, Zhenhua Li, Qing-Fu Zhang, and Sam Kwong · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
On Convexity and Linear Mode Connectivity in Neural Networks
David Yunis, Kumar Kshitij Patel, Pedro Savarese, Gal Vardi, Karen Livescu, Matthew Walter, Jonathan Frankle, and Michael Maire · 2022
Later among the works it cites.
Git Re-Basin: Merging Models modulo Permutation Symmetries, March 2023
Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2023
Later among the works it cites.
AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models, March 2023
Alexandra Chronopoulou, Matthew E. Peters, Alexander Fraser, and Jesse Dodge · 2023
Later among the works it cites.
Pareto Manifold Learning: Tackling multiple tasks via ensembles of single-task models
Nikolaos Dimitriadis, Pascal Frossard, and François Fleuret · 2023
Later among the works it cites.
ZipIt! Merging Models from Different Tasks without Training, May 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ozan Sener and Vladlen Koltun · 2019
Cited alongside, same era.
Towards Impartial Multi-task Learning
Liyang Liu, Yi Li, Zhanghui Kuang, Jing-Hao Xue, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang · 2020
Cited alongside, same era.
Multi-Task Learning with User Preferences: Gradient Descent with Controlled Ascent in Pareto Optimization
Debabrata Mahapatra and Vaibhav Rajan · 2020
Cited alongside, same era.
Optimizing Mode Connectivity via Neuron Alignment
Norman Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk, Prasanna Sattigeri, and Rongjie Lai · 2020
Cited alongside, same era.
Loss Surface Simplexes for Mode Connecting Volumes and Fast Ensembling, November 2021
Gregory W. Benton, Wesley J. Maddox, Sanae Lotfi, and Andrew Gordon Wilson · 2021
Cited alongside, same era.
Openclip, July 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Cited alongside, same era.
Reading Digits in Natural Images with Unsupervised Feature Learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2021
Cited alongside, same era.
George Stoica, Daniel Bolya, Jakob Bjorner, Taylor Hearn, and Judy Hoffman · 2023
Later among the works it cites.
Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models, May 2023
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard · 2023
Later among the works it cites.
Improving Pareto Front Learning via Multi-Sample Hypernetworks
Long P. Hoang, Dung D. Le, Tran Anh Tuan, and Tran Ngoc Thang · 2023
Later among the works it cites.
Editing Models with Task Arithmetic, March 2023
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Later among the works it cites.
Dataless Knowledge Fusion by Merging Weights of Language Models, April 2023
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng · 2023
Later among the works it cites.
Merging Decision Transformers: Weight Averaging for Forming Multi-Task Policies, September 2023
Daniel Lawson and Ahmed H. Qureshi · 2023
Later among the works it cites.
Deep Model Fusion: A Survey, September 2023
Weishi Li, Yong Peng, Miao Zhang, Liang Ding, Han Hu, and Li Shen · 2023
Later among the works it cites.
Sunny Sanyal, Jean Kaddour, Abhishek Kumar, and Sujay Sanghavi · 2023
Later among the works it cites.
Concrete Subspace Learning based Interference Elimination for Multi-task Model Fusion, December 2023
Anke Tang, Li Shen, Yong Luo, Liang Ding, Han Hu, Bo Du, and Dacheng Tao · 2023
Later among the works it cites.
Resolving Interference When Merging Models, June 2023
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal · 2023
Later among the works it cites.
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li · 2023
Later among the works it cites.
Hypervolume Maximization: A Geometric View of Pareto Set Learning
Xiaoyuan Zhang, Xi Lin, Bo Xue, Yifan Chen, and Qingfu Zhang · 2023
Later among the works it cites.
Learn From Model Beyond Fine-Tuning: A Survey, October 2023
Hongling Zheng, Li Shen, Anke Tang, Yong Luo, Han Hu, Bo Du, and Dacheng Tao · 2023
Later among the works it cites.
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts, February 2024
Anke Tang, Li Shen, Yong Luo, Nan Yin, Lefei Zhang, and Dacheng Tao · 2024
Closest in time.
Representation Surgery for Multi-Task Model Merging, February 2024
Enneng Yang, Li Shen, Zhenyi Wang, Guibing Guo, Xiaojun Chen, Xingwei Wang, and Dacheng Tao · 2024
Closest in time.