Fetching the paper…
Reading the bibliography…
In this paper, we introduce a novel approach for addressing the multi-objective optimization problem in large language model merging via black-box multi-objective optimization algorithms.
P.-L. Yu, “Cone convexity, cone extreme points, and nondominated solutions in decision problems with multiobjectives,” Journal of optimization Theory and Applications , vol. 14, pp. 319–377, 1974
1974
Earlier work this paper cites.
G. Brown, “Diversity in neural network ensembles,” Ph.D. dissertation, Citeseer, 2004
2004
Earlier work this paper cites.
E. K. Tang, P. N. Suganthan, and X. Yao, “An analysis of diversity measures,” Machine learning , vol. 65, pp. 247–271, 2006
2006
Earlier work this paper cites.
A. Konak, D. W. Coit, and A. E. Smith, “Multi-objective optimization using genetic algorithms: A tutorial,” Reliability engineering & system safety , vol. 91, no. 9, pp. 992–1007, 2006
2006
Earlier work this paper cites.
M. T. Emmerich, K. C. Giannakoglou, and B. Naujoks, “Single-and multiobjective evolutionary optimization assisted by gaussian random field metamodels,” IEEE Transactions on Evolutionary Computation , vol. 10, no. 4, pp. 421–439, 2006
2006
Earlier work this paper cites.
N. Hansen, “The cma evolution strategy: a comparing review,” Towards a new evolutionary computation: Advances in the estimation of distribution algorithms , pp. 75–102, 2006
2006
Earlier work this paper cites.
X. Yao and M. M. Islam, “Evolving artificial neural network ensembles,” IEEE Computational Intelligence Magazine , vol. 3, no. 1, pp. 31–42, 2008
2008
Earlier work this paper cites.
S. Das and P. N. Suganthan, “Differential evolution: A survey of the state-of-the-art,” IEEE transactions on evolutionary computation , vol. 15, no. 1, pp. 4–31, 2010
2010
Earlier work this paper cites.
T. White, “Sampling generative networks,” arXiv preprint arXiv:1609.04468 , 2016
2016
Earlier work this paper cites.
A. Ly, M. Marsman, J. Verhagen, R. P. Grasman, and E.-J. Wagenmakers, “A tutorial on fisher information,” Journal of Mathematical Psychology , vol. 80, pp. 40–55, 2017
2017
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Daulton, M. Balandat, and E. Bakshy, “Differentiable expected hypervolume improvement for parallel multi-objective bayesian optimization,” Advances in Neural Information Processing Systems , vol. 33, pp. 9851–9864, 2020
2020
Earlier work this paper cites.
K. Shang and H. Ishibuchi, “A new hypervolume-based evolutionary algorithm for many-objective optimization,” IEEE Transactions on Evolutionary Computation , vol. 24, no. 5, pp. 839–852, 2020
2020
Earlier work this paper cites.
M. Balandat, B. Karrer, D. Jiang, S. Daulton, B. Letham, A. G. Wilson, and E. Bakshy, “Botorch: A framework for efficient monte-carlo bayesian optimization,” Advances in neural information processing systems , vol. 33, pp. 21 524–21 538, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba, “Evaluating large language models trained on code,” 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi, “Winogrande: An adversarial winograd schema challenge at scale,” Communications of the ACM , vol. 64, no. 9, pp. 99–106, 2021
2021
Earlier work this paper cites.
S. Daulton, M. Balandat, and E. Bakshy, “Parallel bayesian optimization of multiple noisy objectives with expected hypervolume improvement,” Advances in Neural Information Processing Systems , vol. 34, pp. 2187–2200, 2021
2021
Cited alongside, same era.
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith et al. , “Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,” in International conference on machine learning . PMLR, 2022, pp. 23 965–23 998
2022
Cited alongside, same era.
2022
Cited alongside, same era.
M. S. Matena and C. A. Raffel, “Merging models with fisher-weighted averaging,” Advances in Neural Information Processing Systems , vol. 35, pp. 17 703–17 716, 2022
2023
Later among the works it cites.
O. Contributors, “Opencompass: A universal evaluation platform for foundation models,” https://github.com/open-compass/opencompass
2023
Later among the works it cites.
W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P. Luo, T. Lu, J. Zhou, Y. Qiao et al. , “Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Ivison, Y. Wang, V. Pyatkin, N. Lambert, M. Peters, P. Dasigi, J. Jang, D. Wadden, N. A. Smith, I. Beltagy, and H. Hajishirzi, “Camels in a changing climate: Enhancing lm adaptation with tulu 2,” 2023
2023
Cited alongside, same era.
H. Xu, Y. J. Kim, A. Sharaf, and H. H. Awadalla, “A paradigm shift in machine translation: Boosting translation performance of large language models,” 2023
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
R. Ries. (2024) Power rangers & model merging technology. [Online]. Available: https://www.missioncloud.com/blog/power-rangers-model-merging-technology
2024
Closest in time.
C. O. Goddard. (2024) mergekit. [Online]. Available: https://github.com/arcee-ai/mergekit
2024
Closest in time.
M. Labonne. (2024) Merge large language models with mergekit. [Online]. Available: https://huggingface.co/blog/mlabonne/merge-models
2024
Closest in time.
L. Yu, B. Yu, H. Yu, F. Huang, and Y. Li, “Language models are super mario: Absorbing abilities from homologous models as a free lunch,” in Forty-first International Conference on Machine Learning , 2024
2024
Closest in time.
P. Yadav, D. Tam, L. Choshen, C. A. Raffel, and M. Bansal, “Ties-merging: Resolving interference when merging models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
abacusai. (2024) abacusai/liberated-qwen1.5-7b. [Online]. Available: https://huggingface.co/abacusai/Liberated-Qwen1.5-7B
2024
Closest in time.
YeungNLP. (2024) Yeungnlp/firefly-qwen1.5-en-7b. [Online]. Available: https://huggingface.co/YeungNLP/firefly-qwen1.5-en-7b
2024
Closest in time.
“Qwen2 technical report,” 2024
2024
Closest in time.
VAGOsolutions. (2024) Vagosolutions/sauerkrautlm-1.5b. [Online]. Available: https://huggingface.co/VAGOsolutions/SauerkrautLM-1.5b
2024
Closest in time.
DeepMount00. (2024) Deepmount00/qwen2-1.5b-ita. [Online]. Available: https://huggingface.co/DeepMount00/Qwen2-1.5B-Ita
2024
Closest in time.
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, Q. Lin, and D. Jiang, “Wizardlm: Empowering large pre-trained language models to follow complex instructions,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.