Fetching the paper…
Reading the bibliography…
We propose Model Swarms, a collaborative search algorithm to adapt LLMs via swarm intelligence, the collective behavior guiding individual systems.
An overview of evolutionary algorithms for parameter optimization
Bäck, T. and Schwefel, H.-P · 1993
Earlier work this paper cites.
Particle swarm optimization
Kennedy, J. and Eberhart, R · 1995
Earlier work this paper cites.
Mixture of experts: a literature survey
Masoudnia, S. and Ebrahimpour, R · 2014
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Genetic algorithm-a literature review
Lambora, A., Gupta, K., and Chopra, K · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B. et al · 2020
Earlier work this paper cites.
Classifying the classifier: dissecting the weight space of neural networks
Eilertsen, G., Jönsson, D., Ropinski, T., Unger, J., and Ynnerman, A · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
Wadden, D., Lin, S., Lo, K., Wang, L. L., van Zuylen, M., Cohan, A., and Hajishirzi, H · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P · 2021
Earlier work this paper cites.
Beyond distillation: Task-level mixture-of-experts for efficient inference
Kudugunta, S., Huang, Y., Bapna, A., Krikun, M., Lepikhin, D., Luong, M.-T., and Firat, O · 2021
Earlier work this paper cites.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L · 2021
Earlier work this paper cites.
Hash layers for large sparse models
Roller, S., Sukhbaatar, S., Weston, J., et al · 2021
Earlier work this paper cites.
Stablemoe: Stable routing strategy for mixture of experts
Dai, D., Dong, L., Ma, S., Zheng, B., Sui, Z., Chang, B., and Wei, F · 2022
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Earlier work this paper cites.
Demix layers: Disentangling domains for modular language modeling
Gururangan, S., Lewis, M., Holtzman, A., Smith, N. A., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2022
Earlier work this paper cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
Li, M., Gururangan, S., Dettmers, T., Lewis, M., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., Umapathi, L. K., and Sankarasubbu, M · 2022
Earlier work this paper cites.
Lifting the curse of multilinguality by pre-training modular transformers
Pfeiffer, J., Goyal, N., Lin, X., Li, X., Cross, J., Riedel, S., and Artetxe, M · 2022
Earlier work this paper cites.
Diverse weight averaging for out-of-distribution generalization
Rame, A., Kirchmeyer, M., Rahier, T., Rakotomamonjy, A., Gallinari, P., and Cord, M · 2022
Earlier work this paper cites.
Mixture of attention heads: Selecting attention heads per token
Zhang, X., Shen, Y., Huang, Z., Zhou, J., Rong, W., and Xiong, Z · 2022
Earlier work this paper cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Le, Q. V., Laudon, J., et al · 2022
Earlier work this paper cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S., Hayase, J., and Srinivasa, S · 2023
Earlier work this paper cites.
Code alpaca: An instruction-following llama model for code generation
Chaudhary, S · 2023
Earlier work this paper cites.
Model breadcrumbs: Scaling multi-task model merging with sparse masks
Davari, M. and Belilovsky, E · 2023
Earlier work this paper cites.
From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair nlp models
Feng, S., Park, C. Y., Liu, Y., and Tsvetkov, Y · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Earlier work this paper cites.
Koala: A dialogue model for academic research
Geng, X., Gudibande, A., Liu, H., Wallace, E., Abbeel, P., Levine, S., and Song, D · 2023
Cited alongside, same era.
Scaling expert language models with unsupervised domain discovery
Gururangan, S., Li, M., Lewis, M., Shi, W., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2023
Cited alongside, same era.
Lorahub: Efficient cross-task generalization via dynamic lora composition
Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M · 2023
Cited alongside, same era.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Cited alongside, same era.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
InternLM Team, I · 2023
Cited alongside, same era.
Arcee’s mergekit: A toolkit for merging large language models
Goddard, C., Siriwardhana, S., Ehghaghi, M., Meyers, L., Karpukhin, V., Benedict, B., McQuade, M., and Solawetz, J · 2024
Closest in time.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Guha, N., Nyarko, J., Ho, D., Ré, C., Chilton, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D., Zambrano, D., et al · 2024
Closest in time.
Large language model based multi-agents: A survey of progress and challenges
Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., and Zhang, X · 2024
Closest in time.
Llm multi-agent systems: Challenges and open problems
Han, S., Zhang, Q., Yao, Y., Jin, W., Xu, Z., and He, C · 2024
Closest in time.
Metagpt: Meta programming for a multi-agent collaborative framework
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ivison, H., Wang, Y., Pyatkin, V., Lambert, N., Peters, M., Dasigi, P., Jang, J., Wadden, D., Smith, N. A., Beltagy, I., et al · 2023
Cited alongside, same era.
Exploring the benefits of training expert language models over instruction tuning
Jang, J., Kim, S., Ye, S., Kim, D., Logeswaran, L., Lee, M., Lee, K., and Seo, M · 2023
Cited alongside, same era.
Active retrieval augmented generation
Jiang, Z., Xu, F. F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., and Neubig, G · 2023
Cited alongside, same era.
Smart-llm: Smart multi-agent robot task planning using large language models
Kannan, S. S., Venkatesh, V. L., and Min, B.-C · 2023
Cited alongside, same era.
Neuroevobench: Benchmarking evolutionary optimizers for deep learning applications
Lange, R., Tang, Y., and Tian, Y · 2023
Cited alongside, same era.
Openorca: An open dataset of gpt augmented flan reasoning traces, 2023
Lian, W., Goodson, B., Pentland, E., et al · 2023
Cited alongside, same era.
Multi-agent collaboration: Harnessing the power of intelligent llm agents
Talebirad, Y. and Nadiri, A · 2023
Cited alongside, same era.
Closest in time.
On the resilience of multi-agent systems with malicious agents
Huang, J.-t., Zhou, J., Jin, T., Zhou, X., Chen, Z., Wang, W., Yuan, Y., Sap, M., and Lyu, M. R · 2024
Closest in time.
Ishibashi, Y. and Nishimura, Y · 2024
Closest in time.
Model stock: All we need is just a few fine-tuned models
Jang, D.-H., Yun, S., and Han, D · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., et al · 2024
Closest in time.
Large language models as evolution strategies
Lange, R., Tian, Y., and Tang, Y · 2024
Closest in time.
Moe-llava: Mixture of experts for large vision-language models
Lin, B., Tang, Z., Ye, Y., Cui, J., Zhu, B., Jin, P., Zhang, J., Ning, M., and Yuan, L · 2024
Closest in time.
Tuning language models by proxy
Liu, A., Han, X., Wang, Y., Tsvetkov, Y., Choi, Y., and Smith, N. A · 2024
Closest in time.
Pack of llms: Model fusion at test-time via perplexity optimization
Mavromatis, C., Karypis, P., and Karypis, G · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Warp: On the benefits of weight averaged rewarded policies
Ramé, A., Ferret, J., Vieillard, N., Dadashi, R., Hussenot, L., Cedoz, P.-L., Sessa, P. G., Girgin, S., Douillard, A., and Bachem, O · 2024
Closest in time.
Warm: On the benefits of weight averaged reward models
Rame, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, G., Bachem, O., and Ferret, J · 2024
Closest in time.
Normad: A benchmark for measuring the cultural adaptability of large language models
Rao, A., Yerukola, A., Shah, V., Reinecke, K., and Sap, M · 2024
Closest in time.
Replug: Retrieval-augmented black-box language models
Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t · 2024
Closest in time.
Should we be going mad? a look at multi-agent debate strategies for llms
Smit, A. P., Grinsztajn, N., Duckworth, P., Barrett, T. D., and Pretorius, A · 2024
Closest in time.
Position: A roadmap to pluralistic alignment
Sorensen, T., Moore, J., Fisher, J., Gordon, M. L., Mireshghallah, N., Rytting, C. M., Ye, A., Jiang, L., Lu, X., Dziri, N., et al · 2024
Closest in time.
Llm-based multi-agent reinforcement learning: Current and future directions
Sun, C., Huang, S., and Pompili, D · 2024
Closest in time.
Merging multi-task models via weight-ensembling mixture of experts
Tang, A., Shen, L., Luo, Y., Yin, N., Zhang, L., and Tao, D · 2024
Closest in time.
Knowledge fusion of large language models
Wan, F., Huang, X., Cai, D., Quan, X., Bi, W., and Shi, S · 2024
Closest in time.
Evolutionary computation in the era of large language model: Survey and roadmap
Wu, X., Wu, S.-h., Wu, J., Feng, L., and Tan, K. C · 2024
Closest in time.
Ties-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C. A., and Bansal, M · 2024
Closest in time.
Adamerging: Adaptive model merging for multi-task learning
Yang, E., Wang, Z., Shen, L., Liu, S., Guo, G., Wang, X., and Tao, D · 2024
Closest in time.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y · 2024
Closest in time.
Autodefense: Multi-agent llm defense against jailbreak attacks
Zeng, Y., Wu, Y., Zhang, X., Wang, H., and Wu, Q · 2024
Closest in time.
Competeai: Understanding the competition dynamics of large language model-based agents
Zhao, Q., Wang, J., Zhang, Y., Jin, Y., Zhu, K., Chen, H., and Xie, X · 2024
Closest in time.
Weak-to-strong extrapolation expedites alignment
Zheng, C., Wang, Z., Ji, H., Huang, M., and Peng, N · 2024
Closest in time.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2024
Closest in time.