Fetching the paper…
Reading the bibliography…
In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal instruction data on distributed devices.
Personalized federated learning: A meta-learning approach
Fallah, A.; Mokhtari, A.; and Ozdaglar, A. 2020 · 2002
Earlier work this paper cites.
Adaptive federated optimization
Reddi, S.; Charles, Z.; Zaheer, M.; Garrett, Z.; Rush, K.; Konečnỳ, J.; Kumar, S.; and McMahan, H. B. 2020 · 2003
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Domain separation networks
Bousmalis, K.; Trigeorgis, G.; Silberman, N.; Krishnan, D.; and Erhan, D. 2016 · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C.; Abbeel, P.; and Levine, S. 2017 · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; and y Arcas, B. A. 2017 · 2017
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson. 2019 · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A.; Shekhar, S.; Singh, A. K.; and Chakraborty, A. 2019 · 2019
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Li, T.; Sahu, A. K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V. 2020 · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Earlier work this paper cites.
Local sgd: Unified theory and new efficient methods
Gorbunov, E.; Hanzely, F.; and Richtárik, P. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021 · 2021
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022 · 2022
Cited alongside, same era.
Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning
Xu, Z.; et al. 2022 · 2022
Cited alongside, same era.
St-moe: Designing stable and transferable sparse expert models
Zoph, B.; Bello, I.; Kumar, S.; Du, N.; Huang, Y.; Dean, J.; Shazeer, N.; and Fedus, W. 2022 · 2022
Cited alongside, same era.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Client-Adaptive Cross-Model Reconstruction Network for Modality-Incomplete Multimodal Federated Learning
Xiong, B.; Yang, X.; Song, Y.; Wang, Y.; and Xu, C. 2023 · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Task-Agnostic Privacy-Preserving Representation Learning for Federated Learning against Attribute Inference Attacks
Arevalo, C. A.; Noorbakhsh, S. L.; Dong, Y.; Hong, Y.; and Wang, B. 2024 · 2024
Later among the works it cites.
On Disentanglement of Asymmetrical Knowledge Transfer for Modality-Task Agnostic Federated Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, K.; Zhang, Z.; Zeng, W.; Zhang, R.; Zhu, F.; and Zhao, R. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; et al. 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; et al. 2023 · 2023
Cited alongside, same era.
Fate-llm: A industrial grade federated learning framework for large language models
Fan, T.; Kang, Y.; Ma, G.; Chen, W.; Wei, W.; Fan, L.; and Yang, Q. 2023 · 2023
Cited alongside, same era.
Pfedprompt: Learning personalized prompt for vision-language models in federated learning
Guo, T.; et al. 2023 · 2023
Cited alongside, same era.
Continual instruction tuning for large multimodal models
He, J.; Guo, H.; Tang, M.; and Wang, J. 2023 · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
Chen, J.; and Zhang, A. 2024 · 2024
Later among the works it cites.
FedLPS: Heterogeneous Federated Learning for Multiple Tasks with Local Parameter Sharing
Jia, Y.; Zhang, X.; Beheshti, A.; and Dou, W. 2024 · 2024
Later among the works it cites.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Later among the works it cites.
Multimodal Instruction Tuning with Conditional Mixture of LoRA
Shen, Y.; Xu, Z.; Wang, Q.; Cheng, Y.; Yin, W.; and Huang, L. 2024 · 2024
Later among the works it cites.
Federated adaptive prompt tuning for multi-domain collaborative learning
Su, S.; Yang, M.; Li, B.; and Xue, X. 2024 · 2024
Later among the works it cites.
Umie: Unified multimodal information extraction with instruction tuning
Sun, L.; Zhang, K.; Li, Q.; and Lou, R. 2024 · 2024
Later among the works it cites.
Modality-Collaborative Test-Time Adaptation for Action Recognition
Xiong, B.; Yang, X.; Song, Y.; Wang, Y.; and Xu, C. 2024 · 2024
Later among the works it cites.
Libra: Building Decoupled Vision System on Large Language Models
Xu, Y.; Yang, X.; Song, Y.; and Xu, C. 2024 · 2024
Later among the works it cites.
Openfedllm: Training large language models on decentralized private data via federated learning
Ye, R.; Wang, W.; Chai, J.; Li, D.; Li, Z.; Xu, Y.; Du, Y.; Wang, Y.; and Chen, S. 2024 · 2024
Later among the works it cites.
Towards building the federatedGPT: Federated instruction tuning
Zhang, J.; Vahidian, S.; Kuo, M.; Li, C.; Zhang, R.; Yu, T.; Wang, G.; and Chen, Y. 2024b · 2024
Later among the works it cites.