Fetching the paper…
Reading the bibliography…
Pre-trained large language models (LLMs) need fine-tuning to improve their responsiveness to natural language instructions.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Snapshot ensembles: Train 1, get M for free
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J. E., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Earlier work this paper cites.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Earlier work this paper cites.
Zeroth-order stochastic variance reduction for nonconvex optimization
Liu, S., Kailkhura, B., Chen, P.-Y., Ting, P., Chang, S., and Amini, L · 2018
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V · 2020
Earlier work this paper cites.
FetchSGD: Communication-efficient federated learning with sketching
Rothchild, D., Panda, A., Ullah, E., Ivkin, N., Stoica, I., Braverman, V., Gonzalez, J., and Arora, R · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Gupta, S., and Zettlemoyer, L · 2021
Earlier work this paper cites.
Advances and open problems in federated learning
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K. A., Charles, Z., Cormode, G., Cummings, R., D’Oliveira, R. G. L., Eichner, H., Rouayheb, S. E., Evans, D., Gardner, J., Garrett, Z., Gascón, A., Ghazi, B., Gibbons, P. B., Gruteser, M., Harchaoui, Z., He, C., He, L., Huo, Z., Hutchinson, B., Hsu, J., Jaggi, M., Javidi, T., Joshi, G., Khodak, M., Konečný, J., Korolova, A., Koushanfar, F., Koyejo, S., Lepoint, T., Liu, Y., Mittal, P., Mohri, M., Nock, R., Özgür, A., Pagh, R., Qi, H., Ramage, D., Raskar, R., Raykova, M., Song, D., Song, W., Stich, S. U., Sun, Z., Suresh, A. T., Tramèr, F., Vepakomma, P., Wang, J., Xiong, L., Xu, Z., Yang, Q., Yu, F. X., Yu, H., and Zhao, S · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Communication-efficient decentralized zeroth-order method on heterogeneous data
Li, Z. and Chen, L · 2021
Earlier work this paper cites.
Intrinisic gradient compression for federated learning
Melas-Kyriazi, L. and Wang, F · 2021
Earlier work this paper cites.
Revisiting parameter-efficient tuning: Are we really there yet?
Chen, G., Liu, F., Meng, Z., and Liang, S · 2022
Earlier work this paper cites.
Communication-efficient stochastic zeroth-order optimization for federated learning
Fang, W., Yu, Z., Jiang, Y., Shi, Y., Jones, C. N., and Zhou, Y · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Federated learning on non-IID data silos: An experimental study
Li, Q., Diao, Y., Chen, Q., and He, B · 2022
Cited alongside, same era.
PEFT: State-of-the-art parameter-efficient fine-tuning methods
Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., and Bossan, B · 2022
Cited alongside, same era.
Towards personalized federated learning
Tan, A. Z., Yu, H., Cui, L., and Yang, Q · 2022
Cited alongside, same era.
Will we run out of data? an analysis of the limits of scaling datasets in machine learning
Villalobos, P., Sevilla, J., Heim, L., Besiroglu, T., Hobbhahn, M., and Ho, A · 2022
Cited alongside, same era.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ NLP tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Naik, A., Ashok, A., Dhanasekaran, A. S., Arunkumar, A., Stap, D., Pathak, E., Karamanolakis, G., Lai, H. G., Purohit, I., Mondal, I., Anderson, J., Kuznia, K., Doshi, K., Pal, K. K., Patel, M., Moradshahi, M., Parmar, M., Purohit, M., Varshney, N., Kaza, P. R., Verma, P., Puri, R. S., Karia, R., Doshi, S., Sampat, S. K., Mishra, S., A, S. R., Patro, S., Dixit, T., and Shen, X · 2022
Fine-tuning language models with just forward passes
Malladi, S., Gao, T., Nichani, E., Damian, A., Lee, J. D., Chen, D., and Arora, S · 2023
Closest in time.
FedZeN: Towards superlinear zeroth-order federated learning via incremental hessian estimation
Maritan, A., Dey, S., and Schenato, L · 2023
Closest in time.
Empirical analysis of the strengths and weaknesses of peft techniques for llms
Pu, G., Jain, A., Yin, J., and Kaplan, R · 2023
Closest in time.
Federated zeroth-order optimization using trajectory-informed surrogate gradients
Shu, Y., Lin, X., Dai, Z., and Low, B. K. H · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Cited alongside, same era.
SLoRA: Federated parameter efficient fine-tuning of language models
Babakniya, S., Elkordy, A. R., Ezzeldin, Y. H., Liu, Q., Song, K.-B., El-Khamy, M., and Avestimehr, S · 2023
Cited alongside, same era.
Distributed inference and fine-tuning of large language models over the internet
Borzunov, A., Ryabinin, M., Chumachenko, A., Baranchuk, D., Dettmers, T., Belkada, Y., Samygin, P., and Raffel, C · 2023
Cited alongside, same era.
Federated learning of large language models with parameter-efficient prompt tuning and adaptive optimization
Che, T., Liu, J., Zhou, Y., Ren, J., Zhou, J., Sheng, V., Dai, H., and Dou, D · 2023
Cited alongside, same era.
Free dolly: Introducing the world’s first truly open instruction-tuned LLM, 2023
Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R · 2023
Cited alongside, same era.
QLoRA: Efficient finetuning of quantized LLMs
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Cited alongside, same era.
Towards next-generation intelligent assistants leveraging LLM techniques
Dong, X. L., Moon, S., Xu, Y. E., Malik, K., and Yu, Z · 2023
Cited alongside, same era.
Sun, X., Ji, Y., Ma, B., and Li, X · 2023
Closest in time.
Stanford Alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
Federated fine-tuning of LLMs on the very edge: The good, the bad, the ugly
Woisetschläger, H., Isenko, A., Wang, S., Mayer, R., and Jacobsen, H.-A · 2023
Closest in time.
Training transformers with 4-bit integers
Xi, H., Li, C., Chen, J., and Zhu, J · 2023
Closest in time.
FwdLLM: Efficient FedLLM using forward gradient
Xu, M., Wu, Y., Cai, D., Li, X., and Wang, S · 2023
Closest in time.
Zelikman, E., Huang, Q., Liang, P., Haber, N., and Goodman, N. D · 2023
Closest in time.
FedPETuning: When federated learning meets the parameter-efficient tuning methods of pre-trained language models
Zhang, Z., Yang, Y., Dai, Y., Wang, Q., Yu, Y., Qu, L., and Xu, Z · 2023
Closest in time.
Bai, J., Chen, D., Qian, B., Yao, L., and Li, Y · 2024
Closest in time.
Data-juicer: A one-stop data processing system for large language models
Chen, D., Huang, Y., Ma, Z., Chen, H., Pan, X., Ge, C., Gao, D., Xie, Y., Liu, Z., Gao, J., Li, Y., Ding, B., and Zhou, J · 2024
Closest in time.
On the convergence of zeroth-order federated tuning for large language models
Ling, Z., Chen, D., Yao, L., Li, Y., and Shen, Y · 2024
Closest in time.
BlockDFL: A blockchain-based fully decentralized peer-to-peer federated learning framework
Qin, Z., Yan, X., Zhou, M., and Deng, S · 2024
Closest in time.
EvoFed: Leveraging evolutionary strategies for communication-efficient federated learning
Rahimi, M. M., Bhatti, H. I., Park, Y., Kousar, H., Kim, D.-Y., and Moon, J · 2024
Closest in time.
Towards building the federatedgpt: Federated instruction tuning
Zhang, J., Vahidian, S., Kuo, M., Li, C., Zhang, R., Yu, T., Wang, G., and Chen, Y · 2024
Closest in time.