Fetching the paper…
Reading the bibliography…
This report examines the fine-tuning of Large Language Models (LLMs), integrating theoretical insights with practical applications.
Early stopping-but when?
Lutz Prechelt · 1998
Earlier work this paper cites.
Smote: synthetic minority over-sampling technique
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer · 2002
Earlier work this paper cites.
Cheap and fast—but is it good? evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan O’Connor, Dan Jurafsky, and Andrew Y Ng · 2008
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space, 2013
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2015
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Snorkel: Rapid training data creation with weak supervision
Alexander Ratner, Stephen H Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou · 2017
Earlier work this paper cites.
Fairness in machine learning: Lessons from political philosophy
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Fairness in machine learning: Lessons from political philosophy
Reuben Binns · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Automatic data labeling for supervised learning with applications to visual inspection of mixed-plastic waste
Liang Ding, Philipp Gentner, Artur Duda, Vaibhav Sangtani, Dominik Ziegler, Max Hennen, Siddharth Jain, and Roland Werthschützky · 2019
Earlier work this paper cites.
Machine learning for data preprocessing
Pradeep Rajan, Krishna Vyas, Rajiv Bansal, Ranjan Sharma, and Shubhranshu Mukherjee · 2019
Earlier work this paper cites.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M Khoshgoftaar · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raghavendra Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Cho-Jui Hsieh, and Payal Yadollahpour · 2019
Earlier work this paper cites.
Automated Machine Learning: Methods, Systems, Challenges
Frank Hutter, Lars Kotthoff, and Joaquin Vanschoren · 2019
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks, 2019
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith · 2020
Earlier work this paper cites.
Snorkel: Rapid training data creation with weak supervision
Alexander Ratner, Henry Ehrenberg, Zeshan Hussain, Jared Dunnmon, and Christopher Ré · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2020
Earlier work this paper cites.
Q-bert: Hessian based ultra low precision quantization of bert
Sheng Shen, Zhewei Dong, Xiaocheng Ye, Linjian Ma, Zhewei Li, Zirui Wang, Samyam Rajbhandari, Yuxiong Wang, and Zhen Yang · 2020
Cited alongside, same era.
Deepspeed: Extreme-scale model training for everyone
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2021
Cited alongside, same era.
A survey of data augmentation approaches for nlp
Steven Feng, Varun Gangal, Jinjun Wei, Yashvardhan Chandrasekhar, Yichong Chen, Dani He, Shuyang Huang, Faisal Ladhak, Jiao Lee, Xinyi Li, et al · 2021
Cited alongside, same era.
Parameter-efficient fine-tuning for large models: A comprehensive survey, 2024
Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang · 2024
Closest in time.
Practical Tips for Finetuning LLMs Using LoRA (Low-Rank Adaptation) — magazine.sebastianraschka.com
PhD Sebastian Raschka · 2024
Closest in time.
https://community.analyticsvidhya.com/c/generative-ai-tech-discussion/what-is-qlora
What is QLoRa? — Analytics Vidhya — community.analyticsvidhya.com · 2024
Closest in time.
Dora: Weight-decomposed low-rank adaptation, 2024
Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen · 2024
Closest in time.
Apple intelligence foundation language models, 2024
2024
Closest in time.
Hft: Half fine-tuning for large language models, 2024
Tingfeng Hui, Zhenyu Zhang, Shuohuan Wang, Weiran Xu, Yu Sun, and Hua Wu · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
The stanford natural language inference (snli) corpus
Sebastian Ruder · 2021
Cited alongside, same era.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units, 2021
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning, 2022
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel · 2022
Cited alongside, same era.
Closest in time.
Banishing llm hallucinations requires rethinking generalization, 2024
Johnny Li, Saksham Consul, Eda Zhou, James Wong, Naila Farooqui, Yuxin Ye, Nithyashree Manohar, Zhuxiaona Wei, Tian Wu, Ben Echols, Sharon Zhou, and Gregory Diamos · 2024
Closest in time.
Mixtral of experts, 2024
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
https://developer.nvidia.com/blog/applying-mixture-of-experts-in-llm-architectures/
Applying Mixture of Experts in LLM Architectures — NVIDIA Technical Blog — developer.nvidia.com · 2024
Closest in time.
Mixture-of-agents enhances large language model capabilities, 2024
Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, and James Zou · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model, 2024
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2024
Closest in time.
Is dpo superior to ppo for llm alignment? a comprehensive study, 2024
Shusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye, Weilin Liu, Zhiyu Mei, Guangju Wang, Chao Yu, and Yi Wu · 2024
Closest in time.
Orpo: Monolithic preference optimization without reference model
Jiwoo Hong, Noah Lee, and James Thorne · 2024
Closest in time.
Orpo evaluation: Performance on alpacaeval and mt-bench
Jiwoo Hong, Noah Lee, and James Thorne · 2024
Closest in time.
https://www.linkedin.com/advice/3/what-most-effective-techniques-pruning-0mlef
What are the most effective techniques for pruning ai models? — linkedin.com · 2024
Closest in time.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models, 2024
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li · 2024
Closest in time.
Shieldgemma: Generative ai content moderation based on gemma, 2024
Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, Olivia Sturman, and Oscar Wahltinez · 2024
Closest in time.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms, 2024
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri · 2024
Closest in time.
LLM Deployment Strategies : Its not Magic , Its Logic! — visrow
Vishal Mysore · 2024
Closest in time.
https://aws.amazon.com/blogs/big-data/preprocess-and-fine-tune-llms-quickly-and-cost-effectively-using-amazon-emr-serverless-and-amazon-sagemaker/
Preprocess and fine-tune llms quickly and cost-effectively using amazon emr serverless and amazon sagemaker — aws.amazon.com · 2024
Closest in time.
https://www.run.ai/guides/ai-open-source-projects/nvidia-nemo
Nvidia nemo build and customize your own llms (with tutorial) — run.ai · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team and Rohan Anil et al · 2024
Closest in time.
Efficient multimodal large language models: A survey, 2024
Yizhang Jin, Jian Li, Yexin Liu, Tianjun Gu, Kai Wu, Zhengkai Jiang, Muyang He, Bo Zhao, Xin Tan, Zhenye Gan, Yabiao Wang, Chengjie Wang, and Lizhuang Ma · 2024
Closest in time.
Not all attention is needed: Parameter and computation efficient transfer learning for multi-modal large language models, 2024
Qiong Wu, Weihao Ye, Yiyi Zhou, Xiaoshuai Sun, and Rongrong Ji · 2024
Closest in time.
Memory-space visual prompting for efficient vision-language fine-tuning, 2024
Shibo Jie, Yehui Tang, Ning Ding, Zhi-Hong Deng, Kai Han, and Yunhe Wang · 2024
Closest in time.
Full parameter fine-tuning for large language models with limited resources, 2024
Kai Lv, Yuqing Yang, Tengxiao Liu, Qinghui Gao, Qipeng Guo, and Xipeng Qiu · 2024
Closest in time.
Fine-tuning language models with just forward passes, 2024
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora · 2024
Closest in time.
Pefomed: Parameter efficient fine-tuning of multimodal large language models for medical imaging, 2024
Gang Liu, Jinlong He, Pengfei Li, Genrong He, Zhaolin Chen, and Shenjun Zhong · 2024
Closest in time.
Audio language models and multimodal architecture — prdeepak.babu
Deepak Babu P R · 2024
Closest in time.
A comprehensive overview of large language models, 2024
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian · 2024
Closest in time.
https://rocm.blogs.amd.com/artificial-intelligence/llama2-lora/README.html
Fine-tune llama 2 with lora: Customizing a large language model for question-answering — rocm.blogs.amd.com · 2024
Closest in time.
Understanding llm fine-tuning: Tailoring large language models to your unique requirements — linkedin.com
Aayush Mittal · 2024
Closest in time.
Scaling sparse fine-tuning to large language models, 2024
Alan Ansell, Ivan Vulić, Hannah Sterz, Anna Korhonen, and Edoardo M. Ponti · 2024
Closest in time.
Data-efficient fine-tuning for llm-based recommendation, 2024
Xinyu Lin, Wenjie Wang, Yongqi Li, Shuo Yang, Fuli Feng, Yinwei Wei, and Tat-Seng Chua · 2024
Closest in time.
End-to-end learnable clustering for intent learning in recommendation, 2024
Yue Liu, Shihao Zhu, Jun Xia, Yingwei Ma, Jian Ma, Wenliang Zhong, Xinwang Liu, Guannan Zhang, and Kejun Zhang · 2024
Closest in time.
Federated domain-specific knowledge transfer on large language models using synthetic data, 2024
Haoran Li, Xinyuan Zhao, Dadi Guo, Hanlin Gu, Ziqian Zeng, Yuxing Han, Yangqiu Song, Lixin Fan, and Qiang Yang · 2024
Closest in time.