Fetching the paper…
Reading the bibliography…
Recent studies show that in supervised fine-tuning (SFT) of large language models (LLMs), data quality matters more than quantity.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
Learning with noisy labels
Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A · 2013
Earlier work this paper cites.
Training deep neural networks on noisy labels with bootstrapping
Reed, S., Lee, H., Anguelov, D., Szegedy, C., Erhan, D., and Rabinovich, A · 2014
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
Learning from noisy labels with distillation
Li, Y., Yang, J., Song, Y., Cao, L., Luo, J., and Li, L.-J · 2017
Earlier work this paper cites.
Learning with confident examples: Rank pruning for robust classification with noisy labels
Northcutt, C. G., Wu, T., and Chuang, I. L · 2017
Earlier work this paper cites.
Toward robustness against label noise in training deep discriminative neural networks
Vahdat, A · 2017
Earlier work this paper cites.
Learning from noisy large-scale datasets with minimal supervision
Veit, A., Alldrin, N., Chechik, G., Krasin, I., Gupta, A., and Belongie, S · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Tydi qa: A benchmark for information-seeking question answering in ty pologically di verse languages
Clark, J. H., Choi, E., Collins, M., Garrette, D., Kwiatkowski, T., Nikolaev, V., and Palomaki, J · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Learning with instance-dependent label noise: A sample sieve approach
Cheng, H., Zhu, Z., Li, X., Gong, Y., Sun, X., and Liu, Y · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Earlier work this paper cites.
Can less be more? when increasing-to-balancing label noise rates considered beneficial
Liu, Y. and Wang, J · 2021
Earlier work this paper cites.
Confident learning: Estimating uncertainty in dataset labels
Northcutt, C., Jiang, L., and Chuang, I · 2021
Earlier work this paper cites.
A second-order approach to learning with instance-dependent label noise
Zhu, Z., Liu, T., and Liu, Y · 2021
Cited alongside, same era.
Detecting label errors by using pre-trained language models
Chong, D., Hong, J., and Manning, C · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2022
Cited alongside, same era.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Mindermann, S., Brauner, J. M., Razzak, M. T., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., et al · 2022
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Cited alongside, same era.
How far can camels go? exploring the state of instruction tuning on open resources, 2023
Wang, Y., Ivison, H., Dasigi, P., Hessel, J., Khot, T., Chandu, K. R., Wadden, D., MacMillan, K., Smith, N. A., Beltagy, I., and Hajishirzi, H · 2023
Later among the works it cites.
Wizardlm: Empowering large language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D · 2023
Later among the works it cites.
Dataset quantization
Zhou, D., Wang, K., Gu, J., Peng, X., Lian, D., Zhang, Y., You, Y., and Feng, J · 2023
Later among the works it cites.
Instruction mining: Instruction data selection for tuning large language models
Cao, Y., Kang, Y., Wang, C., and Sun, L · 2024
Later among the works it cites.
Automated data curation for robust language model fine-tuning
Chen, J. and Mueller, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Detecting corrupted labels without training a model to predict
Zhu, Z., Dong, Z., and Liu, Y · 2022
Cited alongside, same era.
Alpagasus: Training a better alpaca with fewer data
Chen, L., Li, S., Yan, J., Wang, H., Gunaratna, K., Yadav, V., Tang, Z., Srinivasan, V., Zhou, T., Huang, H., et al · 2023
Cited alongside, same era.
Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
Databricks · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Estimating label quality and errors in semantic segmentation data via any model
Lad, V. and Mueller, J · 2023
Cited alongside, same era.
Li, M., Zhang, Y., Li, Z., Chen, J., Chen, L., Cheng, N., Wang, J., Zhou, T., and Xiao, J · 2023
Cited alongside, same era.
Liu, W., Zeng, W., He, K., Jiang, Y., and He, J · 2023
Cited alongside, same era.
Provably robust dpo: Aligning language models with noisy feedback
Chowdhury, S. R., Kini, A., and Natarajan, N · 2024
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Mixed preference optimization: Reinforcement learning with data selection and better reference model
Gou, Q. and Nguyen, C.-T · 2024
Later among the works it cites.
Openassistant conversations-democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., et al · 2024
Later among the works it cites.
Not all tokens are what you need for pretraining
Lin, Z., Gou, Z., Gong, Y., Liu, X., Xu, R., Lin, C., Yang, Y., Jiao, J., Duan, N., Chen, W., et al · 2024
Later among the works it cites.
Less: Selecting influential data for targeted instruction tuning
Xia, M., Malladi, S., Gururangan, S., Arora, S., and Chen, D · 2024
Later among the works it cites.
Early stopping against label noise without validation data
Yuan, S., Feng, L., and Liu, T · 2024
Later among the works it cites.
When scaling meets llm finetuning: The effect of data, model and finetuning method
Zhang, B., Liu, Z., Cherry, C., and Firat, O · 2024
Later among the works it cites.
Long is more for alignment: A simple but tough-to-beat baseline for instruction fine-tuning
Zhao, H., Andriushchenko, M., Croce, F., and Flammarion, N · 2024
Later among the works it cites.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2024
Later among the works it cites.
Unmasking and improving data credibility: A study with datasets for training harmless language models
Zhu, Z., Wang, J., Cheng, H., and Liu, Y · 2024
Later among the works it cites.