Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have achieved remarkable progress in reasoning, alignment, and task-specific performance.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Making harmful behaviors unlearnable for large language models, 2023
Xin Zhou, Yi Lu, Ruotian Ma, Tao Gui, Qi Zhang, and Xuanjing Huang · 2023
Earlier work this paper cites.
Superhf: Supervised iterative learning from human feedback, 2023
Gabriel Mukobi, Peter Chatain, Su Fong, Robert Windesheim, Gitta Kutyniok, Kush Bhatia, and Silas Alberti · 2023
Earlier work this paper cites.
Reinforcement learning fine-tuning of language models is biased towards more extractable features, 2023
Diogo Cruz, Edoardo Pona, Alex Holness-Tofts, Elias Schmied, Víctor Abia Alonso, Charlie Griffin, and Bogdan-Ionut Cirstea · 2023
Earlier work this paper cites.
A survey of reinforcement learning from human feedback, 2024
Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier · 2024
Earlier work this paper cites.
The ultimate guide to fine-tuning llms from basics to breakthroughs: An exhaustive review of technologies, research, best practices, applied research challenges and opportunities, 2024
Venkatesh Balavadhani Parthasarathy, Ahtsham Zafar, Aafaq Khan, and Arsalan Shahid · 2024
Earlier work this paper cites.
A survey on knowledge distillation of large language models, 2024
Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou · 2024
Earlier work this paper cites.
Deepseekmath Pushing the limits of mathematical reasoning in open language models, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Cited alongside, same era.
Ai alignment through reinforcement learning from human feedback? contradictions and limitations, 2024
Adam Dahlgren Lindström, Leila Methnani, Lea Krause, Petter Ericson, Íñigo Martínez de Rituerto de Troya, Dimitri Coelho Mollo, and Roel Dobbe · 2024
Cited alongside, same era.
Reinforcement learning enhanced llms: A survey, 2024
Shuhe Wang, Shengyu Zhang, Jie Zhang, Runyi Hu, Xiaoya Li, Tianwei Zhang, Jiwei Li, Fei Wu, Guoyin Wang, and Eduard Hovy · 2024
Cited alongside, same era.
Harmful fine-tuning attacks and defenses for large language models: A survey, 2024
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu · 2024
Cited alongside, same era.
Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback, 2024
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash · 2024
Learning and forgetting unsafe examples in large language models, 2024
Jiachen Zhao, Zhun Deng, David Madras, James Zou, and Mengye Ren · 2024
Later among the works it cites.
Unlock the correlation between supervised fine-tuning and reinforcement learning in training code large language models, 2024
Jie Chen, Xintian Han, Yu Ma, Xun Zhou, and Liang Xiang · 2024
Later among the works it cites.
Q-sft: Q-learning for language models via supervised fine-tuning, 2024
Joey Hong, Anca Dragan, and Sergey Levine · 2024
Later among the works it cites.
A closer look at the limitations of instruction tuning, 2024
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S, Deepali Aneja, Zeyu Jin, Ramani Duraiswami, and Dinesh Manocha · 2024
Later among the works it cites.
Mitigating forgetting in llm supervised fine-tuning and preference learning, 2024
Heshan Fernando, Han Shen, Parikshit Ram, Yi Zhou, Horst Samulowitz, Nathalie Baracaldo, and Tianyi Chen · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Safety-aware fine-tuning of large language models, 2024
Hyeong Kyu Choi, Xuefeng Du, and Yixuan Li · 2024
Cited alongside, same era.
Reinforcing thinking through reasoning-enhanced reward models, 2024
Diji Yang, Linda Zeng, Kezhen Chen, and Yi Zhang · 2024
Cited alongside, same era.
Supervised fine-tuning as inverse reinforcement learning, 2024
Hao Sun · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI and Daya Guo et. al · 2025
Closest in time.