Fetching the paper…
Reading the bibliography…
Direct preference learning offers a promising and computation-efficient beyond supervised fine-tuning (SFT) for improving code generation in coding large language models (LMs).
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A · 1952
Earlier work this paper cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo
Engstrom, L · 2005
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Earlier work this paper cites.
Program synthesis with large language models
Austin, J · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M · 2021
Earlier work this paper cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Le, H · 2022
Earlier work this paper cites.
Achiam, J · 2023
Earlier work this paper cites.
Introducing claude
Anthropic · 2023
Earlier work this paper cites.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G · 2023
Earlier work this paper cites.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R · 2023
Earlier work this paper cites.
Code llama: Open foundation models for code
Roziere, B · 2023
Earlier work this paper cites.
Execution-based code generation using deep reinforcement learning
Shojaee, P · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Team, G · 2023
Cited alongside, same era.
Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf
Xiong, W · 2023
Cited alongside, same era.
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y · 2023
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges
Ahn, J · 2024
Cited alongside, same era.
Stepcoder: Improve code generation with reinforcement learning from compiler feedback
Dou, S · 2024
Robust preference optimization with provable noise tolerance for llms
Liang, X · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation
Lozhkov, A · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
Pal, A · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C · 2024
Closest in time.
Policy filtration in rlhf to fine-tune llm for code generation
Shen, W · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K · 2024
Cited alongside, same era.
Code-optimise: Self-generated preference data for correctness and efficiency
Gee, L · 2024
Cited alongside, same era.
Orpo: Monolithic preference optimization without reference model
Hong, J · 2024
Cited alongside, same era.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J · 2024
Cited alongside, same era.
Qwen2. 5-coder technical report
Hui, B · 2024
Cited alongside, same era.
Towards efficient and exact optimization of language model alignment
Ji, H · 2024
Cited alongside, same era.
A survey on large language models for code generation
Jiang, J · 2024
Cited alongside, same era.
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Tajwar, F · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Tang, Y · 2024
Closest in time.
Self-play preference optimization for language model alignment
Wu, Y · 2024
Closest in time.
Just say what you want: Only-prompting self-rewarding online preference optimization
Xu, R · 2024
Closest in time.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
Ye, C · 2024
Closest in time.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Zhu, Q · 2024
Closest in time.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Zhuo, T. Y · 2024
Closest in time.