Fetching the paper…
Reading the bibliography…
Reward-based alignment methods for large language models (LLMs) face two key limitations: vulnerability to reward hacking, where models exploit flaws in the reward signal; and reliance on brittle, labor-intensive prompt engineering when LLMs are used as reward models.
The Use of Lateral Thinking
E. De Bono · 1971
Earlier work this paper cites.
Metacognition and cognitive monitoring: A new area of cognitive–developmental inquiry
J. H. Flavell · 1979
Earlier work this paper cites.
Biased assimilation and attitude polarization: The effects of prior theories on subsequently considered evidence
C. G. Lord, L. Ross, and M. R. Lepper · 1979
Earlier work this paper cites.
Rhetorical structure theory: A theory of text organization
W. C. Mann and S. A. Thompson · 1987
Earlier work this paper cites.
Metacognition and learning
C. B. McCormick · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
A region of proximal learning model of study time allocation
J. Metcalfe and N. Kornell · 2004
Earlier work this paper cites.
Metacognition and affect: What can metacognitive experiences tell us about the learning process?
A. Efklides · 2006
Earlier work this paper cites.
Metacognition and learning: Conceptual and methodological considerations
M. V. Veenman, B. H. Van Hout-Wolters, and P. Afflerbach · 2006
Earlier work this paper cites.
The hewlett foundation: Automated essay scoring
B. Hamner, J. Morgan, lynnvandev, M. Shermis, and T. V. Ark · 2012
Earlier work this paper cites.
Defining and teaching evaluative thinking: Insights from research on critical thinking
J. Buckley, T. Archibald, M. Hargraves, and W. M. Trochim · 2015
Earlier work this paper cites.
BillSum: A corpus for automatic summarization of US legislation
A. Kornilova and V. Eidelman · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Specification gaming: the flip side of ai ingenuity
V. Krakovna, J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg · 2020
Earlier work this paper cites.
Scruples: A corpus of community ethical judgments on 32,000 real-life anecdotes
N. Lourie, R. L. Bras, and Y. Choi · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec · 2020
Earlier work this paper cites.
Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
T. Everitt, M. Hutter, R. Kumar, and V. Krakovna · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Goal misgeneralization in deep reinforcement learning
L. L. D. Langosco, J. Koch, L. D. Sharkey, J. Pfau, and D. Krueger · 2022
Earlier work this paper cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
H. Le, Y. Wang, A. D. Gotmare, S. Savarese, and S. C. H. Hoi · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
A. Pan, K. Bhatia, and J. Steinhardt · 2022
Cited alongside, same era.
Defining and characterizing reward gaming
J. M. V. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger · 2022
Cited alongside, same era.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2023
Cited alongside, same era.
Reward gaming in conditional text generation
R. Y. Pang, V. Padmakumar, T. Sellam, A. Parikh, and H. He · 2023
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
InfoRM: Mitigating reward hacking in RLHF via information-theoretic reward modeling
Y. Miao, S. Zhang, L. Ding, R. Bao, L. Zhang, and D. Tao · 2024
Later among the works it cites.
Iterative reasoning preference optimization
R. Y. Pang, W. Yuan, K. Cho, H. He, S. Sukhbaatar, and J. Weston · 2024
Later among the works it cites.
Rlhf from heterogeneous feedback via personalization and preference aggregation
C. Park, M. Liu, D. Kong, K. Zhang, and A. Ozdaglar · 2024
Later among the works it cites.
WARM: On the benefits of weight averaged reward models
A. Rame, N. Vieillard, L. Hussenot, R. Dadashi-Tazehozi, G. Cideron, O. Bachem, and J. Ferret · 2024
Later among the works it cites.
Towards understanding sycophancy in language models
M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, E. DURMUS, Z. Hatfield-Dodds, S. R. Johnston, S. M. Kravec, T. Maxwell, S. McCandlish, K. Ndousse, O. Rausch, N. Schiefer, D. Yan, M. Zhang, and E. Perez · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Perez, S. Ringer, K. Lukosiute, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, A. Jones, A. Chen, B. Mann, B. Israel, B. Seethor, C. McKinnon, C. Olah, D. Yan, D. Amodei, D. Amodei, D. Drain, D. Li, E. Tran-Johnson, G. Khundadze, J. Kernion, J. Landis, J. Kerr, J. Mueller, J. Hyun, J. Landau, K. Ndousse, L. Goldberg, L. Lovitt, M. Lucas, M. Sellitto, M. Zhang, N. Kingsland, N. Elhage, N. Joseph, N. Mercado, N. DasSarma, O. Rausch, R. Larson, S. McCandlish, S. Johnston, S. Kravec, S. El Showk, T. Lanham, T. Telleen-Lawton, T. Brown, T. Henighan, T. Hume, Y. Bai, Z. Hatfield-Dodds, J. Clark, S. R. Bowman, A. Askell, R. Grosse, D. Hernandez, D. Ganguli, E. Hubinger, N. Schiefer, and J. Kaplan · 2023
Cited alongside, same era.
Verbosity bias in preference labeling by large language models, 2023
K. Saito, A. Wachi, K. Wataoka, and Y. Akimoto · 2023
Cited alongside, same era.
Iterative dpo alignment
H. Tran, C. Glaze, and B. Hancock · 2023
Cited alongside, same era.
How far can camels go? exploring the state of instruction tuning on open resources
Y. Wang, H. Ivison, P. Dasigi, J. Hessel, T. Khot, K. Chandu, D. Wadden, K. MacMillan, N. A. Smith, I. Beltagy, and H. Hajishirzi · 2023
Cited alongside, same era.
J. Xu, A. Lee, S. Sukhbaatar, and J. Weston · 2023
Cited alongside, same era.
Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün, and S. Hooker · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica · 2024
Cited alongside, same era.
A long way to go: Investigating length correlations in RLHF, 2024
P. Singhal, T. Goyal, J. Xu, and G. Durrett · 2024
Later among the works it cites.
Fine-tuning language models for factuality
K. Tian, E. Mitchell, H. Yao, C. D. Manning, and C. Finn · 2024
Later among the works it cites.
W. Xiong, H. Dong, C. Ye, Z. Wang, H. Zhong, H. Ji, N. Jiang, and T. Zhang · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang, R. Peng, R. Men, R. Gao, R. Lin, S. Wang, S. Bai, S. Tan, T. Zhu, T. Li, T. Liu, W. Ge, X. Deng, X. Zhou, X. Ren, X. Zhang, X. Wei, X. Ren, X. Liu, Y. Fan, Y. Yao, Y. Zhang, Y. Wan, Y. Chu, Y. Liu, Z. Cui, Z. Zhang, Z. Guo, and Z. Fan · 2024
Later among the works it cites.
TS-align: A teacher-student collaborative framework for scalable iterative finetuning of large language models
C. Zhang, C. Tang, D. Chong, K. Shi, G. Tang, F. Jiang, and H. Li · 2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs, 2024
L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y. Sheng · 2024
Later among the works it cites.
Fine-tuning language models from human preferences, 2020
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang · 2025
Closest in time.
Reward shaping to mitigate reward hacking in rlhf
J. Fu, X. Zhao, C. Yao, H. Wang, Q. Han, and Y. Xiao · 2025
Closest in time.
Align to structure: Aligning large language models with structural information, 2025
Z. M. Kim, A. Ramachandran, F. Tavazoee, J.-K. Kim, O. Rokhlenko, and D. Kang · 2025
Closest in time.
The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, 2025
Y. Miao, S. Zhang, L. Ding, Y. Zhang, L. Zhang, and D. Tao · 2025
Closest in time.
C. Park, S. Han, X. Guo, A. Ozdaglar, K. Zhang, and J.-K. Kim · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2025
Closest in time.
Learning to plan & reason for evaluation with thinking-llm-as-a-judge, 2025
S. Saha, X. Li, M. Ghazvininejad, J. Weston, and T. Wang · 2025
Closest in time.
Language models learn to mislead humans via RLHF
J. Wen, R. Zhong, A. Khan, E. Perez, J. Steinhardt, M. Huang, S. R. Bowman, H. He, and S. Feng · 2025
Closest in time.
Self-rewarding language models, 2025
W. Yuan, R. Y. Pang, K. Cho, X. Li, S. Sukhbaatar, J. Xu, and J. Weston · 2025
Closest in time.