Fetching the paper…
Reading the bibliography…
Future advanced AI systems may learn sophisticated strategies through reinforcement learning (RL) that humans cannot understand well enough to safely evaluate.
Reinforcement learning from human reward: Discounting in episodic tasks
W. B. Knox and P. Stone · 2012
Earlier work this paper cites.
Approval-directed agents, 2014
P. Christiano · 2014
Earlier work this paper cites.
Reinforcement learning and the reward engineering principle
D. Dewey · 2014
Earlier work this paper cites.
A toy model of the control problem, 2015
S. Armstrong · 2015
Earlier work this paper cites.
Corrigibility
N. Soares, B. Fallenstein, E. Yudkowsky, and S. Armstrong · 2015
Earlier work this paper cites.
Quantilizers: A safer alternative to maximizers for limited optimization
J. Taylor · 2015
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Faulty reward functions in the wild, 2016
J. Clark and D. Amodei · 2016
Earlier work this paper cites.
Early stopping as nonparametric variational inference
D. Duvenaud, D. Maclaurin, and R. Adams · 2016
Earlier work this paper cites.
Avoiding wireheading with value reinforcement learning
T. Everitt and M. Hutter · 2016
Earlier work this paper cites.
The dependence of effective planning horizon on model accuracy
N. Jiang, A. Kulesza, S. Singh, and R. L. Lewis · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
T. Everitt, V. Krakovna, L. Orseau, M. Hutter, and S. Legg · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
J. MacGlashan, M. K. Ho, R. Loftin, B. Peng, G. Wang, D. L. Roberts, M. E. Taylor, and M. L. Littman · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
What failure looks like, 2019
P. Christiano · 2019
Cited alongside, same era.
The unexpected difficulty of comparing AlphaStar to humans, 2019
R. Korzekwa · 2019
Cited alongside, same era.
Neurosymbolic reinforcement learning with formally verified exploration
G. Anderson, A. Verma, I. Dillig, and S. Chaudhuri · 2020
Cited alongside, same era.
Specification gaming: the flip side of AI ingenuity, 2020
V. Krakovna, J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg · 2020
Cited alongside, same era.
Arguments against myopic training, 2020
R. Ngo · 2020
Cited alongside, same era.
Avoiding tampering incentives in deep rl via decoupled approval
J. Uesato, R. Kumar, V. Krakovna, T. Everitt, R. Ngo, and S. Legg · 2020
A watermark for large language models
J. Kirchenbauer, J. Geiping, Y. Wen, J. Katz, I. Miers, and T. Goldstein · 2023
Later among the works it cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Later among the works it cites.
Preventing language models from hiding their reasoning
F. Roger and R. Greenblatt · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, N. Cheng, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, et al · 2023
Later among the works it cites.
AlphaMath almost zero: Process supervision without process
G. Chen, M. Liao, C. Li, and K. Fan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton · 2021
Cited alongside, same era.
Heuristic-guided reinforcement learning
C.-A. Cheng, A. Kolobov, and A. Swaminathan · 2021
Cited alongside, same era.
Stable-baselines3: Reliable reinforcement learning implementations
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
Estimating and penalizing induced preference shifts in recommender systems
M. Carroll, D. Hadfield-Menell, S. Russell, and A. Dragan · 2022
Cited alongside, same era.
Path-specific objectives for safer agent incentives
S. Farquhar, R. Carey, and T. Everitt · 2022
Cited alongside, same era.
Scalable watermarking for identifying large language model outputs
S. Dathathri, A. See, S. Ghaisas, P.-S. Huang, R. McAdam, J. Welbl, V. Bachani, A. Kaskasoli, R. Stanforth, T. Matejovicova, J. Hayes, N. Vyas, M. A. Merey, J. Brown-Cohen, R. Bunel, B. Balle, T. Cemgil, Z. Ahmed, K. Stacpoole, I. Shumailov, C. Baetu, S. Gowal, D. Hassabis, and P. Kohli · 2024
Later among the works it cites.
Sycophancy to subterfuge: Investigating reward-tampering in large language models
C. Denison, M. MacDiarmid, F. Barez, D. Duvenaud, S. Kravec, S. Marks, N. Schiefer, R. Soklaski, A. Tamkin, et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team · 2024
Later among the works it cites.
GLoRe: When, where, and how to improve LLM reasoning via global and local refinements
A. Havrilla, S. C. Raparthy, C. Nalmpantis, J. Dwivedi-Yu, M. Zhuravinskyi, E. Hambro, and R. Raileanu · 2024
Later among the works it cites.
Learning optimal advantage from preferences and mistaking it for reward
W. B. Knox, S. Hatgis-Kessell, S. O. Adalgeirsson, S. Booth, A. Dragan, P. Stone, and S. Niekum · 2024
Later among the works it cites.
Hidden in plain text: Emergence & mitigation of steganographic collusion in LLMs
Y. Mathew, O. Matthews, R. McCarthy, J. Velja, C. S. de Witt, D. Cope, and N. Schoots · 2024
Later among the works it cites.
Secret collusion among generative AI agents
S. R. Motwani, M. Baranchuk, M. Strohmeier, V. Bolina, P. H. Torr, L. Hammond, and C. S. de Witt · 2024
Later among the works it cites.
Multi-turn reinforcement learning with preference human feedback
L. Shani, A. Rosenberg, A. Cassel, O. Lang, D. Calandriello, A. Zipori, H. Noga, O. Keller, B. Piot, I. Szpektor, A. Hassidim, Y. Matias, and R. Munos · 2024
Later among the works it cites.
Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations
P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui · 2024
Later among the works it cites.
Language models learn to mislead humans via RLHF
J. Wen, R. Zhong, A. Khan, E. Perez, J. Steinhardt, M. Huang, S. R. Bowman, H. He, and S. Feng · 2024
Later among the works it cites.
On targeted manipulation and deception when optimizing LLMs for user feedback
M. Williams, M. Carroll, A. Narang, C. Weisser, B. Murphy, and A. Dragan · 2024
Later among the works it cites.
Rlhs: Mitigating misalignment in rlhf with hindsight simulation
K. Liang, H. Hu, R. Liu, T. L. Griffiths, and J. F. Fisac · 2025
Closest in time.