Fetching the paper…
Reading the bibliography…
Mitigating reward hacking--where AI systems misbehave due to flaws or misspecifications in their learning objectives--remains a key challenge in constructing capable and aligned models.
‘improving ratings’: audit in the british university system
Marilyn Strathern · 1997
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
Jette Randløv and Preben Alstrøm · 1998
Earlier work this paper cites.
Of rats, rice, and race: The great hanoi rat massacre, an episode in french colonial history
Michael G Vann · 2003
Earlier work this paper cites.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?, 2020
Alon Jacovi and Yoav Goldberg · 2004
Earlier work this paper cites.
Detecting spam web pages through content analysis
Alexandros Ntoulas, Marc Najork, Mark Manasse, and Dennis Fetterly · 2006
Earlier work this paper cites.
Where do rewards come from
Satinder Singh, Richard L Lewis, and Andrew G Barto · 2009
Earlier work this paper cites.
Lying, cheating, and teaching to the test
John Gilliom · 2010
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Faulty reward functions in the wild
Jack Clark and Dario Amodei · 2016
Earlier work this paper cites.
Behaviorism is not enough: better recommendations through listening to users
Michael D Ekstrand and Martijn C Willemsen · 2016
Earlier work this paper cites.
Fake it till you make it: Reputation, competition, and yelp review fraud
Michael Luca and Georgios Zervas · 2016
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
R Paulus · 2017
Earlier work this paper cites.
Data-efficient deep reinforcement learning for dexterous manipulation
Ivaylo Popov, Nicolas Heess, Timothy Lillicrap, Roland Hafner, Gabriel Barth-Maron, Matej Vecerik, Thomas Lampe, Yuval Tassa, Tom Erez, and Martin Riedmiller · 2017
Earlier work this paper cites.
Do expiring budgets lead to wasteful year-end spending? evidence from federal procurement
Jeffrey B Liebman and Neale Mahoney · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Bounties, grants, and market-making entrepreneurship
David S Lucas and Caleb S Fuller · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Five points for anger, one for a ‘like’: How facebook’s formula fostered rage and misinformation
Jeremy B. Merril and Will Oremus · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Internet-augmented dialogue generation
Mojtaba Komeili, Kurt Shuster, and Jason Weston · 2021
Earlier work this paper cites.
Eliciting latent knowledge: How to tell if your eyes deceive you, 2021
Paul Christiano, Ajeya Cotra, and Mark Xu · 2021
Earlier work this paper cites.
Defining and characterizing reward gaming
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Cited alongside, same era.
Measuring progress on scalable oversight for large language models
Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
From explicit cot to implicit cot: Learning to internalize cot step by step
Yuntian Deng, Yejin Choi, and Stuart Shieber · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Cited alongside, same era.
Solving math word problems with process- and outcome-based feedback, 2022
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins · 2022
Cited alongside, same era.
Selection-inference: Exploiting large language models for interpretable logical reasoning, 2022
Antonia Creswell, Murray Shanahan, and Irina Higgins · 2022
Cited alongside, same era.
Monitoring for deceptive alignment
Evan Hubinger · 2022
Cited alongside, same era.
Measuring faithfulness in chain-of-thought reasoning, 2023
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez · 2023
Cited alongside, same era.
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang · 2024
Later among the works it cites.
Introduction to ai safety, ethics, and society, 2024
Dan Hendrycks · 2024
Later among the works it cites.
Mechanistic interpretability for ai safety – a review, 2024
Leonard Bereska and Efstratios Gavves · 2024
Later among the works it cites.
Eliciting latent knowledge from quirky language models, 2024
Alex Mallen, Madeline Brumley, Julia Kharchenko, and Nora Belrose · 2024
Later among the works it cites.
Alignment faking in large language models, 2024
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, and Evan Hubinger · 2024
Later among the works it cites.
Sycophancy to subterfuge: Investigating reward-tampering in large language models, 2024
Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, and Evan Hubinger · 2024
Later among the works it cites.
Prover-verifier games improve legibility of llm outputs
Jan Hendrik Kirchner, Yining Chen, Harri Edwards, Jan Leike, Nat McAleese, and Yuri Burda · 2024
Later among the works it cites.
On measuring faithfulness or self-consistency of natural language explanations, 2024
Letitia Parcalabescu and Anette Frank · 2024
Later among the works it cites.
Correlated proxies: A new definition and improved mitigation for reward hacking
Cassidy Laidlaw, Shivam Singhal, and Anca Dragan · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Openai o3-mini system card, February 2025
OpenAI · 2025
Closest in time.
Gemini flash
DeepMind · 2025
Closest in time.
Claude 3.7 sonnet system card, February 2025
Anthropic · 2025
Closest in time.
Deliberative alignment: Reasoning enables safer language models, 2025
Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, Andrea Vallone, Hongyu Ren, Jason Wei, Hyung Won Chung, Sam Toyer, Johannes Heidecke, Alex Beutel, and Amelia Glaese · 2025
Closest in time.
Process-oriented learning with factored cognition, 2022
Ought · 2025
Closest in time.
Coup probes: Catching catastrophes with probes trained off-policy, 2023
Fabien Roger · 2025
Closest in time.
Probes catch sleeper agents, 2023
Anthropic · 2025
Closest in time.
Features as classifiers, 2024
Transformer Circuits · 2025
Closest in time.
Frontier models are capable of in-context scheming, 2025
Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn · 2025
Closest in time.