Fetching the paper…
Reading the bibliography…
Previous work has shown that training "helpful-only" LLMs with reinforcement learning on a curriculum of gameable environments can lead models to generalize to egregious specification gaming, such as editing their own reward function or modifying task checklists to appear more successful.
Using interactive feedback to improve the accuracy and explainability of question answering systems post-deployment
Zichao Li, Prakhar Sharma, Xing Han Lu, Jackie Cheung, and Siva Reddy · 2022
Earlier work this paper cites.
Training language models with language feedback, 2022
Jérémy Scheurer, Jon Ander Campos, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez · 2022
Earlier work this paper cites.
Can language models learn from explanations in context?, 2022
Andrew K. Lampinen, Ishita Dasgupta, Stephanie C. Y. Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L. McClelland, Jane X. Wang, and Felix Hill · 2022
Earlier work this paper cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback, 2023
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Earlier work this paper cites.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Earlier work this paper cites.
Reflection-tuning: Data recycling improves llm instruction-tuning, 2023
Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Heng Huang, Jiuxiang Gu, and Tianyi Zhou · 2023
Earlier work this paper cites.
Investigating the effectiveness of task-agnostic prefix prompt for instruction following, 2023
Seonghyeon Ye, Hyeonbin Hwang, Sohee Yang, Hyeongu Yun, Yireun Kim, and Minjoon Seo · 2023
Cited alongside, same era.
Principle-driven self-alignment of language models from scratch with minimal human supervision, 2023
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan · 2023
Cited alongside, same era.
Sycophancy to subterfuge: Investigating reward-tampering in large language models, 2024
Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, and Evan Hubinger · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
Abhimanyu Dubey et al · 2024
Cited alongside, same era.
The unlocking spell on base LLMs: Rethinking alignment via in-context learning
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi · 2024
Closest in time.
Isr-llm: Iterative self-refined large language model for long-horizon sequential task planning, 2024
Zhehua Zhou, Jiayang Song, Kunpeng Yao, Zhan Shu, and Lei Ma · 2024
Closest in time.
Generative ai for synthetic data generation: Methods, challenges and the future
Xu Guo and Yiqiang Chen · 2024
Closest in time.
Faulty reward functions in the wild
OpenAI · 2024
Closest in time.
A.i. learns to walk
Code Bullet · 2024
Closest in time.
Prompt caching with claude, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Llm2llm: Boosting llms with novel iterative data enhancement
Nicholas Lee, Thanakul Wattanawong, Sehoon Kim, Karttikeya Mangalam, Sheng Shen, Gopala Krishna Anumanchipalli, Michael W. Mahoney, Kurt Keutzer, and Amir Gholami · 2024
Cited alongside, same era.
Self-judge: Selective instruction following with alignment self-evaluation
Hai Ye and Hwee Tou Ng · 2024
Cited alongside, same era.
Specification gaming: the flip side of ai ingenuity
DeepMind · 2024
Cited alongside, same era.
Openai o1 system card
OpenAI · 2024
Cited alongside, same era.
Metareflection: Learning instructions for language agents using past reflections, 2024
Priyanshu Gupta, Shashank Kirtania, Ananya Singha, Sumit Gulwani, Arjun Radhakrishna, Sherry Shi, and Gustavo Soares · 2024
Cited alongside, same era.
Training language models with language feedback at scale, 2024a
Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez
Cited in the paper.
Feedback loops with language models drive in-context reward hacking, 2024a
Alexander Pan, Erik Jones, Meena Jagadeesan, and Jacob Steinhardt
Cited in the paper.
Spontaneous reward hacking in iterative self-refinement, 2024b
Jane Pan, He He, Samuel R. Bowman, and Shi Feng
Cited in the paper.
Anthropic · 2024
Closest in time.
Context caching — gemini api, 2024
Google · 2024
Closest in time.
Honesty to suberfuge - github repository, 2024
Leo Mckee-Reid, Christoph Sträter, Maria Martinez, Joe Needham, and Mikita Balesni · 2024
Closest in time.
Learning to reason with llms, 2024b
OpenAI · 2024
Closest in time.