Fetching the paper…
Reading the bibliography…
Emerging research in Pluralistic Artificial Intelligence (AI) alignment seeks to address how intelligent systems can be designed and deployed in accordance with diverse human needs and values.
Multi-Objective Reinforcement Learning using Sets of Pareto Dominating Policies
Kristof Van Moffaert and Ann Nowé · 2014
Earlier work this paper cites.
Concrete Problems in AI Safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Steering approaches to Pareto-optimal multiobjective reinforcement learning
Peter Vamplew, Rustam Issabekov, Richard Dazeley, Cameron Foale, Adam Berry, Tim Moore, and Douglas Creighton · 2016
Earlier work this paper cites.
Interactive Learning from Policy-Dependent Human Feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman · 2017
Earlier work this paper cites.
Reinforcement learning: an Introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Human-aligned artificial intelligence is a multiobjective problem
Peter Vamplew, Richard Dazeley, Cameron Foale, Sally Firmin, and Jane Mummery · 2018
Earlier work this paper cites.
Dynamic Weights in Multi-Objective Deep Reinforcement Learning
Axel Abels, Diederik M Roijers, Tom Lenaerts, Ann Nowé, and Denis Steckelmacher · 2019
Earlier work this paper cites.
Survey on Applications of Multi-Armed and Contextual Bandits
Djallel Bouneffouf, Irina Rish, and Charu Aggarwal · 2020
Earlier work this paper cites.
Artificial Intelligence, Values, and Alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan · 2020
Earlier work this paper cites.
Alignment for Advanced Machine Learning Systems
Jessica Taylor, Eliezer Yudkowsky, Patrick LaVictoire, and Andrew Critch · 2020
Cited alongside, same era.
A Reinforcement Learning Based Cognitive Empathy Framework for Social Robots
Elahe Bagheri, Oliver Roesler, Hoang Long Cao, and Bram Vanderborght · 2021
Cited alongside, same era.
A Practical Guide to Multi-Objective Reinforcement Learning and Planning
Conor F. Hayes, Roxana Rădulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M. Zintgraf, Richard Dazeley, Fredrik Heintz, Enda Howley, Athirai A. Irissappane, Patrick Mannion, Ann Nowé, Gabriel Ramos, Marcello Restelli, Peter Vamplew, and Diederik M. Roijers · 2021
Cited alongside, same era.
Multi-Objective Decision Making for Trustworthy AI
Patrick Mannion, Fredrik Heintz, Thommen George Karimpanal, and Peter Vamplew · 2021
Cited alongside, same era.
Active learning in robotics: A review of control principles
Annalisa T. Taylor, Thomas A. Berrueta, and Todd D. Murphey · 2021
Cited alongside, same era.
AI apology: interactive multi-objective reinforcement learning for human-aligned AI
Hadassah Harland, Richard Dazeley, Bahareh Nakisa, Francisco Cruz, and Peter Vamplew · 2023
Later among the works it cites.
AI Alignment: A Comprehensive Survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao · 2023
Later among the works it cites.
RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash · 2023
Later among the works it cites.
AI Alignment and Social Choice: Fundamental Limitations and Policy Implications
Abhilash Mishra · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan · 2022
Cited alongside, same era.
A Survey on In-context Learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, and Zhifang Sui · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell, Peter Welinder Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
MORAL: Aligning AI with Human Norms through Multi-Objective Reinforced Active Learning
Markus Peschl, Arkady Zgonnikov, Frans A. Oliehoek, and Luciano C. Siebert · 2022
Cited alongside, same era.
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Cornell Tech, Jérémy Scheurer, Apollo Research, Javier Rando, Eth Zurich, Rachel Freedman, Berkeley Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Micah Carroll, Andi Peng, Phillip Christoffersen, Stewart Slocum, Mit Csail, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Erdem Bıyık, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell · 2023
Cited alongside, same era.
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
Taylor Sorensen, Liwei Jiang, Jena Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, and Yejin Choi
Cited in the paper.
Position: A Roadmap to Pluralistic Alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi
Cited in the paper.
Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Alexandre Rame, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, and Matthieu Cord · 2023
Later among the works it cites.
Towards Scalable Automated Alignment of LLMs: A Survey
Boxi Cao, Keming Lu, Xinyu Lu, Jiawei Chen, Mengjie Ren, Hao Xiang, Peilin Liu, Yaojie Lu, Ben He, Xianpei Han, Le Sun, Hongyu Lin, and Bowen Yu · 2024
Closest in time.
Multi-Objective Reinforcement Learning
Florian Felten · 2024
Closest in time.
Algorithmic Pluralism: A Structural Approach To Equal Opportunity
Shomik Jain, Vinith Suriyakumar, Kathleen Creel, and Ashia Wilson · 2024
Closest in time.
A sociotechnical system perspective on AI
Olya Kudina and Ibo van de Poel · 2024
Closest in time.
Rui Yang, Xiaoman Pan, Feng Luo, Shuang Qiu, Han Zhong, Dong Yu, and Jianshu Chen · 2024
Closest in time.