Fetching the paper…
Reading the bibliography…
We present a novel unified bilevel optimization-based framework, \textsf{PARL}, formulated to address the recently highlighted critical issue of policy alignment in reinforcement learning using utility or preference-based feedback.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
Mathematical programs with optimization problems in the constraints
Jerome Bracken and James T McGill · 1973
Earlier work this paper cites.
Mechanism design
Roger B Myerson · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Mechanism design without games
Leonid Hurwicz · 2003
Earlier work this paper cites.
Convex optimization
Stephen P Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Mechanism design: How to implement social goals
Eric S Maskin · 2008
Earlier work this paper cites.
Convergent temporal-difference learning with arbitrary smooth function approximation
Hamid R. Maei, Csaba Szepesvári, Shalabh Bhatnagar, Doina Precup, David Silver, and Richard S. Sutton · 2009
Earlier work this paper cites.
On integral probability metrics, ϕ \phi -divergences and binary classification, 2009
Bharath K. Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Gert R. G. Lanckriet · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
Real-time reinforcement learning by sequential actor–critics and experience replay
Paweł Wawrzyński · 2009
Earlier work this paper cites.
Finite adaptability in multistage linear optimization
Dimitris Bertsimas and Constantine Caramanis · 2010
Earlier work this paper cites.
Optimality of affine policies in multistage robust optimization
Dimitris Bertsimas, Dan A Iancu, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Market structure and equilibrium
Heinrich Von Stackelberg · 2010
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Brian D. Ziebart, J. Andrew Bagnell, and Anind K. Dey · 2010
Earlier work this paper cites.
Preference-based reinforcement learning: A formal framework and a policy iteration algorithm
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Earlier work this paper cites.
Learning economic parameters from revealed preferences
Maria-Florina Balcan, Amit Daniely, Ruta Mehta, Ruth Urner, and Vijay V Vazirani · 2014
Earlier work this paper cites.
Multistage stochastic optimization , volume 1104
Georg Ch Pflug and Alois Pichler · 2014
Earlier work this paper cites.
Revealed preference, rational inattention, and costly information acquisition
Andrew Caplin and Mark Dean · 2015
Earlier work this paper cites.
Contextual dueling bandits, 2015
Miroslav Dudík, Katja Hofmann, Robert E. Schapire, Aleksandrs Slivkins, and Masrour Zoghi · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models, 2016
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine · 2016
Earlier work this paper cites.
Generative adversarial imitation learning, 2016
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Model-free imitation learning with policy optimization, 2016
Jonathan Ho, Jayesh K. Gupta, and Stefano Ermon · 2016
Earlier work this paper cites.
Watch and learn: Optimizing from revealed preferences feedback
Aaron Roth, Jonathan Ullman, and Zhiwei Steven Wu · 2016
Earlier work this paper cites.
Inverse reinforcement learning from failure
Kyriacos Shiarlis, Joao Messias, and Shimon Whiteson · 2016
Earlier work this paper cites.
Maximum entropy deep inverse reinforcement learning, 2016
Markus Wulfmeier, Peter Ondruska, and Ingmar Posner · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Reinforcement mechanism design
Pingzhong Tang · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, Johannes Fürnkranz, et al · 2017
Cited alongside, same era.
Inference aided reinforcement learning for incentive mechanism design in crowdsourcing
Zehong Hu, Yitao Liang, Jie Zhang, Zhao Li, and Yang Liu · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari, 2018
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Policy optimization with demonstrations
Bingyi Kang, Zequn Jie, and Jiashi Feng · 2018
Cited alongside, same era.
Han Zhong, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2021
Later among the works it cites.
Projection-free stochastic bi-level optimization
Zeeshan Akhtar, Amrit Singh Bedi, Srujan Teja Thomdapu, and Ketan Rajawat · 2022
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Later among the works it cites.
On the hidden biases of policy mirror ascent in continuous action spaces
Amrit Singh Bedi, Souradip Chakraborty, Anjaly Parayil, Brian M Sadler, Pratap Tokekar, and Alec Koppel · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepmind control suite, 2018
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller · 2018
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations, 2019
Daniel S. Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Cited alongside, same era.
Deep hedging
Hans Buehler, Lukas Gonon, Josef Teichmann, and Ben Wood · 2019
Cited alongside, same era.
Convergence of learning dynamics in stackelberg games
Tanner Fiez, Benjamin Chasnov, and Lillian J Ratliff · 2019
Cited alongside, same era.
Parenting: Safe reinforcement learning from human input
Christopher Frye and Ilya Feige · 2019
Cited alongside, same era.
A divergence minimization perspective on imitation learning methods, 2019
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2019
Cited alongside, same era.
Souradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Pratap Tokekar, and Dinesh Manocha · 2022
Later among the works it cites.
A single-timescale method for stochastic bilevel optimization
Tianyi Chen, Yuejiao Sun, Quan Xiao, and Wotao Yin · 2022
Later among the works it cites.
Scaling laws for reward model overoptimization, 2022
Leo Gao, John Schulman, and Jacob Hilton · 2022
Later among the works it cites.
Zero-sum stochastic stackelberg games
Denizalp Goktas, Sadie Zhao, and Amy Greenwald · 2022
Later among the works it cites.
Enhanced bilevel optimization via bregman distance
Feihu Huang, Junyi Li, Shangqian Gao, and Heng Huang · 2022
Later among the works it cites.
Lower bounds and accelerated algorithms for bilevel optimization
Kaiyi Ji and Yingbin Liang · 2022
Later among the works it cites.
Aligning generative language models with human values
Ruibo Liu, Ge Zhang, Xinyu Feng, and Soroush Vosoughi · 2022
Later among the works it cites.
Rewards encoding environment dynamics improves preference-based reinforcement learning, 2022
Katherine Metcalf, Miguel Sarabia, and Barry-John Theobald · 2022
Later among the works it cites.
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2022
Later among the works it cites.
Surf: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning, 2022
Jongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2022
Later among the works it cites.
Stackelberg policy gradient: Evaluating the performance of leaders and followers
Quoc-Liem Vu, Zane Alumbaugh, Ryan Ching, Quanchen Ding, Arnav Mahajan, Benjamin Chasnov, Sam Burden, and Lillian J Ratliff · 2022
Later among the works it cites.
Htron:efficient outdoor navigation with sparse rewards via heavy tailed adaptive reinforce algorithm, 2022
Kasun Weerakoon, Souradip Chakraborty, Nare Karapetyan, Adarsh Jagan Sathyamoorthy, Amrit Singh Bedi, and Dinesh Manocha · 2022
Later among the works it cites.
Projection-free methods for stochastic simple bilevel optimization with convex lower-level problem, 2023
Jincheng Cao, Ruichen Jiang, Nazanin Abolfazli, Erfan Yazdandoost Hamedani, and Aryan Mokhtari · 2023
Closest in time.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Closest in time.
Ai alignment dialogues: An interactive approach to ai alignment in support agents, 2023
Pei-Yu Chen, Myrthe L. Tielman, Dirk K. J. Heylen, Catholijn M. Jonker, and M. Birna van Riemsdijk · 2023
Closest in time.
On momentum-based gradient methods for bilevel optimization with nonconvex lower-level
Feihu Huang · 2023
Closest in time.
A fully first-order method for stochastic bilevel optimization, 2023
Jeongyeol Kwon, Dohyun Kwon, Stephen Wright, and Robert Nowak · 2023
Closest in time.
Averaged method of multipliers for bi-level optimization without lower-level strong convexity
Risheng Liu, Yaohua Liu, Wei Yao, Shangzhi Zeng, and Jin Zhang · 2023
Closest in time.
Bilevel optimization with coupled decision-dependent distributions
Songtao Lu · 2023
Closest in time.
Convergent first-order methods for bi-level optimization and stackelberg games
Chinmay Maheshwari, S Shankar Sasty, Lillian Ratliff, and Eric Mazumdar · 2023
Closest in time.
Dueling rl: Reinforcement learning with trajectory preferences, 2023
Aldo Pacchiano, Aadirupa Saha, and Jonathan Lee · 2023
Closest in time.
Dueling rl: Reinforcement learning with trajectory preferences
Aadirupa Saha, Aldo Pacchiano, and Jonathan Lee · 2023
Closest in time.
Fundamental limitations of alignment in large language models, 2023
Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons, 2023
Banghua Zhu, Jiantao Jiao, and Michael I. Jordan · 2023
Closest in time.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences, 2024
Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Furong Huang, Dinesh Manocha, Amrit Singh Bedi, and Mengdi Wang · 2024
Closest in time.