Fetching the paper…
Reading the bibliography…
Bilevel optimization has been recently applied to many machine learning tasks.
The Theory of Market Economy
H. Stackelberg · 1952
Earlier work this paper cites.
Stochastic games
L. Shapley · 1953
Earlier work this paper cites.
Note on non-cooperative convex games
H. Nikaidô and K. Isoda · 1955
Earlier work this paper cites.
Generalized gradients and applications
F. H. Clarke · 1975
Earlier work this paper cites.
Optimization and non-smooth analysis
F. Clarke · 1983
Earlier work this paper cites.
Mathematical programs with equilibrium constraints
Z. Luo, J. Pang, and D. Ralph · 1996
Earlier work this paper cites.
The actor-critic algorithm as multi-time-scale stochastic approximation
V. Borkar and V. Konda · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P. L Bartlett · 2001
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
M. Littman · 2001
Earlier work this paper cites.
Sensitivity and convergence of uniformly ergodic markov chains
A Y. Mitrophanov · 2005
Earlier work this paper cites.
Implicit functions and solution mappings: A view from variational analysis , volume 616
A. L Dontchev and R T. Rockafellar · 2009
Earlier work this paper cites.
Optimization reformulations of the generalized nash equilibrium problem using nikaido-isoda-type functions
A. Von Heusinger and C. Kanzow · 2009
Earlier work this paper cites.
The exact penalty principle
J. Ye · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
D. Maclaurin, D. Duvenaud, and R. Adams · 2015
Earlier work this paper cites.
Caching incentive design in wireless d2d networks: A stackelberg game approach
Z. Chen, Y. Liu, B. Zhou, and M. Tao · 2016
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
S. Ghadimi, G. Lan, and H. Zhang · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient
F. Pedregosa · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
L. Franceschi, M. Donini, P. Frasconi, and M. Pontil · 2017
Earlier work this paper cites.
A first order method for solving convex bilevel optimization problems
S. Sabach and S. Shtern · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
Bilevel programming for hyperparameter optimization and meta-learning
L. Franceschi, P. Frasconi, S. Salzo, R. Grazzi, and M. Pontil · 2018
Earlier work this paper cites.
Approximation methods for bilevel programming
S. Ghadimi and M. Wang · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
A. Nichol, J. Achiam, and J. Schulman · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
J. Bhandari and D. Russo · 2019
Cited alongside, same era.
On the finite-time convergence of actor-critic algorithm
S. Qiu, Z. Yang, J. Ye, and Z. Wang · 2019
Cited alongside, same era.
Fast global convergence of natural policy gradient methods with entropy regularization
S. Cen, C. Cheng, Y. Chen, Y. Wei, and Y. Chi · 2022
Later among the works it cites.
Independent policy gradient for large-scale markov potential games: Sharper rates, function approximation, and game-agnostic convergence
D. Ding, C. Wei, K. Zhang, and M. Jovanovic · 2022
Later among the works it cites.
Inexact bilevel stochastic gradient methods for constrained and unconstrained lower-level problems
T. Giovannelli, G. Kent, and L. Vicente · 2022
Later among the works it cites.
Will bilevel optimizers benefit from loops
K. Ji, M. Liu, Y. Liang, and L. Ying · 2022
Later among the works it cites.
A fully single loop algorithm for bilevel optimization without hessian inverse
J. Li, B. Gu, and H. Huang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Rajeswaran, C. Finn, S. Kakade, and S. Levine · 2019
Cited alongside, same era.
A perspective on incentive design: Challenges and opportunities
L. J Ratliff, R. Dong, S. Sekar, and T. Fiez · 2019
Cited alongside, same era.
Truncated back-propagation for bilevel optimization
A. Shaban, C. Cheng, N. Hatch, and By. Boots · 2019
Cited alongside, same era.
Global convergence of policy gradient methods to (almost) locally optimal policies
K. Zhang, A. Koppel, H. Zhu, and T. Başar · 2019
Cited alongside, same era.
Reward shaping via meta-learning
H. Zou, T. Ren, D. Yan, H. Su, and J. Zhu · 2019
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2020
Cited alongside, same era.
Coresets via bilevel optimization for continual learning and streaming
Z. Borsos, M. Mutny, and A. Krause · 2020
Cited alongside, same era.
A general descent aggregation framework for gradient-based bi-level optimization
R. Liu, P. Mu, X. Yuan, S. Zeng, and J. Zhang · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
A single-timescale analysis for stochastic approximation with multiple coupled sequences
H. Shen and T. Chen · 2022
Later among the works it cites.
Bome! bilevel optimization made easy: A simple first-order approach
M. Ye, B. Liu, S. Wright, P. Stone, and Q. Liu · 2022
Later among the works it cites.
The ai economist: Taxation policy design via two-level deep multiagent reinforcement learning
S. Zheng, A. Trott, S. Srinivasa, D. Parkes, and R. Socher · 2022
Later among the works it cites.
Optimal algorithms for stochastic bilevel optimization under relaxed smoothness conditions
X. Chen, T. Xiao, and K. Balasubramanian · 2023
Later among the works it cites.
A two-timescale framework for bilevel optimization: Complexity analysis and application to actor-critic
M. Hong, H.-T. Wai, Z. Wang, and Z. Yang · 2023
Later among the works it cites.
On penalty methods for nonconvex bilevel optimization and first-order stochastic approximation
J. Kwon, D. Kwon, S. Wright, and R. Nowak · 2023
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
G. Lan · 2023
Later among the works it cites.
First-order penalty methods for bilevel optimization
Z. Lu and S. Mei · 2023
Later among the works it cites.
Decentralized robust v-learning for solving markov games with model uncertainty
S. Ma, Z. Chen, S. Zou, and Y. Zhou · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D Manning, S. Ermon, and C. Finn · 2023
Later among the works it cites.
On penalty-based bilevel gradient descent method
H. Shen and T. Chen · 2023
Later among the works it cites.
Towards understanding asynchronous advantage actor-critic: Convergence and linear speedup
H. Shen, K. Zhang, M. Hong, and T. Chen · 2023
Later among the works it cites.
Can we find nash equilibria at a linear rate in markov games?
Z. Song, J. Lee, and Z. Yang · 2023
Later among the works it cites.
Achieving 𝒪 ( ϵ − 1.5 ) \mathcal{O}(\epsilon^{-1.5}) complexity in hessian/jacobian-free stochastic bilevel optimization
Y. Yang, P. Xiao, and K. Ji · 2023
Later among the works it cites.
Policy mirror descent for regularized reinforcement learning: A generalized framework with linear convergence
W. Zhan, S. Cen, B. Huang, Y. Chen, J. D Lee, and Y. Chi · 2023
Later among the works it cites.
PARL: A unified framework for policy alignment in reinforcement learning
S. Chakraborty, A. Bedi, A. Koppel, D. Manocha, H. Wang, M. Wang, and F. Huang · 2024
Closest in time.
Last-iterate convergent policy gradient primal-dual methods for constrained mdps
D. Ding, C. Wei, K. Zhang, and A. Ribeiro · 2024
Closest in time.
Slm: A smoothed first-order lagrangian method for structured constrained nonconvex optimization
S. Lu · 2024
Closest in time.
E. Zhang, S. Zhao, T. Wang, S. Hossain, H. Gasztowtt, S. Zheng, D. Parkes, M. Tambe, and Y. Chen · 2024
Closest in time.