Fetching the paper…
Reading the bibliography…
Bilevel optimization, the problem of minimizing a value function which involves the arg-minimum of another function, appears in many areas of machine learning.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Theory of the market economy
Heinrich von Stackelberg · 1952
Earlier work this paper cites.
A simple automatic derivative evaluation program
R. Wengert · 1964
Earlier work this paper cites.
Taylor expansion of the accumulated rounding error
Seppo Linnainmaa · 1976
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C. Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Gradient-Based Optimization of Hyperparameters
Yoshua Bengio · 2000
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
IU E. Nesterov · 2004
Earlier work this paper cites.
Large-Scale Machine Learning with Stochastic Gradient Descent
Léon Bottou · 2010
Earlier work this paper cites.
Generic methods for optimization-based modeling
Justin Domke · 2012
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Numba: A llvm-based python jit compiler
Siu Kwan Lam, Antoine Pitrou, and Stanley Seibert · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient
Fabian Pedregosa · 2016
Cited alongside, same era.
Fast Incremental Method for Nonconvex Optimization
Sashank J. Reddi, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Cited alongside, same era.
Automatic differentiation in Machine Learning: A survey
Atilim Gunes Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind · 2018
Cited alongside, same era.
Optimization methods for large-scale Machine Learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
On the iteration complexity of hypergradient computation
Riccardo Grazzi, Luca Franceschi, Massimiliano Pontil, and Saverio Salzo · 2020
Later among the works it cites.
Projection-Free Algorithm for Stochastic Bi-level Optimization
Zeeshan Akhtar, Amrit Singh Bedi, Srujan Teja Thomdapu, and Ketan Rajawat · 2021
Later among the works it cites.
Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel Problems
Tianyi Chen, Yuejiao Sun, and Wotao Yin · 2021
Later among the works it cites.
Convergence properties of stochastic hypergradients
Riccardo Grazzi, Massimiliano Pontil, and Saverio Salzo · 2021
Later among the works it cites.
A Two-Timescale Framework for Bilevel Optimization: Complexity Analysis and Application to Actor-Critic
Mingyi Hong, Hoi-To Wai, Zhaoran Wang, and Zhuoran Yang · 2021
Later among the works it cites.
BiAdam: Fast Adaptive Bilevel Optimization Methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SPIDER: Near-Optimal Non-Convex Optimization via Stochastic Path Integrated Differential Estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
Bilevel programming for hyperparameter optimization and meta-learning
Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimiliano Pontil · 2018
Cited alongside, same era.
Approximation Methods for Bilevel Programming
Saeed Ghadimi and Mengdi Wang · 2018
Cited alongside, same era.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Cited alongside, same era.
Deep Equilibrium Models
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2019
Cited alongside, same era.
AutoAugment: Learning Augmentation Strategies From Data
Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le · 2019
Cited alongside, same era.
Feihu Huang and Heng Huang · 2021
Later among the works it cites.
Bilevel optimization: Convergence analysis and enhanced design
Kaiyi Ji, Junjie Yang, and Yingbin Liang · 2021
Later among the works it cites.
A Near-Optimal Algorithm for Stochastic Bilevel Optimization via Double-Momentum
Prashant Khanduri, Siliang Zeng, Mingyi Hong, Hoi-To Wai, Zhaoran Wang, and Zhuoran Yang · 2021
Later among the works it cites.
Provably Faster Algorithms for Bilevel Optimization
Junjie Yang, Kaiyi Ji, and Yingbin Liang · 2021
Later among the works it cites.
Amortized Implicit Differentiation for Stochastic Bilevel Optimization
Michael Arbel and Julien Mairal · 2022
Closest in time.
A Single-Timescale Stochastic Bilevel Optimization Method
Tianyi Chen, Yuejiao Sun, and Wotao Yin · 2022
Closest in time.
A Fully Single Loop Algorithm for Bilevel Optimization without Hessian Inverse
Junyi Li, Bin Gu, and Heng Huang · 2022
Closest in time.
Benchopt: Reproducible, efficient and collaborative optimization benchmarks
Thomas Moreau, Mathurin Massias, Alexandre Gramfort, Pierre Ablin, Pierre-Antoine Bannier Benjamin Charlier, Mathieu Dagréou, Tom Dupré la Tour, Ghislain Durif, Cassio F. Dantas, Quentin Klopfenstein, Johan Larsson, En Lai, Tanguy Lefort, Benoit Malézieux, Badr Moufad, Binh T. Nguyen, Alain Rakotomamonjy, Zaccharie Ramzi, Joseph Salmon, and Samuel Vaiter · 2022
Closest in time.
SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit models
Zaccharie Ramzi, Florian Mannel, Shaojie Bai, Jean-Luc Starck, Philippe Ciuciu, and Thomas Moreau · 2022
Closest in time.
CADDA: Class-wise Automatic Differentiable Data Augmentation for EEG Signals
Cédric Rommel, Thomas Moreau, Joseph Paillard, and Alexandre Gramfort · 2022
Closest in time.