Fetching the paper…
Reading the bibliography…
Many important machine learning applications involve regularized nonconvex bi-level optimization.
A topological property of real analytic subsets
Łojasiewicz, S · 1963
Earlier work this paper cites.
Mathematical programs with optimization problems in the constraints
Bracken, J. and McGill, J. T · 1973
Earlier work this paper cites.
Splitting algorithms for the sum of two nonlinear operators
Lions, P.-L. and Mercier, B · 1979
Earlier work this paper cites.
New branch-and-bound rules for linear bilevel programming
Hansen, P., Jaumard, B., and Savard, G · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Multi-step model-agnostic meta-learning: Convergence and improved algorithms
Ji, K., Yang, J., and Liang, Y · 2002
Earlier work this paper cites.
An extended kuhn–tucker approach for linear bilevel programming
Shi, C., Lu, J., and Zhang, G · 2005
Earlier work this paper cites.
Convergence of meta-learning with task-specific adaptation over partial parameters
Ji, K., Lee, J. D., Liang, Y., and Poor, H. V · 2006
Earlier work this paper cites.
The Łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems
Bolte, J., Daniilidis, A., and Lewis, A · 2007
Earlier work this paper cites.
On the convergence of the proximal algorithm for nonsmooth functions involving analytic features
Attouch, H. and Bolte, J · 2009
Earlier work this paper cites.
Variational analysis , volume 317
Rockafellar, R. T. and Wets, R. J.-B · 2009
Earlier work this paper cites.
Bilevel programming algorithms for machine learning model selection
Moore, G. M · 2010
Earlier work this paper cites.
Generic methods for optimization-based modeling
Domke, J · 2012
Earlier work this paper cites.
Convergence of linesearch and trust-region methods using the Kurdyka–Łojasiewicz inequality
Noll, D. and Rondepierre, A · 2013
Earlier work this paper cites.
Proximal alternating linearized minimization for nonconvex and nonsmooth problems
Bolte, J., Sabach, S., and Teboulle, M · 2014
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 2014
Earlier work this paper cites.
Splitting methods with variable metric for Kurdyka–Łojasiewicz functions and general convergence rates
Frankel, P., Garrigos, G., and Peypouquet, J · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Cited alongside, same era.
Gould, S., Fernando, B., Cherian, A., Anderson, P., Cruz, R. S., and Guo, E · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Cited alongside, same era.
Hyperparameter optimization with approximate gradient
Pedregosa, F · 2016
Cited alongside, same era.
Truncated back-propagation for bilevel optimization
Shaban, A., Cheng, C.-A., Hatch, N., and Boots, B · 2019
Later among the works it cites.
Adversarial attacks on graph neural networks via meta learning
Zügner, D. and Günnemann, S · 2019
Later among the works it cites.
On the convergence theory of gradient-based model-agnostic meta-learning algorithms
Fallah, A., Mokhtari, A., and Ozdaglar, A · 2020
Later among the works it cites.
On the iteration complexity of hypergradient computation
Grazzi, R., Franceschi, L., Pontil, M., and Salzo, S · 2020
Later among the works it cites.
Hong, M., Wai, H.-T., Wang, Z., and Yang, Z · 2020
Later among the works it cites.
Improved bilevel model: Fast and optimal algorithm with theoretical guarantee
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhou, Y., Yu, Y., Dai, W., Liang, Y., and Xing, E · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M · 2017
Cited alongside, same era.
Convergence analysis of proximal gradient with momentum for nonconvex optimization
Li, Q., Zhou, Y., Liang, Y., and Varshney, P. K · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R · 2017
Cited alongside, same era.
Meta-learning with differentiable closed-form solvers
Bertinetto, L., Henriques, J. F., Torr, P., and Vedaldi, A · 2018
Cited alongside, same era.
Bilevel programming for hyperparameter optimization and meta-learning
Franceschi, L., Frasconi, P., Salzo, S., Grazzi, R., and Pontil, M · 2018
Cited alongside, same era.
Li, J., Gu, B., and Huang, H · 2020
Later among the works it cites.
On gradient descent ascent for nonconvex-concave minimax problems
Lin, T., Jin, C., and Jordan, M. I · 2020
Later among the works it cites.
A generic first-order algorithmic framework for bi-level programming beyond lower-level singleton
Liu, R., Mu, P., Yuan, X., Zeng, S., and Zhang, J · 2020
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Lorraine, J., Vicol, P., and Duvenaud, D · 2020
Later among the works it cites.
How robust are randomized smoothing based defenses to data poisoning?
Mehra, A., Kailkhura, B., Chen, P.-Y., and Hamm, J · 2020
Later among the works it cites.
Proximal gradient algorithm with momentum and flexible parameter restart for nonconvex optimization
Zhou, Y., Wang, Z., Ji, K., Liang, Y., and Tarokh, V · 2020
Later among the works it cites.
Randomized stochastic variance-reduced methods for stochastic bilevel optimization
Guo, Z. and Yang, T · 2021
Later among the works it cites.
Enhanced bilevel optimization via bregman distance
Huang, F. and Huang, H · 2021
Later among the works it cites.
Bilevel optimization for machine learning: Algorithm design and convergence analysis
Ji, K · 2021
Later among the works it cites.
Lower bounds and accelerated algorithms for bilevel optimization
Ji, K. and Liang, Y · 2021
Later among the works it cites.
Bilevel optimization: Convergence analysis and enhanced design
Ji, K., Yang, J., and Liang, Y · 2021
Later among the works it cites.
Provably faster algorithms for bilevel optimization
Yang, J., Ji, K., and Liang, Y · 2021
Later among the works it cites.