Fetching the paper…
Reading the bibliography…
This paper proposes a new algorithm -- the \underline{S}ingle-timescale Do\underline{u}ble-momentum \underline{St}ochastic \underline{A}pprox\underline{i}matio\underline{n} (SUSTAIN) -- for tackling stochastic unconstrained bilevel optimization problems.
H. V. Stackelberg, The Theory of Market Economy . Oxford University Press, 1952
1952
Earlier work this paper cites.
J. Bracken and J. T. McGill, “Mathematical programs with optimization problems in the constraints,” Operations Research , vol. 21, no. 1, pp. 37–44, 1973
1973
Earlier work this paper cites.
——, “Defense applications of mathematical programs with optimization problems in the constraints,” Operations Research , vol. 22, no. 5, pp. 1086–1096, 1974. [Online]. Available: http://www.jstor.org/stable/169661
1974
Earlier work this paper cites.
J. Bracken, J. E. Falk, and J. T. McGill, “Technical note—the equivalence of two mathematical programs with optimization problems in the constraints,” Operations Research , vol. 22, no. 5, pp. 1102–1104, 1974
1974
Earlier work this paper cites.
W. Rudin, Principles of mathematical analysis , 3rd ed. McGraw-Hill New York, 1976
1976
Earlier work this paper cites.
D. J. White and G. Anandalingam, “A penalty function approach for solving bi-level linear programs,” Journal of Global Optimization , vol. 3, pp. 397–419, 1993
1993
Earlier work this paper cites.
L. Vicente, , G. Savard, and J. Júdice, “Descent approaches for quadratic bilevel programming,” Journal of Optimization Theory and Applications , pp. 379–399, 1994
1994
Earlier work this paper cites.
J. E. Falk and J. Liu, “On bilevel programming, part I: General nonlinear cases,” Mathematical Programming volume , vol. 70, pp. 47–72, 1995
1995
Earlier work this paper cites.
Z.-Q. Luo, J.-S. Pang, and D. Ralph, Mathematical Programs with Equilibrium Constraints . Cambridge University Press, 1996
1996
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms,” in Advances in neural information processing systems , 2000, pp. 1008–1014
2000
Earlier work this paper cites.
S. Dempe, Foundations of bilevel programming . Springer Science & Business Media, 2002
2002
Earlier work this paper cites.
B. Colson, P. Marcotte, and G. Savard, “An overview of bilevel optimization,” Annals of Operations Research , vol. 153, pp. 235–256, 2007
2007
Earlier work this paper cites.
2008
Earlier work this paper cites.
A. Migdalas, P. M. Pardalos, and P. Värbrand, Multilevel optimization: algorithms and applications . Springer Science & Business Media, 2013, vol. 20
2013
Earlier work this paper cites.
2014
Cited alongside, same era.
F. Pedregosa, “Hyperparameter optimization with approximate gradient,” in International conference on machine learning . PMLR, 2016, pp. 737–746
2016
Cited alongside, same era.
O. Vinyals, C. Blundell, T. Lillicrap, k. kavukcuoglu, and D. Wierstra, “Matching networks for one shot learning,” in Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29. Curran Associates, Inc., 2016. [Online]. Available: https://proceedings.neurips.cc/paper/2016/file/90e1357833654983612fb05e3ec9148c-Paper.pdf
2016
Cited alongside, same era.
L. Franceschi, M. Donini, P. Frasconi, and M. Pontil, “Forward and reverse gradient-based hyperparameter optimization,” in Proceedings of the 34th International Conference on Machine Learning - Volume 70 , 2017, p. 1165–1173
A. Shaban, C.-A. Cheng, N. Hatch, and B. Boots, “Truncated back-propagation for bilevel optimization,” 2019
2019
Later among the works it cites.
A. Raghu, M. Raghu, S. Bengio, and O. Vinyals, “Rapid learning or feature reuse? towards understanding the effectiveness of maml,” in ICLR , 2019
2019
Later among the works it cites.
A. Cutkosky and F. Orabona, “Momentum-based variance reduction in non-convex SGD,” in Advances in Neural Information Processing Systems 32 . Curran Associates, Inc., 2019, pp. 15 236–15 245
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Ji, J. Yang, and Y. Liang, “Bilevel optimization: Nonasymptotic analysis and faster algorithms,” 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
S. Sabach and S. Shtern, “A first order method for solving convex bilevel optimization problems,” SIAM J. Optim. , vol. 27, no. 2, pp. 640–660, 2017. [Online]. Available: https://doi.org/10.1137/16M105592X
2017
Cited alongside, same era.
S. Ravi and H. Larochelle, “Optimization as a model for few-shot learning,” in Proceedings of the 5th International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 2017, pp. 1126–1135
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
S. Ghadimi and M. Wang, “Approximation methods for bilevel programming,” 2018
2018
Cited alongside, same era.
C. Fang, C. J. Li, Z. Lin, and T. Zhang, “Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator,” in Advances in Neural Information Processing Systems , 2018, pp. 689–699
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2020
Later among the works it cites.
M. Hong, H.-T. Wai, Z. Wang, and Z. Yang, “A two-timescale framework for bilevel optimization: Complexity analysis and application to actor-critic,” 2020
2020
Later among the works it cites.
R. Liu, P. Mu, X. Yuan, S. Zeng, and J. Zhang, “A generic first-order algorithmic framework for bi-level programming beyond lower-level singleton,” 2020
2020
Later among the works it cites.
J. Li, B. Gu, and H. Huang, “Improved bilevel model: Fast and optimal algorithm with theoretical guarantee,” 2020
2020
Later among the works it cites.
R. Grazzi, M. Pontil, and S. Salzo, “Convergence properties of stochastic hypergradients,” 2020
2020
Later among the works it cites.
R. Grazzi, L. Franceschi, M. Pontil, and S. Salzo, “On the iteration complexity of hypergradient computation,” 2020
2020
Later among the works it cites.
2021
Closest in time.
Z. Guo and T. Yang, “Randomized stochastic variance-reduced methods for stochastic bilevel optimization,” 2021
2021
Closest in time.
R. Liu, J. Gao, J. Zhang, D. Meng, and Z. Lin, “Investigating bi-level optimization for learning and vision from a unified perspective: A survey and beyond,” 2021
2021
Closest in time.