Fetching the paper…
Reading the bibliography…
Bilevel optimization recently has attracted increased interest in machine learning due to its many applications such as hyper-parameter optimization and meta learning.
Y. Censor and A. Lent, “An iterative row-action method for interval convex programming,” Journal of Optimization theory and Applications , vol. 34, no. 3, pp. 321–353, 1981
1981
Earlier work this paper cites.
Y. Censor and S. A. Zenios, “Proximal minimization algorithm withd-functions,” Journal of Optimization Theory and Applications , vol. 73, no. 3, pp. 451–464, 1992
1992
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
B. Colson, P. Marcotte, and G. Savard, “An overview of bilevel optimization,” Annals of operations research , vol. 153, no. 1, pp. 235–256, 2007
2007
Earlier work this paper cites.
G. Kunapuli, K. P. Bennett, J. Hu, and J.-S. Pang, “Classification model selection via bilevel programming,” Optimization Methods & Software , vol. 23, no. 4, pp. 475–489, 2008
2008
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization.” Journal of machine learning research , vol. 12, no. 7, 2011
2011
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Ghadimi, G. Lan, and H. Zhang, “Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization,” Mathematical Programming , vol. 155, no. 1-2, pp. 267–305, 2016
2016
Earlier work this paper cites.
L. Franceschi, P. Frasconi, S. Salzo, R. Grazzi, and M. Pontil, “Bilevel programming for hyperparameter optimization and meta-learning,” in International Conference on Machine Learning . PMLR, 2018, pp. 1568–1577
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
A. Shaban, C.-A. Cheng, N. Hatch, and B. Boots, “Truncated back-propagation for bilevel optimization,” in The 22nd International Conference on Artificial Intelligence and Statistics . PMLR, 2019, pp. 1723–1732
2019
Cited alongside, same era.
A. Cutkosky and F. Orabona, “Momentum-based variance reduction in non-convex sgd,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
P. Khanduri, S. Zeng, M. Hong, H.-T. Wai, Z. Wang, and Z. Yang, “A near-optimal algorithm for stochastic bilevel optimization via double-momentum,” Advances in Neural Information Processing Systems , vol. 34, pp. 30 271–30 283, 2021
2021
Closest in time.
2021
Closest in time.
J. Yang, K. Ji, and Y. Liang, “Provably faster algorithms for bilevel optimization,” Advances in Neural Information Processing Systems , vol. 34, pp. 13 670–13 682, 2021
2021
Closest in time.
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Chen, S. Liu, R. Sun, and M. Hong, “On the convergence of a class of adam-type algorithms for non-convex optimization,” in 7th International Conference on Learning Representations (ICLR) , 2019
2019
Cited alongside, same era.
R. Ward, X. Wu, and L. Bottou, “Adagrad stepsizes: Sharp convergence over nonconvex landscapes,” in International Conference on Machine Learning . PMLR, 2019, pp. 6677–6686
2019
Cited alongside, same era.
X. Li and F. Orabona, “On the convergence of stochastic gradient descent with adaptive stepsizes,” in The 22nd International Conference on Artificial Intelligence and Statistics . PMLR, 2019, pp. 983–992
2019
Cited alongside, same era.
2020
Cited alongside, same era.
J. Zhuang, T. Tang, Y. Ding, S. C. Tatikonda, N. Dvornek, X. Papademetris, and J. Duncan, “Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,” Advances in neural information processing systems , vol. 33, pp. 18 795–18 806, 2020
2020
Cited alongside, same era.
R. Grazzi, L. Franceschi, M. Pontil, and S. Salzo, “On the iteration complexity of hypergradient computation,” in International Conference on Machine Learning . PMLR, 2020, pp. 3748–3758
2020
Cited alongside, same era.
K. Ji, J. Yang, and Y. Liang, “Bilevel optimization: Convergence analysis and enhanced design,” in International Conference on Machine Learning . PMLR, 2021, pp. 4882–4892
2021
Cited alongside, same era.
R. Liu, J. Gao, J. Zhang, D. Meng, and Z. Lin, “Investigating bi-level optimization for learning and vision from a unified perspective: A survey and beyond,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2021
2021
Cited alongside, same era.
T. Chen, Y. Sun, and W. Yin, “Closing the gap: Tighter analysis of alternating stochastic gradient methods for bilevel problems,” Advances in Neural Information Processing Systems , vol. 34, pp. 25 294–25 307, 2021
2021
Closest in time.
R. Liu, Y. Liu, S. Zeng, and J. Zhang, “Towards gradient-based bilevel optimization with non-convex followers and beyond,” Advances in Neural Information Processing Systems , vol. 34, pp. 8662–8675, 2021
2021
Closest in time.
2021
Closest in time.
F. Huang, J. Li, and H. Huang, “Super-adam: faster and universal framework of adaptive gradients,” Advances in Neural Information Processing Systems , vol. 34, pp. 9074–9085, 2021
2021
Closest in time.
2021
Closest in time.
T. Chen, Y. Sun, Q. Xiao, and W. Yin, “A single-timescale method for stochastic bilevel optimization,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2022, pp. 2466–2488
2022
Closest in time.
R. Liu, P. Mu, X. Yuan, S. Zeng, and J. Zhang, “A general descent aggregation framework for gradient-based bi-level optimization,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2022
2022
Closest in time.
K. Ji and Y. Liang, “Lower bounds and accelerated algorithms for bilevel optimization,” Journal of Machine Learning Research , vol. 23, pp. 1–56, 2022
2022
Closest in time.