Fetching the paper…
Reading the bibliography…
Bilevel optimization has become a powerful framework in various machine learning applications including meta-learning, hyperparameter optimization, and network architecture search.
Mean value theorems for vector valued functions
Robert M McLeod · 1965
Earlier work this paper cites.
Mathematical programs with optimization problems in the constraints
Jerome Bracken and James T McGill · 1973
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Y Bengio, S Bengio, and J Cloutier · 1991
Earlier work this paper cites.
New branch-and-bound rules for linear bilevel programming
Pierre Hansen, Brigitte Jaumard, and Gilles Savard · 1992
Earlier work this paper cites.
Meta-neural networks that learn by learning
Devang K Naik and Richard J Mammone · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Some bounds on the complexity of gradients, jacobians, and hessians
Andreas Griewank · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
An extended kuhn–tucker approach for linear bilevel programming
Chenggen Shi, Jie Lu, and Guangquan Zhang · 2005
Earlier work this paper cites.
Efficient multiple hyperparameter learning for log-linear models
Chuan-sheng Foo, Chuong B Do, and Andrew Y Ng · 2008
Earlier work this paper cites.
Classification model selection via bilevel programming
Gautam Kunapuli, Kristin P Bennett, Jing Hu, and Jong-Shi Pang · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Bilevel programming algorithms for machine learning model selection
Gregory M Moore · 2010
Earlier work this paper cites.
Generic methods for optimization-based modeling
Justin Domke · 2012
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Learning constrained task similarities in graphregularized multi-task learning
Rémi Flamary, Alain Rakotomamonjy, and Gilles Gasso · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Gregory Koch, Richard Zemel, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Earlier work this paper cites.
Stephen Gould, Basura Fernando, Anoop Cherian, Peter Anderson, Rodrigo Santa Cruz, and Edison Guo · 2016
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient
Fabian Pedregosa · 2016
Earlier work this paper cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2016
Earlier work this paper cites.
Watch and learn: Optimizing from revealed preferences feedback
Aaron Roth, Jonathan Ullman, and Zhiwei Steven Wu · 2016
Earlier work this paper cites.
Meta-learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, and Daan Wierstra · 2016
Earlier work this paper cites.
Regret bounds for lifelong learning
Pierre Alquier, Massimiliano Pontil, et al · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
One-shot visual imitation learning via meta-learning
Chelsea Finn, Tianhe Yu, Tianhao Zhang, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Earlier work this paper cites.
Meta-SGD: Learning to learn quickly for few-shot learning
Zhenguo Li, Fengwei Zhou, Fei Chen, and Hang Li · 2017
Earlier work this paper cites.
Meta networks
Tsendsuren Munkhdalai and Hong Yu · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Earlier work this paper cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Earlier work this paper cites.
Meta-learning with differentiable closed-form solvers
Luca Bertinetto, Joao F Henriques, Philip Torr, and Andrea Vedaldi · 2018
Earlier work this paper cites.
Federated meta-learning for recommendation
Fei Chen, Zhenhua Dong, Zhenguo Li, and Xiuqiang He · 2018
Earlier work this paper cites.
Incremental learning-to-learn with statistical guarantees
Giulia Denevi, Carlo Ciliberto, Dimitris Stamos, and Massimiliano Pontil · 2018
Earlier work this paper cites.
Learning to learn around a common mean
Giulia Denevi, Carlo Ciliberto, Dimitris Stamos, and Massimiliano Pontil · 2018
Earlier work this paper cites.
SPIDER: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Chelsea Finn and Sergey Levine · 2018
Cited alongside, same era.
Probabilistic model-agnostic meta-learning
Chelsea Finn, Kelvin Xu, and Sergey Levine · 2018
Cited alongside, same era.
DiCE: The infinitely differentiable monte carlo estimator
Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric Xing, and Shimon Whiteson · 2018
Cited alongside, same era.
Bilevel programming for hyperparameter optimization and meta-learning
Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimiliano Pontil · 2018
Cited alongside, same era.
Truncated back-propagation for bilevel optimization
Amirreza Shaban, Ching-An Cheng, Nathan Hatch, and Byron Boots · 2019
Later among the works it cites.
SpiderBoost and momentum: Faster variance reduction algorithms
Zhe Wang, Kaiyi Ji, Yi Zhou, Yingbin Liang, and Vahid Tarokh · 2019
Later among the works it cites.
On lower iteration complexity bounds for the saddle point problems
Junyu Zhang, Mingyi Hong, and Shuzhong Zhang · 2019
Later among the works it cites.
Efficient meta learning via minibatch proximal update
Pan Zhou, Xiaotong Yuan, Huan Xu, Shuicheng Yan, and Jiashi Feng · 2019
Later among the works it cites.
Adversarial attacks on graph neural networks via meta learning
Daniel Zügner and Stephan Günnemann · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saeed Ghadimi and Mengdi Wang · 2018
Cited alongside, same era.
Recasting gradient-based meta-learning as hierarchical bayes
Erin Grant, Chelsea Finn, Sergey Levine, Trevor Darrell, and Thomas Griffiths · 2018
Cited alongside, same era.
Deep bilevel learning
Simon Jenni and Paolo Favaro · 2018
Cited alongside, same era.
Online gradient-based mixtures for transfer modulation in meta-learning
Ghassen Jerfel, Erin Grant, Thomas L Griffiths, and Katherine Heller · 2018
Cited alongside, same era.
Minimax estimation of neural net distance
Kaiyi Ji and Yingbin Liang · 2018
Cited alongside, same era.
Asymptotic miss ratio of lru caching with consistent hashing
Kaiyi Ji, Guocong Quan, and Jian Tan · 2018
Cited alongside, same era.
Reviving and improving recurrent back-propagation
Renjie Liao, Yuwen Xiong, Ethan Fetaya, Lisa Zhang, KiJung Yoon, Xaq Pitkow, Raquel Urtasun, and Richard Zemel · 2018
Cited alongside, same era.
Provable representation learning for imitation learning via bi-level optimization
Sanjeev Arora, Simon S Du, Sham Kakade, Yuping Luo, and Nikunj Saunshi · 2020
Later among the works it cites.
Delta-STN: Efficient bilevel optimization for neural networks using structured response Jacobians
Juhan Bae and Roger Grosse · 2020
Later among the works it cites.
Distribution-agnostic model-agnostic meta-learning
Liam Collins, Aryan Mokhtari, and Sanjay Shakkottai · 2020
Later among the works it cites.
Few-shot learning via learning the representation, provably
Simon S Du, Wei Hu, Sham M Kakade, Jason D Lee, and Qi Lei · 2020
Later among the works it cites.
On the convergence theory of gradient-based model-agnostic meta-learning algorithms
Alireza Fallah, Aryan Mokhtari, and Asuman Ozdaglar · 2020
Later among the works it cites.
Provably convergent policy gradient methods for model-agnostic meta-reinforcement learning
Alireza Fallah, Aryan Mokhtari, and Asuman Ozdaglar · 2020
Later among the works it cites.
On the iteration complexity of hypergradient computation
Riccardo Grazzi, Luca Franceschi, Massimiliano Pontil, and Saverio Salzo · 2020
Later among the works it cites.
Robust stochastic bandit algorithms under probabilistic unbounded adversarial attack
Ziwei Guan, Kaiyi Ji, Donald J Bucci Jr, Timothy Y Hu, Joseph Palombo, Michael Liston, and Yingbin Liang · 2020
Later among the works it cites.
Milenas: Efficient neural architecture search via mixed-level reformulation
Chaoyang He, Haishan Ye, Li Shen, and Tong Zhang · 2020
Later among the works it cites.
Mingyi Hong, Hoi-To Wai, Zhaoran Wang, and Zhuoran Yang · 2020
Later among the works it cites.
Convergence of meta-learning with task-specific adaptation over partial parameters
Kaiyi Ji, Jason D Lee, Yingbin Liang, and H Vincent Poor · 2020
Later among the works it cites.
Learning latent features with pairwise penalties in low-rank matrix completion
Kaiyi Ji, Jian Tan, Jinfeng Xu, and Yuejie Chi · 2020
Later among the works it cites.
History-gradient aided batch size adaptation for variance reduced algorithms
Kaiyi Ji, Zhe Wang, Bowen Weng, Yi Zhou, Wei Zhang, and Yingbin Liang · 2020
Later among the works it cites.
Bilevel optimization: Nonasymptotic analysis and faster algorithms
Kaiyi Ji, Junjie Yang, and Yingbin Liang · 2020
Later among the works it cites.
Multi-step model-agnostic meta-learning: Convergence and improved algorithms
Kaiyi Ji, Junjie Yang, and Yingbin Liang · 2020
Later among the works it cites.
Multi-step estimation for gradient-based meta-learning
Jin-Hwa Kim, Junyoung Park, and Yongseok Choi · 2020
Later among the works it cites.
Improved bilevel model: Fast and optimal algorithm with theoretical guarantee
Junyi Li, Bin Gu, and Heng Huang · 2020
Later among the works it cites.
Near-optimal algorithms for minimax optimization
Tianyi Lin, Chi Jin, Michael Jordan, et al · 2020
Later among the works it cites.
A generic first-order algorithmic framework for bi-level programming beyond lower-level singleton
Risheng Liu, Pan Mu, Xiaoming Yuan, Shangzhi Zeng, and Jin Zhang · 2020
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Later among the works it cites.
Rapid learning or feature reuse? towards understanding the effectiveness of MAML
Aniruddh Raghu, Maithra Raghu, Samy Bengio, and Oriol Vinyals · 2020
Later among the works it cites.
A gradient-based bilevel optimization approach for tuning hyperparameters in machine learning
Ankur Sinha, Tanmay Khandait, and Raja Mohanty · 2020
Later among the works it cites.
ES-MAML: Simple hessian-free meta learning
Xingyou Song, Wenbo Gao, Yuxiang Yang, Choromanski Krzysztof, Aldo Pacchiano, and Yunhao Tang · 2020
Later among the works it cites.
Provable meta-learning of linear representations
Nilesh Tripuraneni, Chi Jin, and Michael I Jordan · 2020
Later among the works it cites.
Global convergence and induced kernels of gradient-based meta-learning with neural nets
Haoxiang Wang, Ruoyu Sun, and Bo Li · 2020
Later among the works it cites.
On the global optimality of model-agnostic meta-learning
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2020
Later among the works it cites.
Hyper-parameter optimization: A review of algorithms and applications
Tong Yu and Hong Zhu · 2020
Later among the works it cites.
Boosting one-point derivative-free online optimization via residual feedback
Yan Zhang, Yi Zhou, Kaiyi Ji, and Michael M Zavlanos · 2020
Later among the works it cites.
Improving the convergence rate of one-point zeroth-order optimization using residual feedback
Yan Zhang, Yi Zhou, Kaiyi Ji, and Michael M Zavlanos · 2020
Later among the works it cites.
Proximal gradient algorithm with momentum and flexible parameter restart for nonconvex optimization
Y Zhou, Z Wang, K Ji, Y Liang, and V Tarokh · 2020
Later among the works it cites.
A single-timescale stochastic bilevel optimization method
Tianyi Chen, Yuejiao Sun, and Wotao Yin · 2021
Closest in time.
On stochastic moving-average estimators for non-convex optimization
Zhishuai Guo, Yi Xu, Wotao Yin, Rong Jin, and Tianbao Yang · 2021
Closest in time.
Randomized stochastic variance-reduced methods for stochastic bilevel optimization
Zhishuai Guo and Tianbao Yang · 2021
Closest in time.
Lower bounds and accelerated algorithms for bilevel optimization
Kaiyi Ji and Yingbin Liang · 2021
Closest in time.
Understanding estimation and generalization error of generative adversarial networks
Kaiyi Ji, Yi Zhou, and Yingbin Liang · 2021
Closest in time.
A near-optimal algorithm for stochastic bilevel optimization via double-momentum
Prashant Khanduri, Siliang Zeng, Mingyi Hong, Hoi-To Wai, Zhaoran Wang, and Zhuoran Yang · 2021
Closest in time.
A value-function-based interior-point method for non-convex bi-level optimization
Risheng Liu, Xuan Liu, Xiaoming Yuan, Shangzhi Zeng, and Jin Zhang · 2021
Closest in time.
BOIL: Towards representation change for few-shot learning
Jaehoon Oh, Hyungjun Yoo, ChangHwan Kim, and Se-Young Yun · 2021
Closest in time.
When will gradient methods converge to max-margin classifier under relu models?
Tengyu Xu, Yi Zhou, Kaiyi Ji, and Yingbin Liang · 2021
Closest in time.
Provably faster algorithms for bilevel optimization
Junjie Yang, Kaiyi Ji, and Yingbin Liang · 2021
Closest in time.