Fetching the paper…
Reading the bibliography…
Bilevel optimization (BLO) is a popular approach with many applications including hyperparameter optimization, neural architecture search, adversarial robustness and model-agnostic meta-learning.
Automatic Hessians by reverse accumulation
Bruce Christianson · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Tutorial on training recurrent neural networks, covering BPPT, RTRL, EKF and the echo state network approach
Herbert Jaeger · 2002
Earlier work this paper cites.
Probability and statistics: The science of uncertainty
Michael J Evans and Jeffrey S Rosenthal · 2004
Earlier work this paper cites.
Enabling user-driven checkpointing strategies in reverse-mode automatic differentiation
Laurent Hascoet and Mauricio Araya-Polo · 2006
Earlier work this paper cites.
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation, Second Edition
A. Griewank and A. Walther · 2008
Earlier work this paper cites.
One shot learning of simple visual concepts
Brenden Lake, Ruslan Salakhutdinov, Jason Gross, and Joshua Tenenbaum · 2011
Earlier work this paper cites.
Iterative methods for computing eigenvalues and eigenvectors
Maysum Panju · 2011
Earlier work this paper cites.
Estimating the Hessian by back-propagating curvature
James Martens, Ilya Sutskever, and Kevin Swersky · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Tobias Glasmachers, Yi Sun, Jan Peters, and Jürgen Schmidhuber · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Stephen Gould, Basura Fernando, Anoop Cherian, Peter Anderson, Rodrigo Santa Cruz, and Edison Guo · 2016
Earlier work this paper cites.
Muprop: Unbiased backpropagation for stochastic neural networks
Shixiang Gu, Sergey Levine, Ilya Sutskever, and Andriy Mnih · 2016
Earlier work this paper cites.
Categorical reparameterization with Gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient
Fabian Pedregosa · 2016
Cited alongside, same era.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, koray kavukcuoglu, and Daan Wierstra · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Cited alongside, same era.
Concrete dropout
Yarin Gal, Jiri Hron, and Alex Kendall · 2017
Cited alongside, same era.
The reversible residual network: Backpropagation without storing activations
Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse · 2017
Cited alongside, same era.
On the convergence theory of gradient-based model-agnostic meta-learning algorithms
Alireza Fallah, Aryan Mokhtari, and Asuman E. Ozdaglar · 2019
Later among the works it cites.
Online meta-learning
Chelsea Finn, Aravind Rajeswaran, Sham M. Kakade, and Sergey Levine · 2019
Later among the works it cites.
Generalized inner loop meta-learning
Edward Grefenstette, Brandon Amos, Denis Yarats, Phu Mon Htut, Artem Molchanov, Franziska Meier, Douwe Kiela, Kyunghyun Cho, and Soumith Chintala · 2019
Later among the works it cites.
Introduction to online convex optimization
Elad Hazan · 2019
Later among the works it cites.
DARTS: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Unbiasing truncated backpropagation through time
Corentin Tallec and Yann Ollivier · 2017
Cited alongside, same era.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
George Tucker, Andriy Mnih, Chris J Maddison, John Lawson, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Neural ordinary differential equations
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Cited alongside, same era.
Learning to defense by learning to attack
Zhehui Chen, Haoming Jiang, Bo Dai, and Tuo Zhao · 2018
Cited alongside, same era.
Graph normalizing flows
Jenny Liu, Aviral Kumar, Jimmy Ba, Jamie Kiros, and Kevin Swersky · 2019
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2019
Later among the works it cites.
Meta-learning with implicit gradients
Aravind Rajeswaran, Chelsea Finn, Sham M Kakade, and Sergey Levine · 2019
Later among the works it cites.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Adversarial attacks on graph neural networks via meta learning
Daniel Zügner and Stephan Günnemann · 2019
Later among the works it cites.
Multi-step model-agnostic meta-learning: Convergence and improved algorithms
Kaiyi Ji, Junjie Yang, and Yingbin Liang · 2020
Closest in time.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Closest in time.
Normalizing flows: An introduction and review of current methods
I. Kobyzev, S. Prince, and M. Brubaker · 2020
Closest in time.
ES-MAML: Simple hessian-free meta learning
Xingyou Song, Wenbo Gao, Yuxiang Yang, Krzysztof Choromanski, Aldo Pacchiano, and Yunhao Tang · 2020
Closest in time.
Drawing early-bird tickets: Toward more efficient training of deep networks
Haoran You, Chaojian Li, Pengfei Xu, Yonggan Fu, Yue Wang, Xiaohan Chen, Richard G. Baraniuk, Zhangyang Wang, and Yingyan Lin · 2020
Closest in time.