Fetching the paper…
Reading the bibliography…
We introduce a framework based on bilevel programming that unifies gradient-based hyperparameter optimization and meta-learning.
Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Networks
Schmidhuber, J. (1992) · 1992
Earlier work this paper cites.
Well-posed optimization problems
Dontchev, A. L. and Zolezzi, T. (1993) · 1993
Earlier work this paper cites.
Learning internal representations
Baxter, J. (1995) · 1995
Earlier work this paper cites.
Multitask learning
Caruana, R. (1998) · 1998
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L. (1998) · 1998
Earlier work this paper cites.
Learning multiple tasks with kernel methods
Evgeniou, T., Micchelli, C. A., and Pontil, M. (2005) · 2005
Earlier work this paper cites.
An overview of bilevel optimization
Colson, B., Marcotte, P., and Savard, G. (2007) · 2007
Earlier work this paper cites.
An efficient method for gradient-based adaptation of hyperparameters in svm models
Keerthi, S. S., Sindhwani, V., and Chapelle, O. (2007) · 2007
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation
Griewank, A. and Walther, A. (2008) · 2008
Earlier work this paper cites.
Classification model selection via bilevel programming
Kunapuli, G., Bennett, K., Hu, J., and Pang, J.-S. (2008) · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Earlier work this paper cites.
Generic Methods for Optimization-Based Modeling
Domke, J. (2012) · 2012
Earlier work this paper cites.
Undoing the damage of dataset bias
Khosla, A., Zhou, T., Malisiewicz, T., Efros, A. A., and Torralba, A. (2012) · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P. (2012) · 2012
Cited alongside, same era.
Practical bilevel optimization: algorithms and applications
Bard, J. F. (2013) · 2013
Cited alongside, same era.
Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures
Bergstra, J., Yamins, D., and Cox, D. D. (2013) · 2013
Cited alongside, same era.
From n to n+ 1: Multiclass transfer incremental learning
Kuzborskij, I., Orabona, F., and Caputo, B. (2013) · 2013
Cited alongside, same era.
Learning constrained task similarities in graph-regularized multi-task learning
Flamary, R., Rakotomamonjy, A., and Gasso, G. (2014) · 2014
Cited alongside, same era.
Beyond Manual Tuning of Hyperparameters
Hutter, F., Lücke, J., and Schmidt-Thieme, L. (2015) · 2015
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., and Wierstra, D. (2016) · 2016
Later among the works it cites.
Automatic differentiation in machine learning: a survey
Baydin, A. G., Pearlmutter, B. A., Radul, A. A., and Siskind, J. M. (2017) · 2017
Later among the works it cites.
On optimal generalizability in parametric learning
Beirami, A., Razaviyayn, M., Shahrampour, S., and Tarokh, V. (2017) · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Later among the works it cites.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M. (2017) · 2017
Later among the works it cites.
Learning to remember rare events
Kaiser, L., Nachum, O., Roy, A., and Bengio, S. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Siamese neural networks for one-shot image recognition
Koch, G., Zemel, R., and Salakhutdinov, R. (2015) · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B. (2015) · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D. K., and Adams, R. P. (2015) · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N. (2016) · 2016
Cited alongside, same era.
Edwards, H. and Storkey, A. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P. (2017) · 2017
Later among the works it cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J. (2017) · 2017
Later among the works it cites.
Meta networks
Munkhdalai, T. and Yu, H. (2017) · 2017
Later among the works it cites.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H. (2017) · 2017
Later among the works it cites.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R. S. (2017) · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
Learned optimizers that scale and generalize
Wichrowska, O., Maheswaranathan, N., Hoffman, M. W., Colmenarejo, S. G., Denil, M., Freitas, N., and Sohl-Dickstein, J. (2017) · 2017
Later among the works it cites.
Incremental learning-to-learn with statistical guarantees
Denevi, G., Ciliberto, C., Stamos, D., and Pontil, M. (2018) · 2018
Closest in time.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P. (2018) · 2018
Closest in time.