Fetching the paper…
Reading the bibliography…
We consider a class of a nested optimization problems involving inner and outer objectives.
Learning internal representations
Baxter, J. (1995) · 1995
Earlier work this paper cites.
Theoretical models of learning to learn
Baxter, J. (1998) · 1998
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L. (1998) · 1998
Earlier work this paper cites.
Algorithmic stability and meta-learning
Maurer, A. (2005) · 2005
Earlier work this paper cites.
An overview of bilevel optimization
Colson, B., Marcotte, P., and Savard, G. (2007) · 2007
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation
Griewank, A. and Walther, A. (2008) · 2008
Earlier work this paper cites.
Learning deep architectures for ai
Bengio, Y. et al. (2009) · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J. S., Bardenet, R., Bengio, Y., and Kégl, B. (2011) · 2011
Earlier work this paper cites.
Model selection for primal SVM
Moore, G., Bergeron, C., and Bennett, K. P. (2011) · 2011
Cited alongside, same era.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Cited alongside, same era.
Generic Methods for Optimization-Based Modeling
Domke, J. (2012) · 2012
Cited alongside, same era.
Making a Science of Model Search: Hyperparameter Optimization in Hundreds of Dimensions for Vision Architectures
Bergstra, J., Yamins, D., and Cox, D. D. (2013) · 2013
Cited alongside, same era.
Beyond Manual Tuning of Hyperparameters
Hutter, F., Lücke, J., and Schmidt-Thieme, L. (2015) · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R. P. (2015) · 2015
Cited alongside, same era.
Hyperparameter optimization with approximate gradient
Pedregosa, F. (2016) · 2016
Later among the works it cites.
Meta-learning with memory-augmented neural networks
Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T. (2016) · 2016
Later among the works it cites.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., and Wierstra, D. (2016) · 2016
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Closest in time.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M. (2017) · 2017
Closest in time.
Meta-Learning with Temporal Convolutions
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N. (2016) · 2016
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Cited alongside, same era.
The benefit of multitask representation learning
Maurer, A., Pontil, M., and Romera-Paredes, B. (2016) · 2016
Cited alongside, same era.
Closest in time.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H. (2017) · 2017
Closest in time.
Learned optimizers that scale and generalize
Wichrowska, O., Maheswaranathan, N., Hoffman, M. W., Colmenarejo, S. G., Denil, M., de Freitas, N., and Sohl-Dickstein, J. (2017) · 2017
Closest in time.