Fetching the paper…
Reading the bibliography…
We propose an evolution strategies-based algorithm for estimating gradients in unrolled computation graphs, called ES-Single.
Evolutionsstrategie: Optimierung technischer Systeme nach Prinzipien der biologischen Evolution
Rechenberg, I · 1973
Earlier work this paper cites.
Evolutionsstrategien für die numerische Optimierung
Schwefel, H.-P. and Schwefel, H.-P · 1977
Earlier work this paper cites.
Learning internal representations by error propagation
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1985
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Backpropagation through time: What it does and how to do it
Werbos, P. J · 1990
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
Williams, R. J. and Peng, J · 1990
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B · 1993
Earlier work this paper cites.
Sensitivity analysis of the MM5 weather model using automatic differentiation
Bischof, C. H., Pusch, G. D., and Knoesel, R · 1996
Earlier work this paper cites.
Design and regularization of neural networks: The optimal use of a validation set
Larsen, J., Hansen, L. K., Svarer, C., and Ohlsson, M · 1996
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Bengio, Y · 2000
Earlier work this paper cites.
Sensitivity analysis of the climate of a chaotic system
Lea, D. J., Allen, M. R., and Haine, T. W · 2000
Earlier work this paper cites.
An adjoint method for the assimilation of statistical characteristics into eddy-resolving ocean models
Köhl, A. and Willebrand, J · 2002
Earlier work this paper cites.
Using a thousand optimization tasks to learn hyperparameter search strategies
Metz, L., Maheswaranathan, N., Sun, R., Freeman, C. D., Poole, B., and Sohl-Dickstein, J · 2002
Earlier work this paper cites.
UCI Machine Learning Repository, 2007
Asuncion, A. and Newman, D · 2007
Earlier work this paper cites.
Efficient multiple hyperparameter learning for log-linear models
Foo, C.-S., Do, C. B., and Ng, A. Y · 2008
Earlier work this paper cites.
Metz, L., Maheswaranathan, N., Freeman, C. D., Poole, B., and Sohl-Dickstein, J · 2009
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J. S., Bardenet, R., Bengio, Y., and Kégl, B · 2011
Earlier work this paper cites.
Generic methods for optimization-based modeling
Domke, J · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Staines, J. and Barber, D · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Monte Carlo Theory, Methods and Examples
Owen, A. B · 2013
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Freeze-thaw Bayesian optimization
Swersky, K., Snoek, J., and Adams, R. P · 2014
Earlier work this paper cites.
Natural evolution strategies
Wierstra, D., Schaul, T., Glasmachers, T., Sun, Y., Peters, J., and Schmidhuber, J · 2014
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Cited alongside, same era.
Gradient estimation using stochastic computation graphs
Schulman, J., Heess, N., Weber, T., and Abbeel, P · 2015
Cited alongside, same era.
Scalable Bayesian optimization using deep neural networks
Snoek, J., Rippel, O., Swersky, K., Kiros, R., Satish, N., Sundaram, N., Patwary, M., Prabhat, M., and Adams, R · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and De Freitas, N · 2016
Cited alongside, same era.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Cited alongside, same era.
Regularizing and optimizing LSTM language models
Merity, S., Keskar, N. S., and Socher, R · 2018
Later among the works it cites.
Meta-learning update rules for unsupervised representation learning
Metz, L., Maheswaranathan, N., Cheung, B., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Approximating real-time recurrent learning with random Kronecker factors
Mujika, A., Meier, F., and Steger, A · 2018
Later among the works it cites.
PIPPS: Flexible model-based policy search robust to the curse of chaos
Parmas, P., Rasmussen, C. E., Peters, J., and Doya, K · 2018
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
Wu, Y., Ren, M., Liao, R., and Grosse, R · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jamieson, K. and Talwalkar, A · 2016
Cited alongside, same era.
Li, K. and Malik, J · 2016
Cited alongside, same era.
Scalable gradient-based tuning of continuous regularization hyperparameters
Luketina, J., Berglund, M., Greff, K., and Raiko, T · 2016
Cited alongside, same era.
Hyperparameter optimization with approximate gradient
Pedregosa, F · 2016
Cited alongside, same era.
The generalized reparameterization gradient
Ruiz, F. R., AUEB, T. R., Blei, D., et al · 2016
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent
Baydin, A. G., Cornish, R., Rubio, D. M., Schmidt, M., and Wood, F · 2017
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M · 2017
Cited alongside, same era.
Optimal Kronecker-sum approximation of real time recurrent learning
Benzing, F., Gauy, M. M., Mujika, A., Martinsson, A., and Steger, A · 2019
Later among the works it cites.
Self-Tuning Networks: Bilevel optimization of hyperparameters using structured best-response functions
MacKay, M., Vicol, P., Lorraine, J., Duvenaud, D., and Grosse, R · 2019
Later among the works it cites.
Guided evolutionary strategies: Augmenting random search with surrogate gradients
Maheswaranathan, N., Metz, L., Tucker, G., Choi, D., and Sohl-Dickstein, J · 2019
Later among the works it cites.
Understanding and correcting pathologies in the training of learned optimizers
Metz, L., Maheswaranathan, N., Nixon, J., Freeman, D., and Sohl-Dickstein, J · 2019
Later among the works it cites.
Parmas, P. and Sugiyama, M · 2019
Later among the works it cites.
Truncated back-propagation for bilevel optimization
Shaban, A., Cheng, C.-A., Hatch, N., and Boots, B · 2019
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Lorraine, J., Vicol, P., and Duvenaud, D · 2020
Later among the works it cites.
Non-greedy gradient-based hyperparameter optimization over long horizons
Micaelli, P. and Storkey, A · 2020
Later among the works it cites.
Efficient and modular implicit differentiation
Blondel, M., Berthet, Q., Cuturi, M., Frostig, R., Hoyer, S., Llinares-López, F., Pedregosa, F., and Vert, J.-P · 2021
Later among the works it cites.
Machine learning–accelerated computational fluid dynamics
Kochkov, D., Smith, J. A., Alieva, A., Wang, Q., Brenner, M. P., and Hoyer, S · 2021
Later among the works it cites.
Optimized finite-build stellarator coils using automatic differentiation
McGreivy, N., Hudson, S. R., and Zhu, C · 2021
Later among the works it cites.
Practical real time recurrent learning with a sparse approximation
Menick, J., Elsen, E., Evci, U., Osindero, S., Simonyan, K., and Graves, A · 2021
Later among the works it cites.
Gradients are not all you need
Metz, L., Freeman, C. D., Schoenholz, S. S., and Kachman, T · 2021
Later among the works it cites.
Learning by directional gradient descent
Silver, D., Goyal, A., Danihelka, I., Hessel, M., and van Hasselt, H · 2021
Later among the works it cites.
Unbiased gradient estimation in unrolled computation graphs with persistent evolution strategies
Vicol, P., Metz, L., and Sohl-Dickstein, J · 2021
Later among the works it cites.
Gradient descent: The ultimate optimizer
Chandra, K., Xie, A., Ragan-Kelley, J., and Meijer, E · 2022
Later among the works it cites.
On Implicit Bias in Overparameterized Bilevel Optimization
Vicol, P., Lorraine, J., Pedregosa, F., Duvenaud, D., and Grosse, R · 2022
Later among the works it cites.
Noise reuse in online evolution strategies
Li, O., Harrison, J., Sohl-Dickstein, J., and Metz, L · 2023
Closest in time.
On Bilevel Optimization without Full Unrolls: Methods and Applications
Vicol, P. A · 2023
Closest in time.