Fetching the paper…
Reading the bibliography…
Unrolled computation graphs arise in many scenarios, including training RNNs, tuning hyperparameters through unrolled optimization, and training learned optimizers.
Evolutionsstrategie: Optimierung technischer Systeme nach Prinzipien der biologischen Evolution
Rechenberg, I · 1973
Earlier work this paper cites.
Learning internal representations by error propagation
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1985
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Backpropagation through time: What it does and how to do it
Werbos, P. J · 1990
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
Williams, R. J. and Peng, J · 1990
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Marcus, M., Santorini, B., and Marcinkiewicz, M. A · 1993
Earlier work this paper cites.
An investigation of the gradient descent process in neural networks
Pearlmutter, B · 1996
Earlier work this paper cites.
Using a thousand optimization tasks to learn hyperparameter search strategies
Metz, L., Maheswaranathan, N., Sun, R., Freeman, C. D., Poole, B., and Sohl-Dickstein, J · 2002
Earlier work this paper cites.
The data-flow equations of checkpointing in reverse automatic differentiation
Dauvergne, B. and Hascoët, L · 2006
Earlier work this paper cites.
UCI machine learning repository, 2007
Asuncion, A. and Newman, D · 2007
Earlier work this paper cites.
Metz, L., Maheswaranathan, N., Freeman, C. D., Poole, B., and Sohl-Dickstein, J · 2009
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
Generic methods for optimization-based modeling
Domke, J · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Staines, J. and Barber, D · 2012
Earlier work this paper cites.
Lecture 6.5—RMSprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Monte Carlo Theory, Methods and Examples
Owen, A. B · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
Freeze-thaw Bayesian optimization
Swersky, K., Snoek, J., and Adams, R. P · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Earlier work this paper cites.
Gradient estimation using stochastic computation graphs
Schulman, J., Heess, N., Weber, T., and Abbeel, P · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and De Freitas, N · 2016
Cited alongside, same era.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Cited alongside, same era.
The CMA evolution strategy: A tutorial
Hansen, N · 2016
Cited alongside, same era.
Non-stochastic best arm identification and hyperparameter optimization
Jamieson, K. and Talwalkar, A · 2016
Cited alongside, same era.
Li, K. and Malik, J · 2016
Cited alongside, same era.
Guided evolutionary strategies: Augmenting random search with surrogate gradients
Maheswaranathan, N., Metz, L., Tucker, G., Choi, D., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Simple random search provides a competitive approach to reinforcement learning
Mania, H., Guy, A., and Recht, B · 2018
Later among the works it cites.
Meta-learning update rules for unsupervised representation learning
Metz, L., Maheswaranathan, N., Cheung, B., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Approximating real-time recurrent learning with random Kronecker factors
Mujika, A., Meier, F., and Steger, A · 2018
Later among the works it cites.
PIPPS: Flexible model-based policy search robust to the curse of chaos
Parmas, P., Rasmussen, C. E., Peters, J., and Doya, K · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Baydin, A. G., Cornish, R., Rubio, D. M., Schmidt, M., and Wood, F · 2017
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M · 2017
Cited alongside, same era.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Learning to optimize neural nets
Li, K. and Malik, J · 2017
Cited alongside, same era.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A · 2017
Cited alongside, same era.
Random gradient-free minimization of convex functions
Nesterov, Y. and Spokoiny, V · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Cited alongside, same era.
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
Wu, Y., Ren, M., Liao, R., and Grosse, R · 2018
Later among the works it cites.
Adaptively truncating backpropagation through time to control gradient bias
Aicher, C., Foti, N. J., and Fox, E. B · 2019
Later among the works it cites.
Efficient optimization of loops and limits with randomized telescoping sums
Beatson, A. and Adams, R. P · 2019
Later among the works it cites.
Optimal Kronecker-sum approximation of real time recurrent learning
Benzing, F., Gauy, M. M., Mujika, A., Martinsson, A., and Steger, A · 2019
Later among the works it cites.
On the variance of unbiased online recurrent optimization
Cooijmans, T. and Martens, J · 2019
Later among the works it cites.
Scheduling the learning rate via hypergradients: New insights and a new algorithm
Donini, M., Franceschi, L., Pontil, M., Majumder, O., and Frasconi, P · 2019
Later among the works it cites.
Generalized inner loop meta-learning
Grefenstette, E., Amos, B., Yarats, D., Htut, P. M., Molchanov, A., Meier, F., Kiela, D., Cho, K., and Chintala, S · 2019
Later among the works it cites.
MacKay, M., Vicol, P., Lorraine, J., Duvenaud, D., and Grosse, R · 2019
Later among the works it cites.
A unified framework of online learning algorithms for training recurrent neural networks
Marschall, O., Cho, K., and Savin, C · 2019
Later among the works it cites.
Understanding and correcting pathologies in the training of learned optimizers
Metz, L., Maheswaranathan, N., Nixon, J., Freeman, D., and Sohl-Dickstein, J · 2019
Later among the works it cites.
Truncated back-propagation for bilevel optimization
Shaban, A., Cheng, C.-A., Hatch, N., and Boots, B · 2019
Later among the works it cites.
Neuroevolution for deep reinforcement learning problems
Ha, D · 2020
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Lorraine, J., Vicol, P., and Duvenaud, D · 2020
Later among the works it cites.
A practical sparse approximation for real time recurrent learning
Menick, J., Elsen, E., Evci, U., Osindero, S., Simonyan, K., and Graves, A · 2020
Later among the works it cites.
Non-greedy gradient-based hyperparameter optimization over long horizons
Micaelli, P. and Storkey, A · 2020
Later among the works it cites.
Variance reduction for evolution strategies via structured control variates
Tang, Y., Choromanski, K., and Kucukelbir, A · 2020
Later among the works it cites.