Deep fried convnets
Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alex Smola, Le Song, and Ziyu Wang · 2015
Later among the works it cites.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Later among the works it cites.
Hypernetworks
Original
David Ha, Andrew Dai, and Quoc V Le · 2016
Later among the works it cites.
Non-stochastic best arm identification and hyperparameter optimization
Kevin Jamieson and Ameet Talwalkar · 2016
Later among the works it cites.
Scalable gradient-based tuning of continuous regularization hyperparameters
Jelena Luketina, Mathias Berglund, Klaus Greff, and Tapani Raiko · 2016
Later among the works it cites.
Unrolled generative adversarial networks
Original
Luke Metz, Ben Poole, David Pfau, and Jascha Sohl-Dickstein · 2016
Later among the works it cites.
Hyperparameter optimization with approximate gradient
Fabian Pedregosa · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Original
Barret Zoph and Quoc V Le · 2016
Later among the works it cites.
SMASH: One-shot model architecture search through hypernetworks
Original
Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston · 2017
Later among the works it cites.
Improved regularization of convolutional neural networks with cutout
Original
Terrance DeVries and Graham W Taylor · 2017
Later among the works it cites.
Beta-VAE: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Later among the works it cites.
Population-based training of neural networks
Original
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Later among the works it cites.
Learning curve prediction with Bayesian neural networks
Aaron Klein, Stefan Falkner, Jost Tobias Springenberg, and Frank Hutter · 2017
Later among the works it cites.
Hyperband: Bandit-based configuration evaluation for hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2017
Later among the works it cites.
The numerics of GANs
Lars Mescheder, Sebastian Nowozin, and Andreas Geiger · 2017
Later among the works it cites.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Bilevel programming for hyperparameter optimization and meta-learning
Original
Luca Franceschi, Paolo Frasconi, Saverio Salzo, and Massimilano Pontil · 2018
Later among the works it cites.
Fast and scalable Bayesian deep learning by weight-perturbation in Adam
Original
Mohammad Emtiyaz Khan, Didrik Nielsen, Voot Tangkaratt, Wu Lin, Yarin Gal, and Akash Srivastava · 2018
Later among the works it cites.
Stochastic hyperparameter optimization through hypernetworks
Original
Jonathan Lorraine and David Duvenaud · 2018
Later among the works it cites.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Later among the works it cites.
Flipout: Efficient pseudo-independent weight perturbations on mini-batches
Yeming Wen, Paul Vicol, Jimmy Ba, Dustin Tran, and Roger Grosse · 2018
Later among the works it cites.