Fetching the paper…
Reading the bibliography…
Identifying optimal values for a high-dimensional set of hyperparameters is a problem that has received growing attention given its importance to large-scale machine learning applications such as neural architecture search.
A note on screening regression equations
Freedman, D. A · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Some PAC-Bayesian theorems
McAllester, D. A · 1999
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
PAC-Bayesian stochastic model selection
McAllester, D. A · 2003
Earlier work this paper cites.
Statistical learning theory and stochastic optimization
Catoni, O · 2004
Earlier work this paper cites.
Tighter PAC-Bayes bounds
Ambroladze, A., Parrado-Hernández, E., and Shawe-Taylor, J. S · 2007
Earlier work this paper cites.
Derivations for linear algebra and optimization, 2007
Duchi, J · 2007
Earlier work this paper cites.
Optimal Transport
Villani, C · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Probabilistic relational reasoning for differential privacy
Barthe, G., Köpf, B., Olmedo, F., and Zanella Béguelin, S · 2012
Earlier work this paper cites.
Generic methods for optimization-based modeling
Domke, J · 2012
Earlier work this paper cites.
PAC-Bayes bounds with data dependent priors
Parrado-Hernández, E., Ambroladze, A., Shawe-Taylor, J., and Sun, S · 2012
Earlier work this paper cites.
Uncertainty quantification in MD simulations. Part II: Bayesian inference of force-field parameters
Rizzi, F., Najm, H. N., Debusschere, B. J., Sargsyan, K., Salloum, M., Adalsteinsson, H., and Knio, O. M · 2012
Earlier work this paper cites.
Lecture 6.5-RMSProp, Coursera: Neural networks for machine learning
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Tighter PAC-Bayes bounds through distribution-dependent priors
Lever, G., Laviolette, F., and Shawe-Taylor, J · 2013
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Dwork, C. and Roth, A · 2014
Cited alongside, same era.
Differential privacy and machine learning: a survey and review
Ji, Z., Lipton, Z. C., and Elkan, C · 2014
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Generalization in adaptive data analysis and holdout reuse
Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., and Roth, A · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Cited alongside, same era.
Privacy for free: Posterior sampling and stochastic gradient Monte Carlo
Bilevel programming for hyperparameter optimization and meta-learning
Franceschi, L., Frasconi, P., Salzo, S., Grazzi, R., and Pontil, M · 2018
Later among the works it cites.
Practical bounds on the error of Bayesian posterior approximations: a nonasymptotic approach
Huggins, J. H., Campbell, T., Kasprzak, M., and Broderick, T · 2018
Later among the works it cites.
Fixing weight decay regularization in Adam, 2018
Loshchilov, I. and Hutter, F · 2018
Later among the works it cites.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Later among the works it cites.
Global convergence of Langevin dynamics based algorithms for nonconvex optimization
Xu, P., Chen, J., Zou, D., and Gu, Q · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, Y.-X., Fienberg, S., and Smola, A · 2015
Cited alongside, same era.
Early stopping as nonparametric variational inference
Duvenaud, D., Maclaurin, D., and Adams, R · 2016
Cited alongside, same era.
Introduction to online convex optimization
Hazan, E · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Scalable gradient-based tuning of continuous regularization hyperparameters
Luketina, J., Berglund, M., Greff, K., and Raiko, T · 2016
Cited alongside, same era.
Max-information, differential privacy, and post-selection hypothesis testing
Rogers, R., Roth, A., Smith, A., and Thakkar, O · 2016
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M · 2017
Cited alongside, same era.
Grefenstette, E., Amos, B., Yarats, D., Htut, P. M., Molchanov, A., Meier, F., Kiela, D., Cho, K., and Chintala, S · 2019
Later among the works it cites.
Towards understanding generalization in gradient-based meta-learning
Guiroy, S., Verma, V., and Pal, C · 2019
Later among the works it cites.
On connecting stochastic gradient MCMC and differential privacy
Li, B., Chen, C., Liu, H., and Carin, L · 2019
Later among the works it cites.
Random search and reproducibility for neural architecture search
Li, L. and Talwalkar, A · 2019
Later among the works it cites.
DARTS: Differentiable architecture search
Liu, H., Simonyan, K., and Yang, Y · 2019
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Lorraine, J., Vicol, P., and Duvenaud, D · 2019
Later among the works it cites.
Understanding and correcting pathologies in the training of learned optimizers
Metz, L., Maheswaranathan, N., Nixon, J., Freeman, D., and Sohl-Dickstein, J · 2019
Later among the works it cites.
Information-theoretic generalization bounds for SGLD via data-dependent estimates
Negrea, J., Haghifam, M., Dziugaite, G. K., Khisti, A., and Roy, D. M · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Truncated back-propagation for bilevel optimization
Shaban, A., Cheng, C.-A., Hatch, N., and Boots, B · 2019
Later among the works it cites.
Can gradient clipping mitigate label noise?
Menon, A. K., Rawat, A. S., Reddi, S. J., and Kumar, S · 2020
Closest in time.
PAC-Bayes Analysis Beyond the Usual Bounds
Rivasplata, O., Kuzborskij, I., Szepesvari, C., and Shawe-Taylor, J · 2020
Closest in time.
Understanding and robustifying differentiable architecture search
Zela, A., Elsken, T., Saikia, T., Marrakchi, Y., Brox, T., and Hutter, F · 2020
Closest in time.