Fetching the paper…
Reading the bibliography…
In most machine learning training paradigms a fixed, often handcrafted, loss function is assumed to be a good proxy for an underlying evaluation metric.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Optimizing area under roc curve with SVMs
Rakotomamonjy, A · 2004
Earlier work this paper cites.
Labeled faces in the wild: A database for studying face recognition in unconstrained environments
Huang, G. B., Ramesh, M., Berg, T., and Learned-Miller, E · 2007
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2009
Earlier work this paper cites.
Self-paced learning for latent variable models
Kumar, M. P., Packer, B., and Koller, D · 2010
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Online and stochastic gradient methods for non-decomposable loss functions
Kar, P., Narasimhan, H., and Jain, P · 2014
Earlier work this paper cites.
Deep learning face representation by joint identification-verification
Sun, Y., Chen, Y., Wang, X., and Tang, X · 2014
Earlier work this paper cites.
Learning face representation from scratch
Yi, D., Lei, Z., Liao, S., and Li, S. Z · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R. P · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., and Philbin, J · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gómez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and de Freitas, N · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Large-margin softmax loss for convolutional neural networks
Liu, W., Wen, Y., Yu, Z., and Yang, M · 2016
Cited alongside, same era.
Deep metric learning via lifted structured feature embedding
Song, H. O., Xiang, Y., Jegelka, S., and Savarese, S · 2016
Cited alongside, same era.
A discriminative feature learning approach for deep face recognition
Wen, Y., Zhang, K., Li, Z., and Qiao, Y · 2016
Searching for activation functions
Ramachandran, P., Zoph, B., and Le, Q. V · 2017
Later among the works it cites.
Learned optimizers that scale and generalize
Wichrowska, O., Maheswaranathan, N., Hoffman, M. W., Colmenarejo, S. G., Denil, M., de Freitas, N., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Sampling matters in deep embedding learning
Wu, C.-Y., Manmatha, R., Smola, A. J., and Krähenbühl, P · 2017
Later among the works it cites.
Learning to teach
Fan, Y., Tian, F., Qin, T., Li, X.-Y., and Liu, T.-Y · 2018
Later among the works it cites.
Deep metric learning with hierarchical triplet loss
Ge, W., Huang, W., Dong, D., and Scott, M. R · 2018
Later among the works it cites.
Deep bilevel learning
Jenni, S. and Favaro, P · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Neural optimizer search with reinforcement learning
Bello, I., Zoph, B., Vasudevan, V., and Le, Q. V · 2017
Cited alongside, same era.
Scalable Learning of Non-Decomposable Objectives
Eban, E., Schain, M., Mackey, A., Gordon, A., Rifkin, R., and Elidan, G · 2017
Cited alongside, same era.
Forward and reverse gradient-based hyperparameter optimization
Franceschi, L., Donini, M., Frasconi, P., and Pontil, M · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., van der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Jiang, L., Zhou, Z., Leung, T., Li, L., and Fei-Fei, L · 2018
Later among the works it cites.
Attention-based ensemble for deep metric learning
Kim, W., Goyal, B., Chawla, K., Lee, J., and Kwon, K · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Later among the works it cites.
Learning to reweight examples for robust deep learning
Ren, M., Zeng, W., Yang, B., and Urtasun, R · 2018
Later among the works it cites.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A · 2018
Later among the works it cites.
CosFace: Large margin Cosine loss for deep face recognition
Wang, H., Wang, Y., Zhou, Z., Ji, X., Li, Z., Gong, D., Zhou, J., and Liu, W · 2018
Later among the works it cites.
Learning to teach with dynamic loss functions
Wu, L., Tian, F., Xia, Y., Fan, Y., Qin, T., Jian-Huang, L., and Liu, T.-Y · 2018
Later among the works it cites.
SNAS: stochastic neural architecture search
Xie, S., Zheng, H., Liu, C., and Lin, L · 2019
Closest in time.
Autoloss: Learning discrete schedule for alternate optimization
Xu, H., Zhang, H., Hu, Z., Liang, X., Salakhutdinov, R., and Xing, E · 2019
Closest in time.