Fetching the paper…
Reading the bibliography…
Memorization in over-parameterized neural networks could severely hurt generalization in the presence of mislabeled examples.
Hochreiter, S., Schmidhuber, J.: SIMPLIFYING NEURAL NETS BY DISCOVERING FLAT MINIMA. In: Tesauro, G., Touretzky, D.S., Leen, T.K. (eds.) Advances in Neural Information Processing Systems 7, pp. 529–536. MIT Press (1995)
1995
Earlier work this paper cites.
Brodley, C.E., Friedl, M.A., Others: Identifying and eliminating mislabeled training instances. In: Proceedings of the National Conference on Artificial Intelligence. pp. 799–805 (1996)
1996
Earlier work this paper cites.
Cretu, G.F., Stavrou, A., Locasto, M.E., Stolfo, S.J., Keromytis, A.D.: Casting out demons: Sanitizing training data for anomaly sensors. In: Security and Privacy, 2008. SP 2008. IEEE Symposium on. pp. 81–95. IEEE (2008)
2008
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: ImageNet large scale visual recognition challenge. International journal of computer vision 115
2015
Earlier work this paper cites.
Xiao, T., Xia, T., Yang, Y., Huang, C., Wang, X.: Learning from massive noisy labeled data for image classification. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2691–2699 (2015)
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Krishna, R.A., Hata, K., Chen, S., Kravitz, J., Shamma, D.A., Fei-Fei, L., Bernstein, M.S.: Embracing error to enable rapid crowdsourcing. In: Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems. pp. 3167–3179. CHI ’16, ACM, New York, NY, USA (2016). https://doi.org/10.1145/2858036.2858115
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Shrivastava, A., Gupta, A., Girshick, R.: Training region-based object detectors with online hard example mining. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 761–769 (2016)
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
Zagoruyko, S., Komodakis, N.: Wide residual networks. arXiv preprint arXiv:1605.07146 (May 2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Veit, A., Alldrin, N., Chechik, G., Krasin, I., Gupta, A., Belongie, S.J.: Learning from noisy Large-Scale datasets with minimal supervision. In: CVPR. pp. 6575–6583 (2017)
2017
Later among the works it cites.
Zhang, H., Cisse, M., Dauphin, Y.N., Lopez-Paz, D.: mixup: Beyond empirical risk minimization (October 2017)
2017
Later among the works it cites.
Chaudhari, P., Soatto, S.: Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks. In: 2018 Information Theory and Applications Workshop (ITA). pp. 1–10. IEEE (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ghosh, A., Kumar, H., Sastry, P.S.: Robust loss functions under label noise for deep neural networks. In: AAAI. pp. 1919–1925 (2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Hoffer, E., Hubara, I., Soudry, D.: Train longer, generalize better: closing the generalization gap in large batch training of neural networks. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processing Systems 30, pp. 1731–1741. Curran Associates, Inc. (2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Mandt, S., Hoffman, M.D., Blei, D.M.: Stochastic gradient descent as approximate bayesian inference. The Journal of Machine Learning Research 18
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Guo, S., Huang, W., Zhang, H., Zhuang, C., Dong, D., Scott, M.R., Huang, D.: Curriculumnet: Weakly supervised learning from large-scale web images. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 135–150 (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Lee, K.H., He, X., Zhang, L., Yang, L.: Cleannet: Transfer learning for scalable image classifier training with label noise. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 5447–5456 (2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Zhang, Z., Sabuncu, M.: Generalized cross entropy loss for training deep neural networks with noisy labels. In: Advances in neural information processing systems. pp. 8778–8788 (2018)
2018
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
Yao, J., Wu, H., Zhang, Y., Tsang, I.W., Sun, J.: Safeguarded dynamic label regression for noisy supervision. AAAI (2019)
2019
Later among the works it cites.