Fetching the paper…
Reading the bibliography…
Deep Learning (DL) powered by Deep Neural Networks (DNNs) has revolutionized various domains, yet understanding the intricacies of DNN decision-making and learning processes remains a significant challenge.
1905
Earlier work this paper cites.
1907
Earlier work this paper cites.
V. N. Vapnik, “Adaptive and learning systems for signal processing communications, and control,” Statistical learning theory , 1998
1998
Earlier work this paper cites.
P. L. Bartlett and S. Mendelson, “Rademacher and gaussian complexities: Risk bounds and structural results,” in International Conference on Computational Learning Theory . Springer, 2001, pp. 224–240
2001
Earlier work this paper cites.
W. J. Reed, “The Pareto, Zipf and other power laws,” Economics Letters , vol. 74, no. 1, pp. 15–19, Dec. 2001
2001
Earlier work this paper cites.
S. Mukherjee, P. Niyogi, T. Poggio, and R. Rifkin, “Statistical learning: Stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization,” 2002
2002
Earlier work this paper cites.
O. Bousquet and A. Elisseeff, “Stability and generalization,” The Journal of Machine Learning Research , vol. 2, pp. 499–526, 2002
2002
Earlier work this paper cites.
T. Poggio, R. Rifkin, S. Mukherjee, and P. Niyogi, “General conditions for predictivity in learning theory,” Nature , vol. 428, no. 6981, pp. 419–422, 2004
2004
Earlier work this paper cites.
2012
Earlier work this paper cites.
C. Dwork, M. Hardt, T. Pitassi, O. Reingold, and R. Zemel, “Fairness through awareness,” in Proceedings of the 3rd innovations in theoretical computer science conference , 2012, pp. 214–226
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Advances in Neural Information Processing Systems , vol. 25. Curran Associates, Inc., 2012
2012
Earlier work this paper cites.
R. Zemel, Y. Wu, K. Swersky, T. Pitassi, and C. Dwork, “Learning fair representations,” in International conference on machine learning . PMLR, 2013, pp. 325–333
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in Advances in Neural Information Processing Systems , vol. 27. Curran Associates, Inc., 2014
2014
Earlier work this paper cites.
M. Feldman, S. A. Friedler, J. Moeller, C. Scheidegger, and S. Venkatasubramanian, “Certifying and removing disparate impact,” in proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining , 2015, pp. 259–268
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Hardt, E. Price, and N. Srebro, “Equality of opportunity in supervised learning,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
D. Arpit, S. Jastrzębski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, and S. Lacoste-Julien, “A Closer Look at Memorization in Deep Networks,” in Proceedings of the 34th International Conference on Machine Learning . PMLR, Jul. 2017, pp. 233–242
2017
Earlier work this paper cites.
D. Krueger*, N. Ballas*, S. Jastrzebski*, D. Arpit*, M. S. Kanwal, T. Maharaj, E. Bengio, A. Fischer, and A. Courville, “Deep Nets Don’t Learn via Memorization,” Feb. 2017
2017
Earlier work this paper cites.
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein, “SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability,” in Advances in Neural Information Processing Systems , vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems , vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
R. Shokri, M. Stronati, C. Song, and V. Shmatikov, “Membership Inference Attacks Against Machine Learning Models,” in 2017 IEEE Symposium on Security and Privacy (SP) , pp. 3–18
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 618–626
2017
Earlier work this paper cites.
R. C. Fong and A. Vedaldi, “Interpretable explanations of black boxes by meaningful perturbation,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 3429–3437
2017
Earlier work this paper cites.
F. Tramer, V. Atlidakis, R. Geambasu, D. Hsu, J.-P. Hubaux, M. Humbert, A. Juels, and H. Lin, “Fairtest: Discovering unwarranted associations in data-driven applications,” in 2017 IEEE European Symposium on Security and Privacy (EuroS&P) . IEEE, 2017, pp. 401–416
2017
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell, “Overcoming catastrophic forgetting in neural networks,” Proceedings of the National Academy of Sciences , vol. 114, no. 13, pp. 3521–3526, Mar. 2017
2017
Earlier work this paper cites.
M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in International conference on machine learning . PMLR, 2017, pp. 3319–3328
2017
Earlier work this paper cites.
X. Chen, C. Liu, B. Li, K. Lu, and D. Song, “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning,” Dec. 2017
2017
Earlier work this paper cites.
Y. Tian, K. Pei, S. Jana, and B. Ray, “Deeptest: Automated testing of deep-neural-network-driven autonomous cars,” in Proceedings of the 40th international conference on software engineering , 2018, pp. 303–314
2018
Earlier work this paper cites.
A. Achille, M. Rovere, and S. Soatto, “Critical learning periods in deep networks,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
A. Morcos, M. Raghu, and S. Bengio, “Insights on representational similarity in neural networks with canonical correlation,” in Advances in Neural Information Processing Systems , vol. 31. Curran Associates, Inc., 2018
2018
Earlier work this paper cites.
S. Chatterjee, “Learning and Memorization,” in Proceedings of the 35th International Conference on Machine Learning . PMLR, Jul. 2018, pp. 755–763
2018
Earlier work this paper cites.
S. Yeom, I. Giacomelli, M. Fredrikson, and S. Jha, “Privacy Risk in Machine Learning: Analyzing the Connection to Overfitting,” in 2018 IEEE 31st Computer Security Foundations Symposium (CSF) , Jul. 2018, pp. 268–282
2018
Earlier work this paper cites.
Y. Long, V. Bindschaedler, L. Wang, D. Bu, X. Wang, H. Tang, C. A. Gunter, and K. Chen, “Understanding Membership Inferences on Well-Generalized Learning Models,” Feb. 2018
2018
Earlier work this paper cites.
A. Sablayrolles, M. Douze, C. Schmid, and H. Jégou, “D\’ej\‘a Vu: An empirical evaluation of the memorization properties of ConvNets,” Sep. 2018
2018
Earlier work this paper cites.
L. Schmidt, S. Santurkar, D. Tsipras, K. Talwar, and A. Madry, “Adversarially Robust Generalization Requires More Data,” in Advances in Neural Information Processing Systems , vol. 31. Curran Associates, Inc., 2018
2018
Earlier work this paper cites.
R. Guidotti, A. Monreale, S. Ruggieri, F. Turini, F. Giannotti, and D. Pedreschi, “A survey of methods for explaining black box models,” ACM computing surveys (CSUR) , vol. 51, no. 5, pp. 1–42, 2018
2018
Earlier work this paper cites.
H. Ritter, A. Botev, and D. Barber, “Online Structured Laplace Approximations For Overcoming Catastrophic Forgetting,” May 2018
2018
Earlier work this paper cites.
Z. Li and D. Hoiem, “Learning without Forgetting,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, no. 12, pp. 2935–2947, Dec. 2018
2018
Earlier work this paper cites.
B. Han, Q. Yao, X. Yu, G. Niu, M. Xu, W. Hu, I. W. Tsang, and M. Sugiyama, “Co-teaching: Robust training of deep neural networks with extremely noisy labels,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , ser. NIPS’18. Red Hook, NY, USA: Curran Associates Inc., Dec. 2018, pp. 8536–8546
2018
Earlier work this paper cites.
V. Pondenkandath, M. Alberti, S. Puran, R. Ingold, and M. Liwicki, “Leveraging Random Label Memorization for Unsupervised Pre-Training,” Nov. 2018
2018
Earlier work this paper cites.
B. Han, Q. Yao, X. Yu, G. Niu, M. Xu, W. Hu, I. W. Tsang, and M. Sugiyama, “Co-teaching: Robust training of deep neural networks with extremely noisy labels,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , ser. NIPS’18. Red Hook, NY, USA: Curran Associates Inc., Dec. 2018, pp. 8536–8546
2018
Earlier work this paper cites.
L. Jiang, Z. Zhou, T. Leung, L.-J. Li, and L. Fei-Fei, “MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels,” in Proceedings of the 35th International Conference on Machine Learning . PMLR, Jul. 2018, pp. 2304–2313
2018
Cited alongside, same era.
N. Carlini, C. Liu, Ú. Erlingsson, J. Kos, and D. Song, “The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks,” 2019
2019
Cited alongside, same era.
J. Gu and V. Tresp, “Neural Network Memorization Dissection,” Nov. 2019
2019
Cited alongside, same era.
J. Frankle, D. J. Schwab, and A. S. Morcos, “The Early Phase of Neural Network Training,” in International Conference on Learning Representations , Sep. 2019
2019
Cited alongside, same era.
A. Ansuini, A. Laio, J. H. Macke, and D. Zoccolan, “Intrinsic dimension of data representations in deep neural networks,” in Advances in Neural Information Processing Systems , vol. 32. Curran Associates, Inc., 2019
D. Patel and P. S. Sastry, “Memorization in Deep Neural Networks: Does the Loss Function Matter?” in Advances in Knowledge Discovery and Data Mining , ser. Lecture Notes in Computer Science, K. Karlapalem, H. Cheng, N. Ramakrishnan, R. K. Agrawal, P. K. Reddy, J. Srivastava, and T. Chakraborty, Eds. Cham: Springer International Publishing, 2021, pp. 131–142
2021
Later among the works it cites.
E. Kharitonov, M. Baroni, and D. Hupkes, “How BPE Affects Memorization in Transformers,” Dec. 2021
2021
Later among the works it cites.
H. Liu, Y. Yang, and X. Wang, “Overcoming Catastrophic Forgetting in Graph Neural Networks,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 10, pp. 8653–8661, May 2021
2021
Later among the works it cites.
D. Yogatama, C. de Masson d’Autume, and L. Kong, “Adaptive Semiparametric Language Models,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 362–373, Apr. 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Y. Li, C. Wei, and T. Ma, “Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks,” in Advances in Neural Information Processing Systems , vol. 32. Curran Associates, Inc., 2019
2019
Cited alongside, same era.
M. Toneva, A. Sordoni, R. T. des Combes, A. Trischler, Y. Bengio, and G. J. Gordon, “An Empirical Study of Example Forgetting during Deep Neural Network Learning,” Nov. 2019
2019
Cited alongside, same era.
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis, “Generalization through Memorization: Nearest Neighbor Language Models,” in International Conference on Learning Representations , Sep. 2019
2019
Cited alongside, same era.
X. Yu, B. Han, J. Yao, G. Niu, I. Tsang, and M. Sugiyama, “How does disagreement help generalization against label corruption?” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 7164–7173. [Online]. Available: https://proceedings.mlr.press/v97/yu19b.html
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Cui, Z. Gu, D. Mahajan, L. van der Maaten, S. Belongie, and S.-N. Lim, “Measuring Dataset Granularity,” Dec. 2019
2019
Cited alongside, same era.
V. Feldman and C. Zhang, “What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation,” in Advances in Neural Information Processing Systems , vol. 33. Curran Associates, Inc., 2020, pp. 2881–2891
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Zhang, Z. Li, S. Yan, X. He, and J. Sun, “Distribution Alignment: A Unified Framework for Long-Tail Visual Recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 2361–2370
2021
Later among the works it cites.
A. Althnian, D. AlSaeed, H. Al-Baity, A. Samha, A. B. Dris, N. Alzakari, A. Abou Elwafa, and H. Kurdi, “Impact of Dataset Size on Classification Performance: An Empirical Evaluation in the Medical Domain,” Applied Sciences , vol. 11, no. 2, p. 796, Jan. 2021
2021
Later among the works it cites.
P. Dhariwal and A. Nichol, “Diffusion Models Beat GANs on Image Synthesis,” in Advances in Neural Information Processing Systems , vol. 34. Curran Associates, Inc., 2021, pp. 8780–8794
2021
Later among the works it cites.
K. Tirumala, A. Markosyan, L. Zettlemoyer, and A. Aghajanyan, “Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language Models,” Advances in Neural Information Processing Systems , vol. 35, pp. 38 274–38 290, Dec. 2022
2022
Later among the works it cites.
S. Anagnostidis, G. Bachmann, L. Noci, and T. Hofmann, “The Curious Case of Benign Memorization,” in The Eleventh International Conference on Learning Representations , Sep. 2022
2022
Later among the works it cites.
N. Carlini, M. Jagielski, C. Zhang, N. Papernot, A. Terzis, and F. Tramer, “The Privacy Onion Effect: Memorization is Relative,” Advances in Neural Information Processing Systems , vol. 35, pp. 13 263–13 276, Dec. 2022
2022
Later among the works it cites.
N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tramer, “Membership Inference Attacks From First Principles,” Apr. 2022
2022
Later among the works it cites.
Y. Dong, K. Xu, X. Yang, T. Pang, Z. Deng, H. Su, and J. Zhu, “Exploring Memorization in Adversarial Training,” Mar. 2022
2022
Later among the works it cites.
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini, “Deduplicating Training Data Makes Language Models Better,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 8424–8445
2022
Later among the works it cites.
C. Agarwal, D. D’souza, and S. Hooker, “Estimating Example Difficulty Using Variance of Gradients,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 368–10 378
2022
Later among the works it cites.
Y. Cao, Z. Chen, M. Belkin, and Q. Gu, “Benign Overfitting in Two-layer Convolutional Neural Networks,” Advances in Neural Information Processing Systems , vol. 35, pp. 25 237–25 250, Dec. 2022
2022
Later among the works it cites.
M. Jagielski, O. Thakkar, F. Tramer, D. Ippolito, K. Lee, N. Carlini, E. Wallace, S. Song, A. G. Thakurta, N. Papernot, and C. Zhang, “Measuring Forgetting of Memorized Training Examples,” in The Eleventh International Conference on Learning Representations , Sep. 2022
2022
Later among the works it cites.
P. Maini, S. Garg, Z. Lipton, and J. Z. Kolter, “Characterizing Datapoints via Second-Split Forgetting,” Advances in Neural Information Processing Systems , vol. 35, pp. 30 044–30 057, Dec. 2022
2022
Later among the works it cites.
Z. Zhou, J. Yao, Y.-F. Wang, B. Han, and Y. Zhang, “Contrastive Learning with Boosted Memorization,” in Proceedings of the 39th International Conference on Machine Learning . PMLR, Jun. 2022, pp. 27 367–27 377
2022
Later among the works it cites.
Y. Wu, M. N. Rabe, D. Hutchins, and C. Szegedy, “Memorizing Transformers,” Mar. 2022
2022
Later among the works it cites.
Y. Tay, V. Tran, M. Dehghani, J. Ni, D. Bahri, H. Mehta, Z. Qin, K. Hui, Z. Zhao, J. Gupta, T. Schuster, W. W. Cohen, and D. Metzler, “Transformer Memory as a Differentiable Search Index,” Advances in Neural Information Processing Systems , vol. 35, pp. 21 831–21 843, Dec. 2022
2022
Later among the works it cites.
K. Meng, D. Bau, A. Andonian, and Y. Belinkov, “Locating and editing factual associations in gpt,” Advances in Neural Information Processing Systems , vol. 35, pp. 17 359–17 372, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
E. Mitchell, C. Lin, A. Bosselut, C. D. Manning, and C. Finn, “Memory-based model editing at scale,” in International Conference on Machine Learning . PMLR, 2022, pp. 15 817–15 831
2022
Later among the works it cites.
2023
Later among the works it cites.
Y. Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang, L. Lu, X. Jia, Q. Liu, J. Dai, Y. Qiao, and H. Li, “Planning-oriented autonomous driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2023, pp. 17 853–17 862
2023
Later among the works it cites.
Y. Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang et al. , “Planning-oriented autonomous driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 17 853–17 862
2023
Later among the works it cites.
M. Li, B. Lin, Z. Chen, H. Lin, X. Liang, and X. Chang, “Dynamic graph enhanced contrastive learning for chest x-ray report generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 3334–3343
2023
Later among the works it cites.
N. Carlini, J. Hayes, M. Nasr, M. Jagielski, V. Sehwag, F. Tramèr, B. Balle, D. Ippolito, and E. Wallace, “Extracting Training Data from Diffusion Models,” in 32nd USENIX Security Symposium (USENIX Security 23) , 2023, pp. 5253–5270
2023
Later among the works it cites.
P. Maini, M. C. Mozer, H. Sedghi, Z. C. Lipton, J. Z. Kolter, and C. Zhang, “Can Neural Network Memorization Be Localized?” Jul. 2023
2023
Later among the works it cites.
D. A. Nguyen, R. Levie, J. Lienen, G. Kutyniok, and E. Hüllermeier, “Memorization-Dilation: Modeling Neural Collapse Under Label Noise,” Apr. 2023
2023
Later among the works it cites.
X. Li, Q. Li, Z. Hu, and X. Hu, “On the Privacy Effect of Data Enhancement via the Lens of Memorization,” Feb. 2023
2023
Later among the works it cites.
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang, “Quantifying Memorization Across Neural Language Models,” Mar. 2023
2023
Later among the works it cites.
H. Xu, X. Liu, W. Wang, Z. Liu, A. K. Jain, and J. Tang, “How does the Memorization of Neural Networks Impact Adversarial Robust Models?” in Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , ser. KDD ’23. New York, NY, USA: Association for Computing Machinery, Aug. 2023, pp. 2801–2812
2023
Later among the works it cites.
R. Dwivedi, D. Dave, H. Naik, S. Singhal, R. Omer, P. Patel, B. Qian, Z. Wen, T. Shah, G. Morgan et al. , “Explainable ai (xai): Core ideas, techniques, and solutions,” ACM Computing Surveys , vol. 55, no. 9, pp. 1–33, 2023
2023
Later among the works it cites.
J. Zhang, Y. Hong, and Q. Zhao, “Memorization Weights for Instance Reweighting in Adversarial Training,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 9, pp. 11 228–11 236, Jun. 2023
2023
Later among the works it cites.
A. S. Dhanjal and W. Singh, “A comprehensive survey on automatic speech recognition using neural networks,” Multimedia Tools and Applications , vol. 83, no. 8, pp. 23 367–23 412, 2024
2024
Closest in time.
2024
Closest in time.
P. Hase, M. Bansal, B. Kim, and A. Ghandeharioun, “Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
C. Shao and Y. Feng, “Overcoming Catastrophic Forgetting beyond Continual Learning: Balanced Training for Neural Machine Translation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 2023–2036
2036
Closest in time.