Fetching the paper…
Reading the bibliography…
Recent studies suggest that deep learning models inductive bias towards favoring simpler features may be one of the sources of shortcut learning.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, R. T., Pavlick, E., and Linzen, T. (2019) · 1902
Earlier work this paper cites.
Understanding neural networks via feature visualization: A survey
Nguyen, A., Yosinski, J., and Clune, J. (2019) · 1904
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2019) · 1907
Earlier work this paper cites.
What do compressed deep neural networks forget?
Hooker, S., Courville, A., Clark, G., Dauphin, Y., and Frome, A. (2019) · 1911
Earlier work this paper cites.
A proposal for the dartmouth summer research project on artificial intelligence, august 31, 1955
McCarthy, J., Minsky, M. L., Rochester, N., and Shannon, C. E. (1956) · 1955
Earlier work this paper cites.
On tables of random numbers
Kolmogorov, A. N. (1963) · 1963
Earlier work this paper cites.
A formal theory of inductive inference. part i
Solomonoff, R. J. (1964) · 1964
Earlier work this paper cites.
Three approaches to the quantitative definition of information
Kolmogorov, A. N. (1965) · 1965
Earlier work this paper cites.
On the length of programs for computing finite binary sequences: statistical considerations
Chaitin, G. J. (1969) · 1969
Earlier work this paper cites.
Universal sequential search problems
Levin, L. A. (1973) · 1973
Earlier work this paper cites.
Algorithmic information theory
Chaitin, G. J. (1977) · 1977
Earlier work this paper cites.
Randomness conservation inequalities; information and independence in mathematical theories
Levin, L. A. (1984) · 1984
Earlier work this paper cites.
Applications of time-bounded kolmogorov complexity in complexity theory
Allender, E. (1992) · 1992
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Hochreiter, S. and Schmidhuber, J. (1994) · 1994
Earlier work this paper cites.
Discovering solutions with low kolmogorov complexity and high generalization capability
Schmidhuber, J. (1995) · 1995
Earlier work this paper cites.
Discovering neural nets with low kolmogorov complexity and high generalization capability
Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Learning the parts of objects by non-negative matrix factorization
Lee, D. D. and Seung, H. S. (1999) · 1999
Earlier work this paper cites.
Algorithmic information theory
Grünwald, P. D., Vitányi, P. M., et al. (2008) · 2008
Earlier work this paper cites.
An introduction to Kolmogorov complexity and its applications
Li, M., Vitányi, P., et al. (2008) · 2008
Earlier work this paper cites.
Algorithmic probability: Theory and applications
Solomonoff, R. J. (2009) · 2009
Earlier work this paper cites.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. (2020) · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. (2011) · 2011
Earlier work this paper cites.
Nonnegative matrix factorization: A comprehensive review
Wang, Y.-X. and Zhang, Y.-J. (2012) · 2012
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A. (2013) · 2013
Earlier work this paper cites.
Short lists for shortest descriptions in short time
Teutsch, J. (2014) · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R. (2014) · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W. (2015) · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Multifaceted feature visualization: Uncovering the different types of features learned by each neuron in deep neural networks
Nguyen, A., Yosinski, J., and Clune, J. (2016) · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A. (2017) · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Fong, R. C. and Vedaldi, A. (2017) · 2017
Earlier work this paper cites.
Interpretation of neural networks is fragile
Ghorbani, A., Abid, A., and Zou, J. (2017) · 2017
Earlier work this paper cites.
Superhuman accuracy on the snemi3d connectomics challenge
Lee, K., Zung, J., Li, P., Jain, V., and Seung, H. S. (2017) · 2017
Earlier work this paper cites.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L. (2017) · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M. (2017) · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q. (2017) · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B. (2018) · 2018
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Ancona, M., Ceolini, E., Öztireli, C., and Gross, M. (2018) · 2018
Cited alongside, same era.
Linear algebraic structure of word senses, with applications to polysemy
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A. (2018) · 2018
Cited alongside, same era.
Deep convolutional networks do not classify based on global object shape
Baker, N., Lu, H., Erlikhman, G., and Kellman, P. J. (2018) · 2018
Cited alongside, same era.
Short lists with short programs in short time
Bauwens, B., Makhlin, A., Vereshchagin, N., and Zimand, M. (2018) · 2018
Cited alongside, same era.
The description length of deep learning models
Blier, L. and Ollivier, Y. (2018) · 2018
The effectiveness of feature attribution methods and its correlation with automatic evaluation scores
Nguyen, G., Kim, D., and Nguyen, A. (2021) · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., and Dosovitskiy, A. (2021) · 2021
Later among the works it cites.
Counterfactual explanations can be manipulated
Slack, D., Hilgard, A., Lakkaraju, H., and Singh, S. (2021) · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Wightman, R., Touvron, H., and Jégou, H. (2021) · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.-H., Tay, F. E., Feng, J., and Yan, S. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks
Chattopadhay, A., Sarkar, A., Howlader, P., and Balasubramanian, V. N. (2018) · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al. (2018) · 2018
Cited alongside, same era.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Pérez, G. V., Camargo, C. Q., and Louis, A. A. (2018) · 2018
Cited alongside, same era.
Rise: Randomized input sampling for explanation of black-box models
Petsiuk, V., Das, A., and Saenko, K. (2018) · 2018
Cited alongside, same era.
Analyzing image segmentation for connectomics
Plaza, S. M. and Funke, J. (2018) · 2018
Cited alongside, same era.
On the learning dynamics of deep neural networks
Tachet, R., Pezeshki, M., Shabanian, S., Courville, A., and Bengio, Y. (2018) · 2018
Cited alongside, same era.
Beyer, L., Izmailov, P., Kolesnikov, A., Caron, M., Kornblith, S., Zhai, X., Minderer, M., Tschannen, M., Alabdulmohsin, I., and Pavetic, F. (2022) · 2022
Later among the works it cites.
Toy models of superposition
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C. (2022) · 2022
Later among the works it cites.
HIVE: Evaluating the human interpretability of visual explanations
Kim, S. S. Y., Meister, N., Ramaswamy, V. V., Fong, R., and Russakovsky, O. (2022) · 2022
Later among the works it cites.
Relaxing the kolmogorov structure function for realistic computational constraints
Lee, Y., Finn, C., and Ermon, S. (2022) · 2022
Later among the works it cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. (2022) · 2022
Later among the works it cites.
Theory and applications of probabilistic kolmogorov complexity
Lu, Z. and Oliveira, I. C. (2022) · 2022
Later among the works it cites.
Making sense of dependence: Efficient black-box explanations using dependence measure
Novello, P., Fel, T., and Vigouroux, D. (2022) · 2022
Later among the works it cites.
Kolmogorov complexity and algorithmic randomness
Shen, A., Uspensky, V. A., and Vereshchagin, N. (2022) · 2022
Later among the works it cites.
Salient imagenet: How to discover spurious features in deep learning?
Singla, S. and Feizi, S. (2022) · 2022
Later among the works it cites.
From attribution maps to human-understandable explanations through concept relevance propagation
Achtibat, R., Dreyer, M., Eisenbraun, I., Bosse, S., Wiegand, T., Samek, W., and Lapuschkin, S. (2023) · 2023
Later among the works it cites.
Dissecting neural network robustness proofs
Banerjee, D., Singh, A., and Singh, G. (2023) · 2023
Later among the works it cites.
Understanding deep neural networks through the lens of their non-linearity
Bouniot, Q., Redko, I., Mallasto, A., Laclau, C., Arndt, K., Struckmeier, O., Heinonen, M., Kyrki, V., and Kaski, S. (2023) · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C. (2023b) · 2023
Later among the works it cites.
Goldblum, M., Finzi, M., Rowan, K., and Wilson, A. G. (2023) · 2023
Later among the works it cites.
Concept discovery and dataset exploration with singular value decomposition
Graziani, M., Nguyen, A.-p., O’Mahony, L., Müller, H., and Andrearczyk, V. (2023) · 2023
Later among the works it cites.
On the foundations of shortcut learning
Hermann, K. L., Mobahi, H., Fel, T., and Mozer, M. C. (2023) · 2023
Later among the works it cites.
Neurosurgeon: A toolkit for subnetwork analysis
Lepori, M. A., Pavlick, E., and Serre, T. (2023) · 2023
Later among the works it cites.
Patches are all you need?
Trockman, A. and Kolter, J. Z. (2023) · 2023
Later among the works it cites.
Multi-dimensional concept discovery (mcd): A unifying framework with completeness guarantees
Vielhaben, J., Blücher, S., and Strodthoff, N. (2023) · 2023
Later among the works it cites.
Going beyond neural network feature similarity: The network feature complexity and its interpretation using category theory
Chen, Y., Zhou, Z., and Yan, J. (2024) · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J. (2024) · 2024
Closest in time.
Feature accentuation: Revealing’what’features respond to in natural images
Hamblin, C., Fel, T., Saha, S., Konkle, T., and Alvarez, G. (2024) · 2024
Closest in time.
Deep networks always grok and here is why
Humayun, A. I., Balestriero, R., and Baraniuk, R. (2024) · 2024
Closest in time.
Learned feature representations are biased by complexity, learning order, position, and more
Lampinen, A. K., Chan, S. C., and Hermann, K. (2024) · 2024
Closest in time.
Diffused redundancy in pre-trained representations
Nanda, V., Speicher, T., Dickerson, J., Gummadi, K., Feizi, S., and Weller, A. (2024) · 2024
Closest in time.
Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task
Okawa, M., Lubana, E. S., Dick, R., and Tanaka, H. (2024) · 2024
Closest in time.
Emergence of hidden capabilities: Exploring learning dynamics in concept space
Park, C. F., Okawa, M., Lee, A., Lubana, E. S., and Tanaka, H. (2024) · 2024
Closest in time.
Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders
Rajamanoharan, S., Lieberum, T., Sonnerat, N., Conmy, A., Varma, V., Kramár, J., and Nanda, N. (2024) · 2024
Closest in time.
Sharpness-aware minimization enhances feature quality via balanced learning
Springer, J. M., Nagarajan, V., and Raghunathan, A. (2024) · 2024
Closest in time.
Simplicity bias of transformers to learn low sensitivity functions
Vasudeva, B., Fu, D., Zhou, T., Kau, E., Huang, Y., and Sharan, V. (2024) · 2024
Closest in time.