Fetching the paper…
Reading the bibliography…
We study neural network loss landscapes through the lens of mode connectivity, the observation that minimizers of neural networks retrieved via training on a dataset are connected via simple paths of low loss.
On the algebraic structure of feedforward network weight spaces
Hecht-Nielsen, R · 1990
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2016
Earlier work this paper cites.
Unsupervised feature extraction by time-contrastive learning and nonlinear ica
Hyvarinen, A. and Morioka, H · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Earlier work this paper cites.
Deep coral: Correlation alignment for deep domain adaptation
Sun, B. and Saenko, K · 2016
Earlier work this paper cites.
Bridging mode connectivity in loss landscapes and adversarial robustness
Zhao, P., Chen, P.-Y., Das, P., Ramamurthy, K. N., and Lin, X · 2016
Earlier work this paper cites.
Balestriero, R · 2017
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2017
Earlier work this paper cites.
Nonlinear ICA of temporally dependent stationary sources
Hyvarinen, A. and Morioka, H · 2017
Earlier work this paper cites.
The loss surface of deep and wide neural networks
Nguyen, Q. and Hein, M · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Peters, J., Janzing, D., and Schölkopf, B · 2017
Earlier work this paper cites.
Cognitive psychology for deep neural networks: A shape bias case study
Ritter, S., Barrett, D. G., Santoro, A., and Botvinick, M. M · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Earlier work this paper cites.
Mad max: Affine spline insights into deep learning
Balestriero, R. and Baraniuk, R · 2018
Earlier work this paper cites.
A spline theory of deep learning
Balestriero, R. et al · 2018
Earlier work this paper cites.
Recognition in terra incognita
Beery, S., Van Horn, G., and Perona, P · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Earlier work this paper cites.
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Earlier work this paper cites.
Excessive invariance causes adversarial vulnerability
Jacobsen, J.-H., Behrmann, J., Zemel, R., and Bethge, M · 2018
Earlier work this paper cites.
Domain generalization via conditional invariant representations
Li, Y., Gong, M., Tian, X., Liu, T., and Tao, D · 2018
Earlier work this paper cites.
Optimization landscape and expressivity of deep CNNs
Nguyen, Q. and Hein, M · 2018
Earlier work this paper cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Nguyen, Q., Mukkamala, M. C., and Hein, M · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., Combes, R. T. d., Trischler, A., Bengio, Y., and Gordon, G. J · 2018
Earlier work this paper cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Valle-Perez, G., Camargo, C. Q., and Louis, A. A · 2018
Earlier work this paper cites.
A max-affine spline perspective of recurrent neural networks
Wang, Z., Balestriero, R., and Baraniuk, R · 2018
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 2019
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Earlier work this paper cites.
Rethinking imagenet pre-training
He, K., Girshick, R., and Dollár, P · 2019
Earlier work this paper cites.
What do compressed deep neural networks forget?
Hooker, S., Courville, A., Clark, G., Dauphin, Y., and Frome, A · 2019
Earlier work this paper cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Kuditipudi, R., Wang, X., Lee, H., Zhang, Y., Li, Z., Hu, W., Ge, R., and Arora, S · 2019
Earlier work this paper cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Locatello, F., Bauer, S., Lucic, M., Raetsch, G., Gelly, S., Schölkopf, B., and Bachem, O · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 2019
Cited alongside, same era.
Do deep neural networks learn shallow learnable examples first?
Mangalam, K. and Prabhu, V. U · 2019
Cited alongside, same era.
Model similarity mitigates test set overuse
Mania, H., Miller, J., Schmidt, L., Hardt, M., and Recht, B · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, R. T., Pavlick, E., and Linzen, T · 2019
Cited alongside, same era.
Convergence of gradient descent on separable data
Nacson, M. S., Lee, J., Gunasekar, S., Savarese, P. H. P., Srebro, N., and Soudry, D · 2019
Cited alongside, same era.
Shape or texture: Understanding discriminative features in cnns
Islam, M. A., Kowal, M., Esser, P., Jia, S., Ommer, B., Derpanis, K. G., and Bruce, N · 2021
Later among the works it cites.
Causal autoregressive flows
Khemakhem, I., Monti, R., Leech, R., and Hyvarinen, A · 2021
Later among the works it cites.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics
Kunin, D., Sagastuy-Brena, J., Ganguli, S., Yamins, D. L., and Tanaka, H · 2021
Later among the works it cites.
Predicting inductive biases of pre-trained models
Lovering, C., Jha, R., Linzen, T., and Pavlick, E · 2021
Later among the works it cites.
How do quadratic regularizers prevent catastrophic forgetting: The role of interpolation
Lubana, E. S., Trivedi, P., Koutra, D., and Dick, R. P · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nakkiran, P., Kalimeris, D., Kaplun, G., Edelman, B., Yang, T., Barak, B., and Zhang, H · 2019
Cited alongside, same era.
On connected sublevel sets in deep learning
Nguyen, Q · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., and Courville, A · 2019
Cited alongside, same era.
Are disentangled representations helpful for abstract visual reasoning?
Van Steenkiste, S., Locatello, F., Schmidhuber, J., and Bachem, O · 2019
Cited alongside, same era.
On the transfer of disentangled representations in realistic settings
Dittadi, A., Träuble, F., Locatello, F., Wüthrich, M., Agrawal, V., Winther, O., Bauer, S., and Schölkopf, B · 2020
Cited alongside, same era.
Underspecification presents challenges for credibility in modern machine learning
D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al · 2020
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M · 2020
Cited alongside, same era.
When are solutions connected in deep networks?
Nguyen, Q., Bréchet, P., and Mondelli, M · 2021
Later among the works it cites.
Editing a classifier by rewriting its prediction rules
Santurkar, S., Tsipras, D., Elango, M., Bau, D., Torralba, A., and Madry, A · 2021
Later among the works it cites.
Towards causal representation learning
Schölkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y · 2021
Later among the works it cites.
Which shortcut cues will dnns choose? a study from the parameter-space perspective
Scimeca, L., Oh, S. J., Chun, S., Poli, M., and Yun, S · 2021
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Simsek, B., Ged, F., Jacot, A., Spadaro, F., Hongler, C., Gerstner, W., and Brea, J · 2021
Later among the works it cites.
Noether’s learning dynamics: Role of symmetry breaking in neural networks
Tanaka, H. and Kunin, D · 2021
Later among the works it cites.
Designing counterfactual generators using deep model inversion
Thiagarajan, J., Narayanaswamy, V. S., Rajan, D., Liang, J., Chaudhari, A., and Spanias, A · 2021
Later among the works it cites.
Self-supervised learning with data augmentations provably isolates content from style
Von Kügelgen, J., Sharma, Y., Gresele, L., Brendel, W., Schölkopf, B., Besserve, M., and Locatello, F · 2021
Later among the works it cites.
A fine-grained analysis on distribution shift
Wiles, O., Gowal, S., Stimberg, F., Alvise-Rebuffi, S., Ktena, I., Cemgil, T., et al · 2021
Later among the works it cites.
Learning neural network subspaces
Wortsman, M., Horton, M. C., Guestrin, C., Farhadi, A., and Rastegari, M · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S · 2022
Closest in time.
Ensemble of averages: Improving model selection and boosting performance in domain generalization
Arpit, D., Wang, H., Zhou, Y., and Xiong, C · 2022
Closest in time.
Distinguishing rule and exemplar-based generalization in learning systems
Dasgupta, I., Grant, E., and Griffiths, T · 2022
Closest in time.
Linear connectivity reveals generalization strategies
Juneja, J., Bansal, R., Cho, K., Sedoc, J., and Saphra, N · 2022
Closest in time.
When do flat minima optimizers work?
Kaddour, J., Liu, L., Silva, R., and Kusner, M · 2022
Closest in time.
Modeling the data-generating process is necessary for out-of-distribution generalization
Kaur, J. N., Kiciman, E., and Sharma, A · 2022
Closest in time.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P · 2022
Closest in time.
Characterizing Datapoints via Second-Split Forgetting
Maini, P., Garg, S., Lipton, Z. C., and Kolter, J. Z · 2022
Closest in time.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Closest in time.
An empirical analysis of memorization in fine-tuned autoregressive language models
Mireshghallah, F., Uniyal, A., Wang, T., Evans, D. K., and Berg-Kirkpatrick, T · 2022
Closest in time.
Fast model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2022
Closest in time.
Measuring representational robustness of neural networks through shared invariances
Nanda, V., Speicher, T., Kolling, C., Dickerson, J. P., Gummadi, K., and Weller, A · 2022
Closest in time.
Pittorino, F., Ferraro, A., Perugini, G., Feinauer, C., Baldassi, C., and Zecchina, R · 2022
Closest in time.
Diverse weight averaging for out-of-distribution generalization
Rame, A., Kirchmeyer, M., Rahier, T., Rakotomamonjy, A., Gallinari, P., and Cord, M · 2022
Closest in time.
Spherical perspective on learning with normalization layers
Roburin, S., de Mont-Marin, Y., Bursuc, A., Marlet, R., Pérez, P., and Aubry, M · 2022
Closest in time.
Rosenfeld, E., Ravikumar, P., and Risteski, A · 2022
Closest in time.
Predicting is not understanding: Recognizing and addressing underspecification in machine learning
Teney, D., Peyrard, M., and Abbasnejad, E · 2022
Closest in time.
Augmentations in graph contrastive learning: Current methodological flaws & towards better practices
Trivedi, P., Lubana, E. S., Yan, Y., Yang, Y., and Koutra, D · 2022
Closest in time.
Truncated normal distribution, 2022
Truncated Gaussian Distribution · 2022
Closest in time.
Task-specific skill localization in fine-tuned language models
Panigrahi, A., Saunshi, N., Zhao, H., and Arora, S · 2023
Closest in time.