The intriguing role of module criticality in the generalization of deep networks
N. Chatterji, B. Neyshabur, and H. Sedghi · 2020
Later among the works it cites.
Linear Mode Connectivity and the Lottery Ticket Hypothesis
J. Frankle, G. K. Dziugaite, D. Roy, and M. Carbin · 2020
Later among the works it cites.
COGS: A Compositional Generalization Challenge Based on Semantic Interpretation
N. Kim and T. Linzen · 2020
Later among the works it cites.
Berts of a feather do not generalize together: Large variability in generalization across models with similar test set performance
Original
R. T. McCoy, J. Min, and T. Linzen · 2020
Later among the works it cites.
What is being transferred in transfer learning?
B. Neyshabur, H. Sedghi, and C. Zhang · 2020
Later among the works it cites.
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)
A. Warstadt, Y. Zhang, X. Li, H. Liu, and S. R. Bowman · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush · 2020
Later among the works it cites.
Loss Surface Simplexes for Mode Connecting Volumes and Fast Ensembling
Original
G. W. Benton, W. J. Maddox, S. Lotfi, and A. G. Wilson · 2021
Later among the works it cites.
The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks
Original
R. Entezari, H. Sedghi, O. Saukh, and B. Neyshabur · 2021
Later among the works it cites.
Natural adversarial examples
D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song · 2021
Later among the works it cites.
Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization
S. Jastrzebski, D. Arpit, O. Astrand, G. B. Kerg, H. Wang, C. Xiong, R. Socher, K. Cho, and K. J. Geras · 2021
Later among the works it cites.
Predicting inductive biases of pre-trained models
C. Lovering, R. Jha, T. Linzen, and E. Pavlick · 2021
Later among the works it cites.
Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization, Oct. 2021
Original
J. Miller, R. Taori, A. Raghunathan, S. Sagawa, P. W. Koh, V. Shankar, P. Liang, Y. Carmon, and L. Schmidt · 2021
Later among the works it cites.
The MultiBERTs: BERT Reproductions for Robustness Analysis
T. Sellam, S. Yadlowsky, I. Tenney, J. Wei, N. Saphra, A. D’Amour, T. Linzen, J. Bastings, I. R. Turc, J. Eisenstein, D. Das, and E. Pavlick · 2021
Later among the works it cites.
Git Re-Basin: Merging Models modulo Permutation Symmetries, Sept. 2022
Original
S. K. Ainsworth, J. Hayase, and S. Srinivasa · 2022
Closest in time.
On the Maximum Hessian Eigenvalue and Generalization
Original
S. Kaur, J. Cohen, and Z. C. Lipton · 2022
Closest in time.
Fine-tuning can distort pretrained features and underperform out-of-distribution
A. Kumar, A. Raghunathan, R. M. Jones, T. Ma, and P. Liang · 2022
Closest in time.
The Multiscale Structure of Neural Network Loss Functions: The Effect on Optimization and Origin
Original
C. Ma, L. Wu, and L. Ying · 2022
Closest in time.
Can Neural Nets Learn the Same Model Twice? Investigating Reproducibility and Double Descent from the Decision Boundary Perspective
Original
G. Somepalli, L. Fowl, A. Bansal, P. Yeh-Chiang, Y. Dar, R. Baraniuk, M. Goldblum, and T. Goldstein · 2022
Closest in time.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Original
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt · 2022
Closest in time.
How grammatical is character-level neural machine translation? assessing MT quality with contrastive translation pairs
R. Sennrich · 2060
Closest in time.