Fetching the paper…
Reading the bibliography…
Learning representations of well-trained neural network models holds the promise to provide an understanding of the inner workings of those models.
Pattern Recognition and Machine Learning
Bishop, C. M · 2006
Earlier work this paper cites.
HyperNetworks, 2016
Ha, D., Dai, A., and Le, Q. V · 2016
Earlier work this paper cites.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Generating Neural Networks with Neural Networks
Deutsch, L · 2018
Earlier work this paper cites.
Essentially No Barriers in Neural Network Energy Landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F · 2018
Earlier work this paper cites.
Tune: A Research Platform for Distributed Model Selection and Training
Liaw, R., Liang, E., Nishihara, R., Moritz, P., Gonzalez, J. E., and Stoica, I · 2018
Earlier work this paper cites.
Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates, May 2018
Smith, L. N. and Topin, N · 2018
Earlier work this paper cites.
MetaInit: Initializing learning by learning to initialize
Dauphin, Y. N. and Schoenholz, S · 2019
Earlier work this paper cites.
Large Scale Structure of Neural Network Loss Landscapes
Fort, S. and Jastrzebski, S · 2019
Earlier work this paper cites.
Linear Mode Connectivity and the Lottery Ticket Hypothesis
Frankle, J., Dziugaite, G., Roy, D. M., and Carbin, M · 2019
Earlier work this paper cites.
Predicting the Generalization Gap in Deep Networks with Margin Distributions
Jiang, Y., Krishnan, D., Mobahi, H., and Bengio, S · 2019
Earlier work this paper cites.
Similarity of Neural Network Representations Revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G · 2019
Earlier work this paper cites.
HyperVAE: A Minimum Description Length Variational Hyper-Encoding Network
Nguyen, P., Tran, T., Gupta, S., Rana, S., and Dam, H.-C · 2019
Earlier work this paper cites.
On Connected Sublevel Sets in Deep Learning
Nguyen, Q. N · 2019
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
HyperGAN: A Generative Model for Diverse, Performant Neural Networks
Ratzlaff, N. and Fuxin, L · 2019
Earlier work this paper cites.
Towards Task and Architecture-Independent Generalization Gap Predictors
Yak, S., Gonzalvo, J., and Mazzawi, H · 2019
Earlier work this paper cites.
Graph HyperNetworks for Neural Architecture Search
Zhang, C., Ren, M., and Urtasun, R · 2019
Cited alongside, same era.
Computing the Testing Error Without a Testing Set
Corneanu, C. A., Escalera, S., and Martinez, A. M · 2020
Cited alongside, same era.
Classifying the classifier: Dissecting the weight space of neural networks
Eilertsen, G., Jönsson, D., Ropinski, T., Unger, J., and Ynnerman, A · 2020
Cited alongside, same era.
Heavy-tailed Universality predicts trends in test accuracies for very large pre-trained deep neural networks
Martin, C. H. and Mahoney, M. W · 2020
Cited alongside, same era.
Predicting Neural Network Accuracy from Weights
Unterthiner, T., Keysers, D., Gelly, S., Bousquet, O., and Tolstikhin, I · 2020
Cited alongside, same era.
Loss Surface Simplexes for Mode Connecting Volumes and Fast Ensembling
Benton, G. W., Maddox, W. J., Lotfi, S., and Wilson, A. G · 2021
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, June 2022
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Meta-Learning via Classifier(-free) Diffusion Guidance
Nava, E., Kobayashi, S., Yin, Y., Katzschmann, R. K., and Grewe, B. F · 2022
Later among the works it cites.
Learning to Learn with Generative Models of Neural Network Checkpoints, September 2022
Peebles, W., Radosavovic, I., Brooks, T., Efros, A. A., and Malik, J · 2022
Later among the works it cites.
Hyper-Representations for Pre-Training and Transfer Learning
Schürholt, K., Knyazev, B., Giró-i-Nieto, X., and Borth, D · 2022
Later among the works it cites.
Yang, Y., Theisen, R., Hodgkinson, L., Gonzalez, J. E., Ramchandran, K., Martin, C. H., and Mahoney, M. W · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Parameter Prediction for Unseen Deep Architectures
Knyazev, B., Drozdzal, M., Taylor, G. W., and Romero-Soriano, A · 2021
Cited alongside, same era.
On Monotonic Linear Interpolation of Neural Network Parameters
Lucas, J. R., Bae, J., Zhang, M. R., Fort, S., Zemel, R., and Grosse, R. B · 2021
Cited alongside, same era.
Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning
Martin, C. H. and Mahoney, M. W · 2021
Cited alongside, same era.
Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
Martin, C. H., Peng, T. S., and Mahoney, M. W · 2021
Cited alongside, same era.
Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction
Schürholt, K., Kostadinov, D., and Borth, D · 2021
Cited alongside, same era.
Scaling Local Self-Attention for Parameter Efficient Visual Backbones
Vaswani, A., Ramachandran, P., Srinivas, A., Parmar, N., Hechtman, B., and Shlens, J · 2021
Cited alongside, same era.
HyperTransformer: Model Generation for Supervised and Semi-Supervised Few-Shot Learning
Zhmoginov, A., Sandler, M., and Vladymyrov, M · 2022
Later among the works it cites.
Set-based Neural Network Encoding, May 2023
Andreis, B., Bedionita, S., and Hwang, S. J · 2023
Later among the works it cites.
On Privileged and Convergent Bases in Neural Network Representations, July 2023
Brown, D., Vyas, N., and Bansal, Y · 2023
Later among the works it cites.
Deep Learning on Implicit Neural Representations of Shapes, February 2023
De Luigi, L., Cardace, A., Spezialetti, R., Ramirez, P. Z., Salti, S., and Di Stefano, L · 2023
Later among the works it cites.
Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models?
Knyazev, B., Hwang, D., and Lacoste-Julien, S · 2023
Later among the works it cites.
FFCV: Accelerating Training by Removing Data Bottleneck
Leclerc, G., Ilyas, A., Engstrom, L., Park, S. M., Salman, H., and Madry, A · 2023
Later among the works it cites.
Singular Value Representation: A New Graph Perspective On Neural Networks, February 2023
Meller, D. and Berkouk, N · 2023
Later among the works it cites.
Compact and Optimal Deep Learning with Recurrent Parameter Generators
Wang, J., Chen, Y., Yu, S. X., Cheung, B., and LeCun, Y · 2023
Later among the works it cites.
Neural Networks Are Graphs!Graph Neural Networks for Equivariant Processing of Neural Networks
Zhang, D. W., Kofinas, M., Zhang, Y., Chen, Y., Burghouts, G. J., and Snoek, C. G. M · 2023
Later among the works it cites.
Graph neural networks for learning equivariant representations of neural networks
Kofinas, M., Knyazev, B., Zhang, Y., Chen, Y., Burghouts, G. J., Gavves, E., Snoek, C. G. M., and Zhang, D. W · 2024
Closest in time.
Osmr/imgclsmob, January 2024
Sémery, O · 2024
Closest in time.