Fetching the paper…
Reading the bibliography…
Can we identify the weights of a neural network by probing its input-output mapping? At first glance, this problem seems to have many solutions because of permutation, overparameterisation and activation function symmetries.
Learning internal representations by error propagation
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Neural net algorithms that learn in polynomial time from examples and queries
Baum, E. B · 1991
Earlier work this paper cites.
Cones of matrices and set-functions and 0–1 optimization
Lovász, L. and Schrijver, A · 1991
Earlier work this paper cites.
Uniqueness of the weights for minimal feedforward nets with a given input-output map
Sussmann, H. J · 1992
Earlier work this paper cites.
Reconstructing a neural net from its output
Fefferman, C · 1994
Earlier work this paper cites.
On-line learning in soft committee machines
Saad, D. and Solla, S. A · 1995
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Fukumizu, K. and Amari, S.-i · 2000
Earlier work this paper cites.
Global optimization with polynomials and the problem of moments
Lasserre, J. B · 2001
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Feature selection with the boruta package
Kursa, M. B. and Rudnicki, W. R · 2010
Earlier work this paper cites.
Algorithms for hierarchical clustering: an overview
Murtagh, F. and Contreras, P · 2012
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Janzamin, M., Sedghi, H., and Anandkumar, A · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Data-free parameter pruning for deep neural networks
Srinivas, S. and Babu, R. V · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Network trimming: A data-driven neuron pruning approach towards efficient deep architectures
Hu, H., Peng, R., Tai, Y.-W., and Tang, C.-K · 2016
Earlier work this paper cites.
Stealing machine learning models via prediction { \{ APIs } \}
Tramèr, F., Zhang, F., Juels, A., Reiter, M. K., and Ristenpart, T · 2016
Earlier work this paper cites.
Self-normalizing neural networks
Klambauer, G., Unterthiner, T., Mayr, A., and Hochreiter, S · 2017
Earlier work this paper cites.
Searching for activation functions
Ramachandran, P., Zoph, B., and Le, Q. V · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Zhong, K., Song, Z., Jain, P., Bartlett, P. L., and Dhillon, I. S · 2017
Earlier work this paper cites.
Adversarial attacks and defences: A survey
Chakraborty, A., Alam, M., Dey, V., Chattopadhyay, A., and Mukhopadhyay, D · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Earlier work this paper cites.
The loss landscape of overparameterized neural networks
Cooper, Y · 2018
Earlier work this paper cites.
Born again neural networks
Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and Anandkumar, A · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
The landscape of empirical risk for nonconvex losses
Mei, S., Bai, Y., and Montanari, A · 2018
Cited alongside, same era.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2019
Cited alongside, same era.
Robust and resource efficient identification of shallow neural networks by fewest samples
Fornasier, M., Vybíral, J., and Daubechies, I · 2021
Later among the works it cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Later among the works it cites.
Data-independent structured pruning of neural networks via coresets
Mussay, B., Feldman, D., Zhou, S., Braverman, V., and Osadchy, M · 2021
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Şimşek, B., Ged, F., Jacot, A., Spadaro, F., Hongler, C., Gerstner, W., and Brea, J · 2021
Later among the works it cites.
Does knowledge distillation really work?
Stanton, S., Izmailov, P., Kirichenko, P., Alemi, A. A., and Wilson, A. G · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust and resource-efficient identification of two hidden layer neural networks
Fornasier, M., Klock, T., and Rauchensteiner, M · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Cited alongside, same era.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Goldt, S., Advani, M., Saxe, A. M., Krzakala, F., and Zdeborová, L · 2019
Cited alongside, same era.
Mish: A self regularized non-monotonic activation function
Misra, D · 2019
Cited alongside, same era.
On the connection between learning two-layer neural networks and tensor decomposition
Mondelli, M. and Montanari, A · 2019
Cited alongside, same era.
Fundamental bounds on learning performance in neural circuits
Raman, D. V., Rotondo, A. P., and O’Leary, T · 2019
Cited alongside, same era.
Sun, J · 2021
Later among the works it cites.
Affine symmetries and neural network identifiability
Vlačić, V. and Bölcskei, H · 2021
Later among the works it cites.
A survey on neural network interpretability
Zhang, Y., Tiňo, P., Leonardis, A., and Tang, K · 2021
Later among the works it cites.
An initial alignment between neural network and target is needed for gradient descent to learn
Abbe, E., Cornacchia, E., Hazla, J., and Marquis, C · 2022
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S · 2022
Later among the works it cites.
Knowledge distillation: A good teacher is patient and consistent
Beyer, L., Zhai, X., Royer, A., Markeeva, L., Anil, R., and Kolesnikov, A · 2022
Later among the works it cites.
Otov2: Automatic, generic, user-friendly
Chen, T., Liang, L., Tianyu, D., Zhu, Z., and Zharkov, I · 2022
Later among the works it cites.
Toy models of superposition
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C · 2022
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2022
Later among the works it cites.
Finite sample identification of wide shallow neural networks with biases
Fornasier, M., Klock, T., Mondelli, M., and Rauchensteiner, M · 2022
Later among the works it cites.
Repair: Renormalizing permuted activations for interpolation repair
Jordan, K., Sedghi, H., Saukh, O., Entezari, R., and Neyshabur, B · 2022
Later among the works it cites.
I know what you trained last summer: A survey on stealing machine learning models and defences
Oliynyk, D., Mayer, R., and Rauber, A · 2022
Later among the works it cites.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
Rotskoff, G. and Vanden-Eijnden, E · 2022
Later among the works it cites.
An embedding of relu networks and an analysis of their identifiability
Stock, P. and Gribonval, R · 2022
Later among the works it cites.
Neural network identifiability for a family of sigmoidal nonlinearities
Vlačić, V. and Bölcskei, H · 2022
Later among the works it cites.
Brea, J., Martinelli, F., Şimşek, B., and Gerstner, W · 2023
Closest in time.
Stable recovery of entangled weights: Towards robust identification of deep neural networks from minimal samples
Fiedler, C., Fornasier, M., Klock, T., and Rauchensteiner, M · 2023
Closest in time.
Hidden symmetries of relu networks
Grigsby, E., Lindsey, K., and Rolnick, D · 2023
Closest in time.
To reverse engineer an entire nervous system
Haspel, G., Boyden, E. S., Brown, J., Church, G., Cohen, N., Fang-Yen, C., Flavell, S., Goodman, M. B., Hart, A. C., Hobert, O., et al · 2023
Closest in time.
Stealing part of a production language model
Carlini, N., Paleka, D., Dvijotham, K. D., Steinke, T., Hayase, J., Cooper, A. F., Lee, K., Jagielski, M., Nasr, M., Conmy, A., et al · 2024
Closest in time.
Kan: Kolmogorov-arnold networks
Liu, Z., Wang, Y., Vaidya, S., Ruehle, F., Halverson, J., Soljačić, M., Hou, T. Y., and Tegmark, M · 2024
Closest in time.
A petavoxel fragment of human cerebral cortex reconstructed at nanoscale resolution
Shapson-Coe, A., Januszewski, M., Berger, D. R., Pope, A., Wu, Y., Blakely, T., Schalek, R. L., Li, P. H., Wang, S., Maitin-Shepard, J., et al · 2024
Closest in time.
Should under-parameterized student networks copy or average teacher weights?
Şimşek, B., Bendjeddou, A., Gerstner, W., and Brea, J · 2024
Closest in time.