Fetching the paper…
Reading the bibliography…
Recent work has sought to understand the behavior of neural networks by comparing representations between layers and between different trained models.
A method for synthesis of factor analysis studies
Tucker, L. R · 1951
Earlier work this paper cites.
Numerical methods for computing angles between linear subspaces
Björck, Å. and Golub, G. H · 1973
Earlier work this paper cites.
A unifying tool for linear multivariate statistical methods: the RV-coefficient
Robert, P. and Escoufier, Y · 1976
Earlier work this paper cites.
Canonical ridge and econometrics of joint production
Vinod, H. D · 1976
Earlier work this paper cites.
Matrix correlation
Ramsay, J., ten Berge, J., and Styan, G · 1984
Earlier work this paper cites.
Second order properties of error surfaces: Learning time and generalization
LeCun, Y., Kanter, I., and Solla, S. A · 1991
Earlier work this paper cites.
On the geometry of feedforward neural network error surfaces
Chen, A. M., Lu, H.-m., and Hecht-Nielsen, R · 1993
Earlier work this paper cites.
The canonical correlations of matrix pairs and their numerical computation
Golub, G. H. and Zha, H · 1995
Earlier work this paper cites.
Representation is representation of similarities
Edelman, S · 1998
Earlier work this paper cites.
Content and cluster analysis: assessing representational similarity in neural systems
Laakso, A. and Cottrell, G · 2000
Earlier work this paper cites.
Distributed and overlapping representations of faces and objects in ventral temporal cortex
Haxby, J. V., Gobbini, M. I., Furey, M. L., Ishai, A., Schouten, J. L., and Pietrini, P · 2001
Earlier work this paper cites.
On kernel-target alignment
Cristianini, N., Shawe-Taylor, J., Elisseeff, A., and Kandola, J. S · 2002
Earlier work this paper cites.
The geometry of kernel canonical correlation analysis
Kuss, M. and Graepel, T · 2003
Earlier work this paper cites.
Measuring statistical dependence with Hilbert-Schmidt norms
Gretton, A., Bousquet, O., Smola, A., and Schölkopf, B · 2005
Earlier work this paper cites.
Tucker’s congruence coefficient as a meaningful index of factor similarity
Lorenzo-Seva, U. and Ten Berge, J. M · 2006
Earlier work this paper cites.
Supervised feature selection via dependence estimation
Song, L., Smola, A., Gretton, A., Borgwardt, K. M., and Bedo, J · 2007
Earlier work this paper cites.
Functional compartmentalization and viewpoint generalization within the macaque face-processing system
Freiwald, W. A. and Tsao, D. Y · 2010
Earlier work this paper cites.
Canonical correlation clarified by singular value decomposition, 2011
Press, W. H · 2011
Earlier work this paper cites.
The representation of biological classes in the human brain
Connolly, A. C., Guntupalli, J. S., Gors, J., Hanke, M., Halchenko, Y. O., Wu, Y.-C., Abdi, H., and Haxby, J. V · 2012
Cited alongside, same era.
Algorithms for learning kernels based on centered alignment
Cortes, C., Mohri, M., and Rostamizadeh, A · 2012
Cited alongside, same era.
Equivalence of distance-based and RKHS-based statistics in hypothesis testing
Sejdinovic, D., Sriperumbudur, B., Gretton, A., and Fukumizu, K · 2013
Cited alongside, same era.
Deep supervised, but not unsupervised, models may explain it cortical representation
Khaligh-Razavi, S.-M. and Kriegeskorte, N · 2014
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Cited alongside, same era.
Performance-optimized hierarchical models predict neural responses in higher visual cortex
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Advani, M. S. and Saxe, A. M · 2017
Later among the works it cites.
Density estimation using real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2017
Later among the works it cites.
A learned representation for artistic style
Dumoulin, V., Shlens, J., and Kudlur, M · 2017
Later among the works it cites.
SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Offline bilingual word vectors, orthogonal transformations and the inverted softmax
Smith, S. L., Turban, D. H., Hamblin, S., and Hammerla, N. Y · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yamins, D. L., Hong, H., Cadieu, C. F., Solomon, E. A., Seibert, D., and DiCarlo, J. J · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Convergent learning: Do different neural networks learn the same representations?
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J · 2015
Cited alongside, same era.
Asymmetrically weighted CCA and hierarchical kernel sentence embedding for multimodal retrieval
Mroueh, Y., Marcheret, E., and Goel, V · 2015
Cited alongside, same era.
FitNets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2015
Cited alongside, same era.
Introduction to Regression Procedures
SAS Institute · 2015
Cited alongside, same era.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2015
Cited alongside, same era.
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
Dynamics of learning in MLP: Natural gradient and singularity revisited
Amari, S.-i., Ozeki, T., Karakida, R., Yoshida, Y., and Okada, M · 2018
Later among the works it cites.
Findings of the 2018 Conference on Machine Translation (WMT18)
Bojar, O., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Huck, M., Koehn, P., and Monz, C · 2018
Later among the works it cites.
i-RevNet: Deep invertible networks
Jacobsen, J.-H., Smeulders, A. W., and Oyallon, E · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Later among the works it cites.
Deep neural networks as gaussian processes
Lee, J., Sohl-dickstein, J., Pennington, J., Novak, R., Schoenholz, S., and Bahri, Y · 2018
Later among the works it cites.
Insights on representational similarity in neural networks with canonical correlation
Morcos, A., Raghu, M., and Bengio, S · 2018
Later among the works it cites.
Skip connections eliminate singularities
Orhan, E. and Pitkow, X · 2018
Later among the works it cites.
Tensor2tensor for neural machine translation
Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A. N., Gouws, S., Jones, L., Kaiser, Ł., Kalchbrenner, N., Parmar, N., et al · 2018
Later among the works it cites.
Towards understanding learning representations: To what extent do different neural networks learn the same representation
Wang, L., Hu, L., Gu, J., Wu, Y., Hu, Z., He, K., and Hopcroft, J. E · 2018
Later among the works it cites.
Deep convolutional networks as shallow Gaussian processes
Garriga-Alonso, A., Rasmussen, C. E., and Aitchison, L · 2019
Closest in time.
Bayesian deep convolutional networks with many channels are Gaussian processes
Novak, R., Xiao, L., Bahri, Y., Lee, J., Yang, G., Abolafia, D. A., Pennington, J., and Sohl-dickstein, J · 2019
Closest in time.