Fetching the paper…
Reading the bibliography…
Measuring the similarity of different representations of neural architectures is a fundamental task and an open research challenge for the machine learning community.
Relations Between Two Sets of Variates
Harold Hotelling · 1936
Earlier work this paper cites.
Measuring statistical dependence with Hilbert-Schmidt norms
Arthur Gretton, Olivier Bousquet, Alex Smola, and Bernhard Schölkopf · 2005
Earlier work this paper cites.
Measuring and testing dependence by correlation of distances
Gábor J. Székely, Maria L. Rizzo, and Nail K. Bakirov · 2007
Earlier work this paper cites.
Representational similarity analysis - connecting the branches of systems neuroscience
Nikolaus Kriegeskorte, Marieke Mur, and Peter Bandettini · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John Hopcroft · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Cultural shift or linguistic drift? comparing two computational measures of semantic change
William L. Hamilton, Jure Leskovec, and Dan Jurafsky · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Revisiting semi-supervised learning with graph embeddings
Zhilin Yang, William Cohen, and Ruslan Salakhudinov · 2016
Earlier work this paper cites.
Inductive representation learning on large graphs
William L. Hamilton, Rex Ying, and Jure Leskovec · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Earlier work this paper cites.
SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Beware of the beginnings: intermediate and higher-level representations in deep neural networks are strongly affected by weight initialization
Johannes Mehrer, Nikolaus Kriegeskorte, and Tim Kietzmann · 2018
Earlier work this paper cites.
Insights on representational similarity in neural networks with canonical correlation
Ari Morcos, Maithra Raghu, and Samy Bengio · 2018
Earlier work this paper cites.
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2018
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
On the Dimensionality of Word Embedding
Zi Yin and Yuanyuan Shen · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Fast graph representation learning with PyTorch Geometric
Matthias Fey and Jan E. Lenssen · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Cited alongside, same era.
On the downstream performance of compressed word embeddings
Avner May, Jian Zhang, Tri Dao, and Christopher Ré · 2019
Cited alongside, same era.
Anatomy of catastrophic forgetting: Hidden representations and task semantics
Vinay Venkatesh Ramasesh, Ethan Dyer, and Maithra Raghu · 2021
Later among the works it cites.
Using distance on the riemannian manifold to compare representations in brain and in models
Mahdiyar Shahbazi, Ali Shirali, Hamid Aghajan, and Hamed Nili · 2021
Later among the works it cites.
All bark and no bite: Rogue dimensions in transformer language models obscure representational quality
William Timkey and Marten van Schijndel · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Generalized shape metrics on neural representations
Alex H. Williams, Erin Kunz, Simon Kornblith, and Scott Linderman · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
EDA: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou · 2019
Cited alongside, same era.
Position-aware graph neural networks
Jiaxuan You, Rex Ying, and Jure Leskovec · 2019
Cited alongside, same era.
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec · 2020
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Cited alongside, same era.
What happens to BERT embeddings during fine-tuning?
Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney · 2020
Cited alongside, same era.
Representation topology divergence: A method for comparing neural network representations
Serguei Barannikov, Ilya Trofimov, Nikita Balabin, and Evgeny Burnaev · 2022
Later among the works it cites.
GULP: a prediction-based metric between representations
Enric Boix-Adserà, Hannah Lawrence, George Stepaniants, and Philippe Rigollet · 2022
Later among the works it cites.
Deconfounded representation similarity for comparison of neural networks
Tianyu Cui, Yogesh Kumar, Pekka Marttinen, and Samuel Kaski · 2022
Later among the works it cites.
On the inadequacy of CKA as a measure of similarity in deep learning
MohammadReza Davari, Stefan Horoi, Amine Natik, Guillaume Lajoie, Guy Wolf, and Eugene Belilovsky · 2022
Later among the works it cites.
Building and interpreting deep similarity models
Oliver Eberle, Jochen Büttner, Florian Kräutli, Klaus-Robert Müller, Matteo Valleriani, and Grégoire Montavon · 2022
Later among the works it cites.
The multiBERTs: BERT reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Raluca Turc, Jacob Eisenstein, Dipanjan Das, and Ellie Pavlick · 2022
Later among the works it cites.
Towards understanding the instability of network embedding
Chenxu Wang, Wei Rao, Wenna Guo, Pinghui Wang, Jun Liu, and Xiaohong Guan · 2022
Later among the works it cites.
Obstacles to inferring mechanistic similarity using representational similarity analysis
Marin Dujmović, Jeffrey S Bowers, Federico Adolfi, and Gaurav Malhotra · 2023
Later among the works it cites.
Similarity of neural network models: A survey of functional and representational measures
Max Klabunde, Tobias Schumacher, Markus Strohmaier, and Florian Lemmerich · 2023
Later among the works it cites.
Getting aligned on representational alignment
Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C. Love, Christopher J. Cueva, Erin Grant, Iris Groen, Jascha Achterberg, Joshua B. Tenenbaum, Katherine M. Collins, Katherine L. Hermann, Kerem Oktar, Klaus Greff, Martin N. Hebart, Nathan Cloos, Nikolaus Kriegeskorte, Nori Jacoby, Qiuyi Zhang, Raja Marjieh, Robert Geirhos, Sherol Chen, Simon Kornblith, Sunayana Rane, Talia Konkle, Thomas P. O’Connell, Thomas Unterthiner, Andrew K. Lampinen, Klaus-Robert Müller, Mariya Toneva, and Thomas L. Griffiths · 2023
Later among the works it cites.
Does representation similarity capture function similarity?
Lucas Hayne, Heejung Jung, and R. Carter · 2024
Closest in time.
Sparse autoencoders reveal universal feature spaces across large language models
Michael Lan, Philip Torr, Austin Meek, Ashkan Khakzar, David Krueger, and Fazl Barez · 2024
Closest in time.
Conclusions about neural network to brain alignment are profoundly impacted by the similarity measure
Ansh Soni, Sudhanshu Srivastava, Konrad Kording, and Meenakshi Khosla · 2024
Closest in time.
SmolLM2: When smol goes big – data-centric training of a small language model
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, et al · 2025
Closest in time.