Fetching the paper…
Reading the bibliography…
Pre-trained language encoders -- functions that represent text as vectors -- are an integral component of many NLP tasks.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
A generalized solution of the orthogonal procrustes problem
Peter H. Schönemann. 1966 · 1966
Earlier work this paper cites.
Canonical ridge and econometrics of joint production
Hrishikesh D. Vinod. 1976 · 1976
Earlier work this paper cites.
Data structures and algorithms for nearest neighbor search in general metric spaces
Peter N. Yianilos. 1993 · 1993
Earlier work this paper cites.
A Course in Metric Geometry
Dmitri Burago, Yuri Burago, and Sergei Ivanov. 2001 · 2001
Earlier work this paper cites.
Performance guarantees for hierarchical clustering
Sanjoy Dasgupta and Philip M. Long. 2005 · 2002
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Algebra
Serge Lang. 2002 · 2002
Earlier work this paper cites.
Metric spaces, generalized logic, and closed categories
F. W. Lawvere. 2002 · 2002
Earlier work this paper cites.
Canonical correlation analysis: An overview with application to learning methods
David Roi Hardoon, Sándor Szedmák, and John Shawe-Taylor. 2004 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Measuring statistical dependence with Hilbert-Schmidt norms
Arthur Gretton, Olivier Bousquet, Alex Smola, and Bernhard Schölkopf. 2005 · 2005
Earlier work this paper cites.
Numerical Recipes: The Art of Scientific Computing , 3rd edition
William H. Press, Saul A. Teukolsky, William T. Vetterling, and Brian P. Flannery. 2007 · 2007
Earlier work this paper cites.
Representational similarity analysis - connecting the branches of systems neuroscience
Nikolaus Kriegeskorte, Marieke Mur, and Peter Bandettini. 2008 · 2008
Earlier work this paper cites.
The impact of triangular inequality violations on medoid-based clustering
Saaid Baraty, Dan A. Simovici, and Catalin Zara. 2011 · 2011
Earlier work this paper cites.
The Statistical Theory of Shape
C.G. Small. 2012 · 2012
Cited alongside, same era.
Non-Hausdorff Topology and Domain Theory: Selected Topics in Point-Set Topology
Jean Goubault-Larrecq. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Convergent learning: Do different neural networks learn the same representations?
Li, Yixuan and Yosinski, Jason and Clune, Jeff and Lipson, Hod and Hopcroft, John. 2015 · 2015
Cited alongside, same era.
Survey on distance metric learning and dimensionality reduction in data mining
Fei Wang and Jimeng Sun. 2015 · 2015
Cited alongside, same era.
A mathematical theory for clustering in metric spaces
C. Chang, W. Liao, Y. Chen, and L. Liou. 2016 · 2016
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine A. Heller, Dan I. Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yi-An Ma, Cory Y. McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin G. Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, and D. Sculley. 2020 · 2020
Later among the works it cites.
BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance
R. Thomas McCoy, Junghyun Min, and Tal Linzen. 2020 · 2020
Later among the works it cites.
Revisiting model stitching to compare neural representations
Yamini Bansal, Preetum Nakkiran, and Boaz Barak. 2021 · 2021
Later among the works it cites.
Similarity and matching of neural network representations
Adrián Csiszárik, Péter Kőrösi-Szabó, Ákos Matszangosz, Gergely Papp, and Dániel Varga. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Diachronic word embeddings reveal statistical laws of semantic change
William L. Hamilton, Jure Leskovec, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
On the properties of the softmax function with application in game theory and reinforcement learning
Bolin Gao and Lacra Pavel. 2017 · 2017
Cited alongside, same era.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017 · 2017
Cited alongside, same era.
Insights on representational similarity in neural networks with canonical correlation
Ari Morcos, Maithra Raghu, and Samy Bengio. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. 2019 · 2019
Cited alongside, same era.
Grounding representation similarity through statistical testing
Frances Ding, Jean-Stanislas Denain, and Jacob Steinhardt. 2021 · 2021
Later among the works it cites.
Using distance on the Riemannian manifold to compare representations in brain and in models
Mahdiyar Shahbazi, Ali Shirali, Hamid Aghajan, and Hamed Nili. 2021 · 2021
Later among the works it cites.
Generalized shape metrics on neural representations
Alex H. Williams, Erin Kunz, Simon Kornblith, and Scott Linderman. 2021 · 2021
Later among the works it cites.
Are larger pretrained language models uniformly better? Comparing performance at the instance level
Ruiqi Zhong, Dhruba Ghosh, Dan Klein, and Jacob Steinhardt. 2021 · 2021
Later among the works it cites.
Gulp: a prediction-based metric between representations
Enric Boix-Adsera, Hannah Lawrence, George Stepaniants, and Philippe Rigollet. 2022 · 2022
Later among the works it cites.
The MultiBERTs: BERT reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, et al. 2022 · 2022
Later among the works it cites.
Formal aspects of language modeling
Ryan Cotterell, Anej Svete, Clara Meister, Tianyu Liu, and Li Du. 2023 · 2023
Later among the works it cites.
A measure-theoretic characterization of tight language models
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Similarity of neural network models: A survey of functional and representational measures
Max Klabunde, Tobias Schumacher, Markus Strohmaier, and Florian Lemmerich. 2023 · 2023
Later among the works it cites.
All roads lead to Rome? Exploring the invariance of transformers’ representations
Yuxin Ren, Qipeng Guo, Zhijing Jin, Shauli Ravfogel, Mrinmaya Sachan, Bernhard Schölkopf, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.