Fetching the paper…
Reading the bibliography…
Given the success of Large Language Models (LLMs), there has been considerable interest in studying the properties of model activations.
Dimensionality compression and expansion in deep neural networks
Stefano Recanatesi, Matthew Farrell, Madhu Advani, Timothy Moore, Guillaume Lajoie, and Eric Shea-Brown · 1906
Earlier work this paper cites.
Representation degeneration problem in training natural language generation models
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu · 1907
Earlier work this paper cites.
Kawin Ethayarajh · 1909
Earlier work this paper cites.
What do you mean, bert? assessing BERT as a distributional semantics model
Timothee Mickus, Denis Paperno, Mathieu Constant, and Kees van Deemter · 1911
Earlier work this paper cites.
Regularized discriminant analysis
Jerome H. Friedman · 1989
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia · 2001
Earlier work this paper cites.
Embedding compression with isotropic iterative quantization
Siyu Liao, Jie Chen, Yanzhi Wang, Qinru Qiu, and Bo Yuan · 2001
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Isobn: Fine-tuning BERT with isotropic batch normalization
Wenxuan Zhou, Bill Yuchen Lin, and Xiang Ren · 2005
Earlier work this paper cites.
R1-pca: Rotational invariant l1-norm principal component analysis for robust subspace factorization
Chris Ding, Ding Zhou, Xiaofeng He, and Hongyuan Zha · 2006
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Random walks on context spaces: Towards an explanation of the mysteries of semantic word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2015
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
Estimating the intrinsic dimension of datasets by a minimal neighborhood information
Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio · 2017
Cited alongside, same era.
All-but-the-top: Simple and effective postprocessing for word representations
Jiaqi Mu, Suma Bhat, and Pramod Viswanath · 2017
Cited alongside, same era.
Classification and geometry of general perceptual manifolds
SueYeon Chung, Daniel D. Lee, and Haim Sompolinsky · 2018
Cited alongside, same era.
Revisiting representation degeneration problem in language modeling
Zhong Zhang, Chongming Gao, Cong Xu, Rui Miao, Qinli Yang, and Junming Shao · 2020
Later among the works it cites.
Low anisotropy sense retrofitting (LASeR) : Towards isotropic and sense enriched representations
Geetanjali Bihani and Julia Rayz · 2021
Later among the works it cites.
Too much in common: Shifting of embeddings in transformer language models and its implications
Daniel Biś, Maksim Podkorytov, and Xiuwen Liu · 2021
Later among the works it cites.
Isotropy in the contextual embedding space: Clusters and manifolds
Xingyu Cai, Jiaji Huang, Yuchen Bian, and Kenneth Church · 2021
Later among the works it cites.
BERT busters: Outlier dimensions that disrupt transformers
Olga Kovaleva, Saurabh Kulshreshtha, Anna Rogers, and Anna Rumshisky · 2021
Later among the works it cites.
Learning to remove: Towards isotropic pre-trained BERT embedding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decorrelated batch normalization
Lei Huang, Dawei Yang, Bo Lang, and Jia Deng · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Cited alongside, same era.
Intrinsic Dimension of Data Representations in Deep Neural Networks
Alessio Ansuini, Alessandro Laio, Jakob H. Macke, and Davide Zoccolan · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations, 2020
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2020
Cited alongside, same era.
Yuxin Liang, Rui Cao, Jie Zheng, Jie Ren, and Ling Gao · 2021
Later among the works it cites.
A cluster-based approach for improving isotropy in contextual embedding space
Sara Rajaee and Mohammad Taher Pilehvar · 2021
Later among the works it cites.
William Timkey and Marten van Schijndel · 2021
Later among the works it cites.
IsoScore: Measuring the uniformity of embedding space utilization
William Rudman, Nate Gillman, Taylor Rayne, and Carsten Eickhoff · 2022
Later among the works it cites.
Effect of post-processing on contextualized word representations
Hassan Sajjad, Firoj Alam, Fahim Dalvi, and Nadir Durrani · 2022
Later among the works it cites.
Isotropy, clusters, and classifiers, 2024
Timothee Mickus, Stig-Arne Grönroos, and Joseph Attieh · 2024
Closest in time.