Fetching the paper…
Reading the bibliography…
Whether embedding spaces use all their dimensions equally, i.e., whether they are isotropic, has been a recent subject of discussion.
A dendrite method for cluster analysis
Tadeusz Caliński and Jerzy Harabasz. 1974 · 1974
Earlier work this paper cites.
Well-separated clusters and optimal fuzzy partitions
J. C. Dunn. 1974 · 1974
Earlier work this paper cites.
A cluster separation measure
David L. Davies and Donald W. Bouldin. 1979 · 1979
Earlier work this paper cites.
Comparing partitions
Lawrence Hubert and Phipps Arabie. 1985 · 1985
Earlier work this paper cites.
Silhouettes: A graphical aid to the interpretation and validation of cluster analysis
Peter J. Rousseeuw. 1987 · 1987
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
Christiane Fellbaum. 1998 · 1998
Earlier work this paper cites.
NLTK: The natural language toolkit
Steven Bird and Edward Loper. 2004 · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
All-but-the-top: Simple and effective postprocessing for word representations
Jiaqi Mu and Pramod Viswanath. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Some new deformation formulas about variance and covariance
Yuli Zhang, Huaiyu Wu, and Lei Cheng. 2012 · 2018
Cited alongside, same era.
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Kawin Ethayarajh. 2019 · 2019
Cited alongside, same era.
Pytorch: an imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Measuring the mixing of contextual information in the transformer
Javier Ferrando, Gerard I. Gállego, and Marta R. Costa-jussà. 2022b · 2022
Later among the works it cites.
Semeval-2022 task 1: CODWOE – comparing dictionaries and word embeddings
Timothee Mickus, Kees Van Deemter, Mathieu Constant, and Denis Paperno. 2022b · 2022
Later among the works it cites.
GlobEnc: Quantifying global token attribution by incorporating the whole encoder layer in transformers
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, and Mohammad Taher Pilehvar. 2022 · 2022
Later among the works it cites.
IsoScore: Measuring the uniformity of embedding space utilization
William Rudman, Nate Gillman, Taylor Rayne, and Carsten Eickhoff. 2022 · 2022
Later among the works it cites.
Is anisotropy truly harmful? a case study on text clustering
Mira Ait-Saada and Mohamed Nadif. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Cited alongside, same era.
Isotropy in the contextual embedding space: Clusters and manifolds
Xingyu Cai, Jiaji Huang, Yuchen Bian, and Kenneth Church. 2021 · 2021
Cited alongside, same era.
Normalizing flows: An introduction and review of current methods
Ivan Kobyzev, Simon J.D. Prince, and Marcus A. Brubaker. 2021 · 2021
Cited alongside, same era.
Learning to remove: Towards isotropic pre-trained BERT embedding
Yuxin Liang, Rui Cao, Jie Zheng, Jie Ren, and Ling Gao. 2021 · 2021
Cited alongside, same era.
How does fine-tuning affect the geometry of embedding space: A case study on isotropy
Sara Rajaee and Mohammad Taher Pilehvar. 2021b · 2021
Cited alongside, same era.
All bark and no bite: Rogue dimensions in transformer language models obscure representational quality
William Timkey and Marten van Schijndel. 2021 · 2021
Cited alongside, same era.
On isotropy calibration of transformer models
Yue Ding, Karolis Martinkus, Damian Pascual, Simon Clematide, and Roger Wattenhofer. 2022 · 2022
Cited alongside, same era.
Exploring anisotropy and outliers in multilingual language models for cross-lingual semantic sentence similarity
Katharina Haemmerl, Alina Fastowski, Jindřich Libovický, and Alexander Fraser. 2023 · 2023
Later among the works it cites.
Isotropic representation can improve dense retrieval
Euna Jung, Jungwon Park, Jaekeol Choi, Sungyoon Kim, and Wonjong Rhee. 2023 · 2023
Later among the works it cites.
Why bother with geometry? on the relevance of linear decompositions of transformer embeddings
Timothee Mickus and Raúl Vázquez. 2023 · 2023
Later among the works it cites.
Token-wise decomposition of autoregressive language model hidden states for analyzing model predictions
Byung-Doh Oh and William Schuler. 2023 · 2023
Later among the works it cites.
Stable anisotropic regularization
William Rudman and Carsten Eickhoff. 2023 · 2023
Later among the works it cites.
Leverage points in modality shifts: Comparing language-only and multimodal word representations
Alexey Tikhonov, Lisa Bylinina, and Denis Paperno. 2023 · 2023
Later among the works it cites.
Local interpretation of transformer based on linear decomposition
Sen Yang, Shujian Huang, Wei Zou, Jianbing Zhang, Xinyu Dai, and Jiajun Chen. 2023 · 2023
Later among the works it cites.