Fetching the paper…
Reading the bibliography…
Understanding what defines a good representation in large language models (LLMs) is fundamental to both theoretical understanding and practical applications.
On measures of entropy and information
Alfréd Rényi · 1961
Earlier work this paper cites.
The dip test of unimodality
John A Hartigan and Pamela M Hartigan · 1985
Earlier work this paper cites.
Measures of entropy from data using infinitely divisible kernels
Luis Gonzalo Sanchez Giraldo, Murali Rao, and Jose C Principe · 2014
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Scholkopf and Alexander J Smola · 2018
Earlier work this paper cites.
Von neumann entropy from unitarity
Paul Boes, Jens Eisert, Rodrigo Gallego, Markus P Müller, and Henrik Wilming · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F Liu, Matt Gardner, Yonatan Belinkov, Matthew E Peters, and Noah A Smith · 2019
Earlier work this paper cites.
Nlp augmentation
Edward Ma · 2019
Earlier work this paper cites.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Emergence of separable manifolds in deep language representations
Jonathan Mamou, Hang Le, Miguel A Del Rio, Cory Stephenson, Hanlin Tang, Yoon Kim, and SueYeon Chung · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Understanding neural networks with logarithm determinant entropy estimator
Zhanghao Zhouyin and Ding Liu · 2021
Cited alongside, same era.
α \alpha -ReQ: Assessing representation quality in self-supervised learning by measuring eigenspectrum decay
Kumar K Agrawal, Arnab Kumar Mondal, Arna Ghosh, and Blake Richards · 2022
Cited alongside, same era.
Information theory with kernel methods
Francis Bach · 2022
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
MTEB: Massive text embedding benchmark
DiME: Maximizing mutual information by a difference of matrix-based entropies
Oscar Skean, Jhoan Keider Hoyos Osorio, Austin J Brockmeier, and Luis Gonzalo Sanchez Giraldo · 2023
Later among the works it cites.
LLM2Vec: Large language models are secretly powerful text encoders
Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy · 2024
Closest in time.
Not all layers of llms are necessary during inference
Siqi Fan, Xin Jiang, Xiang Li, Xuying Meng, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang, and Zhongyuan Wang · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2024
Closest in time.
Exploring concept depth: How large language models acquire knowledge at different layers?
Mingyu Jin, Qinkai Yu, Jingyuan Huang, Qingcheng Zeng, Zhenting Wang, Wenyue Hua, Haiyan Zhao, Kai Mei, Yanda Meng, Kaize Ding, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2022
Cited alongside, same era.
Information flow in deep neural networks
Ravid Shwartz-Ziv · 2022
Cited alongside, same era.
Reverse engineering self-supervised learning
Ido Ben-Shaul, Ravid Shwartz-Ziv, Tomer Galanti, Shai Dekel, and Yann LeCun · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Cited alongside, same era.
Guillotine regularization: Why removing layers is needed to improve generalization in self-supervised learning
Florian Bordes, Randall Balestriero, Quentin Garrido, Adrien Bardes, and Pascal Vincent · 2023
Cited alongside, same era.
RankMe: Assessing the downstream performance of pretrained self-supervised representations by their rank
Quentin Garrido, Randall Balestriero, Laurent Najman, and Yann Lecun · 2023
Cited alongside, same era.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2023
Cited alongside, same era.
Vedang Lad, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
Bm25s: Orders of magnitude faster lexical search via eager sparse scoring, 2024
Xing Han Lù · 2024
Closest in time.
Eliciting latent knowledge from quirky language models
Alex Troy Mallen and Nora Belrose · 2024
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2024
Closest in time.
To compress or not to compress—self-supervised learning and information theory: A review
Ravid Shwartz Ziv and Yann LeCun · 2024
Closest in time.
FroSSL: Frobenius norm minimization for self-supervised learning
Oscar Skean, Aayush Dhakal, Nathan Jacobs, and Luis Gonzalo Sanchez Giraldo · 2024
Closest in time.
LiDAR: Sensing linear probing performance in joint embedding ssl architectures
Vimal Thilak, Chen Huang, Omid Saremi, Laurent Dinh, Hanlin Goh, Preetum Nakkiran, Joshua M Susskind, and Etai Littwin · 2024
Closest in time.
Ai medical chatbot dataset, 2024
Ruslan Magana Vsevolodovna · 2024
Closest in time.
Large language model evaluation via matrix entropy
Lai Wei, Zhiquan Tan, Chenghai Li, Jindong Wang, and Weiran Huang · 2024
Closest in time.