Fetching the paper…
Reading the bibliography…
From extracting features to generating text, the outputs of large language models (LLMs) typically rely on the final layers, following the conventional wisdom that earlier layers capture only low-level cues.
On measures of entropy and information
Rényi, A · 1961
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Roy, O. and Vetterli, M · 2007
Earlier work this paper cites.
Measures of entropy from data using infinitely divisible kernels
Giraldo, L. G. S., Rao, M., and Principe, J. C · 2014
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Scholkopf, B. and Smola, A. J · 2018
Earlier work this paper cites.
Von neumann entropy from unitarity
Boes, P., Eisert, J., Gallego, R., Müller, M. P., and Wilming, H · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Liu, N. F., Gardner, M., Belinkov, Y., Peters, M. E., and Smith, N. A · 2019
Earlier work this paper cites.
NLP Augmentation, 2019
Ma, E · 2019
Earlier work this paper cites.
Opening the black box of deep neural networks via information
Shwartz-Ziv, R. and Tishby, N · 2019
Earlier work this paper cites.
BERT rediscovers the classical nlp pipeline
Tenney, I., Das, D., and Pavlick, E · 2019
Earlier work this paper cites.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Voita, E., Sennrich, R., and Titov, I · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
On identifiability in transformers
Brunner, G., Liu, Y., Pascual, D., Richter, O., Ciaramita, M., and Wattenhofer, R · 2020
Earlier work this paper cites.
Generative pretraining from pixels
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I · 2020
Earlier work this paper cites.
Emergence of separable manifolds in deep language representations
Mamou, J., Le, H., Del Rio, M. A., Stephenson, C., Tang, H., Kim, Y., and Chung, S · 2020
Earlier work this paper cites.
Contrastive multiview coding
Tian, Y., Krishnan, D., and Isola, P · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Understanding neural networks with logarithm determinant entropy estimator
Zhouyin, Z. and Liu, D · 2021
Earlier work this paper cites.
α \alpha -ReQ: Assessing representation quality in self-supervised learning by measuring eigenspectrum decay
Agrawal, K. K., Mondal, A. K., Ghosh, A., and Richards, B · 2022
Cited alongside, same era.
Information theory with kernel methods
Bach, F · 2022
Cited alongside, same era.
BeIT: Bert pre-training of image transformers
Bao, H., Dong, L., Piao, S., and Wei, F · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al · 2022
Cited alongside, same era.
MTEB: Massive text embedding benchmark
Muennighoff, N., Tazi, N., Magne, L., and Reimers, N · 2022
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2024
Later among the works it cites.
Training large language models to reason in a continuous latent space
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Later among the works it cites.
Exploring concept depth: How large language models acquire knowledge at different layers?
Jin, M., Yu, Q., Huang, J., Zeng, Q., Wang, Z., Hua, W., Zhao, H., Mei, K., Meng, Y., Ding, K., et al · 2024
Later among the works it cites.
The remarkable robustness of LLMs: Stages of inference?
Lad, V., Gurnee, W., and Tegmark, M · 2024
Later among the works it cites.
Eliciting latent knowledge from quirky language models
Mallen, A. T. and Belrose, N · 2024
Later among the works it cites.
Implicit regularization of deep residual networks towards neural odes
Marion, P., Wu, Y.-H., Sander, M. E., and Biau, G · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Information flow in deep neural networks
Shwartz-Ziv, R · 2022
Cited alongside, same era.
Neural representational geometry underlies few-shot concept learning
Sorscher, B., Ganguli, S., and Sompolinsky, H · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Cited alongside, same era.
Guillotine regularization: Why removing layers is needed to improve generalization in self-supervised learning
Bordes, F., Balestriero, R., Garrido, Q., Bardes, A., and Vincent, P · 2023
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2023
Cited alongside, same era.
RankMe: Assessing the downstream performance of pretrained self-supervised representations by their rank
Garrido, Q., Balestriero, R., Najman, L., and Lecun, Y · 2023
Cited alongside, same era.
Later among the works it cites.
DINOv2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al · 2024
Later among the works it cites.
The geometry of categorical and hierarchical concepts in large language models
Park, K., Choe, Y. J., Jiang, Y., and Veitch, V · 2024
Later among the works it cites.
The shape of learning: Anisotropy and intrinsic dimensions in transformer-based models
Razzhigaev, A., Mikhalchuk, M., Goncharova, E., Oseledets, I., Dimitrov, D., and Kuznetsov, A · 2024
Later among the works it cites.
FroSSL: Frobenius norm minimization for self-supervised learning
Skean, O., Dhakal, A., Jacobs, N., and Giraldo, L. G. S · 2024
Later among the works it cites.
LiDAR: Sensing linear probing performance in joint embedding ssl architectures
Thilak, V., Huang, C., Saremi, O., Dinh, L., Goh, H., Nakkiran, P., Susskind, J. M., and Littwin, E · 2024
Later among the works it cites.
Diff-eRank: A novel rank-based metric for evaluating large language models
Wei, L., Tan, Z., Li, C., Wang, J., and Huang, W · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Later among the works it cites.
Qwen2.5 technical report
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., and Qiu, Z · 2024
Later among the works it cites.
Layer by layer: Uncovering where multi-task learning happens in instruction-tuned large language models
Zhao, Z., Ziser, Y., and Cohen, S. B · 2024
Later among the works it cites.
Seq-VCR: Preventing collapse in intermediate transformer representations for enhanced reasoning
Arefin, M. R., Subbaraj, G., Gontier, N., LeCun, Y., Rish, I., Shwartz-Ziv, R., and Pal, C · 2025
Closest in time.
Why do LLMs attend to the first token?
Barbero, F., Arroyo, A., Gu, X., Perivolaropoulos, C., Bronstein, M., Veličković, P., and Pascanu, R · 2025
Closest in time.
Emergence of a high-dimensional abstraction phase in language transformers
Cheng, E., Doimo, D., Kervadec, C., Macocco, I., Yu, J., Laio, A., and Baroni, M · 2025
Closest in time.
Do language models use their depth efficiently?
Csordás, R., Manning, C. D., and Potts, C · 2025
Closest in time.
Deepseek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Multimodal autoregressive pre-training of large vision encoders
Fini, E., Shukor, M., Li, X., Dufter, P., Klein, M., Haldimann, D., Aitharaju, S., da Costa, V. G. T., Béthune, L., Gan, Z., et al · 2025
Closest in time.
When attention sink emerges in language models: An empirical view
Gu, X., Pang, T., Du, C., Liu, Q., Zhang, F., Du, C., Wang, Y., and Lin, M · 2025
Closest in time.
The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in transformer training
Saponati, M., Sager, P., Aceituno, P. V., Stadelmann, T., and Grewe, B · 2025
Closest in time.