Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) drive current AI breakthroughs despite very little being known about their internal representations.
The intrinsic dimensionality of signal collections
Bennett, R · 1969
Earlier work this paper cites.
The use of the area under the roc curve in the evaluation of machine learning algorithms
Bradley, A. P · 1997
Earlier work this paper cites.
Accurate minkowski sum approximation of polyhedral models
Varadhan, G. and Manocha, D · 2004
Earlier work this paper cites.
Better fine-tuning by reducing representational collapse
Aghajanyan, A., Shrivastava, A., Gupta, A., Goyal, N., Zettlemoyer, L., and Gupta, S · 2008
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Zettlemoyer, L., and Gupta, S · 2012
Earlier work this paper cites.
Intrinsic dimension estimation: Relevant techniques and a benchmark framework
Campadelli, P., Casiraghi, E., Ceruti, C., and Rozza, A · 2015
Earlier work this paper cites.
Toxic comment classification challenge, 2017
Adams, C., Jeffrey, S., Julia, E., Lucas, D., Mark, M., Nithum, and Will, C · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Balestriero, R. and Baraniuk, R. G · 2018
Earlier work this paper cites.
A spline theory of deep learning
Balestriero, R. et al · 2018
Earlier work this paper cites.
Challenges for toxic comment classification: An in-depth error analysis
Van Aken, B., Risch, J., Krestel, R., and Löser, A · 2018
Earlier work this paper cites.
The geometry of deep networks: Power diagram subdivision
Balestriero, R., Cosentino, R., Aazhang, B., and Baraniuk, R · 2019
Earlier work this paper cites.
Mad max: Affine spline insights into deep learning
Balestriero, R. and Baraniuk, R. G · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Hatebert: Retraining bert for abusive language detection in english
Caselli, T., Basile, V., Mitrović, J., and Granitzer, M · 2020
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained bert networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Ravichander, A., Belinkov, Y., and Hovy, E · 2020
Cited alongside, same era.
Graph construction from data by non-negative kernel regression
Shekkizhar, S. and Ortega, A · 2020
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al · 2021
Cited alongside, same era.
Generalizable implicit hate speech detection using contrastive learning
Kim, Y., Park, S., and Han, Y.-S · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Noci, L., Anagnostidis, S., Biggio, L., Orvieto, A., Singh, S. P., and Lucchi, A · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Later among the works it cites.
Toxicity detection with generative prompt-based inference
Wang, Y.-S. and Chang, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The low-dimensional linear geometry of contextualized word representations
Hernandez, E. and Andreas, J · 2021
Cited alongside, same era.
Understanding dimensional collapse in contrastive self-supervised learning
Jing, L., Vincent, P., LeCun, Y., and Tian, Y · 2021
Cited alongside, same era.
Hatexplain: A benchmark dataset for explainable hate speech detection
Mathew, B., Saha, P., Yimam, S. M., Biemann, C., Goyal, P., and Mukherjee, A · 2021
Cited alongside, same era.
The intrinsic dimension of images and its impact on learning
Pope, P., Zhu, C., Abdelkader, A., Goldblum, M., and Goldstein, T · 2021
Cited alongside, same era.
Challenges in automated debiasing for toxic language detection
Zhou, X · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y · 2022
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Later among the works it cites.
Transformers learn through gradual rank increase
Boix-Adsera, E., Littwin, E., Abbe, E., Bengio, S., and Susskind, J · 2023
Closest in time.
What did you learn to hate? a topic-oriented analysis of generalization in hate speech detection
Bourgeade, T., Chiril, P., Benamara, F., and Moriceau, V · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations
Chughtai, B., Chan, L., and Nanda, N · 2023
Closest in time.
Free dolly: Introducing the world’s first truly open instruction-tuned llm
Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R · 2023
Closest in time.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Closest in time.
Uncovering hidden geometry in transformers via disentangling position and context
Song, J. and Zhong, Y · 2023
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Mimetic initialization of self-attention layers
Trockman, A. and Kolter, J. Z · 2023
Closest in time.
Explainability for large language models: A survey
Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H., Wang, S., Yin, D., and Du, M · 2023
Closest in time.