Fetching the paper…
Reading the bibliography…
We introduce the Concept Bottleneck Large Language Model (CB-LLM), a pioneering approach to creating inherently interpretable Large Language Models (LLMs).
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia
Lehmann, J., Isele, R., Jakob, M., Jentzsch, A., Kontokostas, D., Mendes, P. N., Hellmann, S., Morsey, M., van Kleef, P., Auer, S., and Bizer, C · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J. J., and LeCun, Y · 2015
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Li, J., Chen, X., Hovy, E. H., and Jurafsky, D · 2016
Earlier work this paper cites.
Representation of linguistic form and function in recurrent neural networks
Kádár, Á., Chrupala, G., and Alishahi, A · 2017
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C. J., Wexler, J., Viégas, F. B., and Sayres, R · 2018
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep NLP models
Dalvi, F., Durrani, N., Sajjad, H., Belinkov, Y., Bau, A., and Glass, J. R · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
The emergence of number and syntax units in LSTM language models
Lakretz, Y., Kruszewski, G., Desbordes, T., Hupkes, D., Dehaene, S., and Baroni, M · 2019
Cited alongside, same era.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Concept bottleneck models
Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P · 2020
Cited alongside, same era.
On the pitfalls of analyzing individual neurons in language models
Antverg, O. and Belinkov, Y · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Language models can explain neurons in language models, 2023
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Later among the works it cites.
Are large language models post hoc explainers?
Kroeger, N., Ley, D., Krishna, S., Agarwal, C., and Lakkaraju, H · 2023
Later among the works it cites.
The importance of prompt tuning for automated neuron explanations
Lee, J., Oikarinen, T., Chatha, A., Chang, K.-C., Chen, Y., and Weng, T.-W · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mpnet: Masked and permuted pre-training for language understanding
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T · 2020
Cited alongside, same era.
SimCSE: Simple contrastive learning of sentence embeddings
Gao, T., Yao, X., and Chen, D · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Clip-dissect: Automatic description of neuron representations in deep vision networks
Oikarinen, T. P. and Weng, T · 2023
Later among the works it cites.
Label-free concept bottleneck models
Oikarinen, T. P., Das, S., Nguyen, L. M., and Weng, T · 2023
Later among the works it cites.
Post-hoc concept bottleneck models
Yüksekgönül, M., Wang, M., and Zou, J · 2023
Later among the works it cites.