Fetching the paper…
Reading the bibliography…
How do large language models (LLMs) obtain their answers? The ability to explain and control an LLM's reasoning process is key for reliability, transparency, and future model developments.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
Understanding deep image representations by inverting them
Mahendran, A. and Vedaldi, A · 2015
Earlier work this paper cites.
Generating visual explanations
Hendricks, L. A., Akata, Z., Rohrbach, M., Donahue, J., Schiele, B., and Darrell, T · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Nguyen, A. M., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J · 2016
Earlier work this paper cites.
”why should I trust you?”: Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Selvaraju, R. R., Das, A., Vedantam, R., Cogswell, M., Parikh, D., and Batra, D · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S · 2017
Earlier work this paper cites.
Plug & play generative networks: Conditional iterative generation of images in latent space
Nguyen, A., Clune, J., Bengio, Y., Dosovitskiy, A., and Yosinski, J · 2017
Earlier work this paper cites.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F. B., and Wattenberg, M · 2017
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Earlier work this paper cites.
Textworld: A learning environment for text-based games, 2019
Côté, M.-A., Ákos Kádár, Yuan, X., Kybartas, B., Barnes, T., Fine, E., Moore, J., Tao, R. Y., Hausknecht, M., Asri, L. E., Adada, M., Tay, W., and Trischler, A · 2019
Cited alongside, same era.
Quantifying attention flow in transformers
Abnar, S. and Zuidema, W · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Concept bottleneck models
Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P · 2020
Cited alongside, same era.
Modifying memories in transformer models, 2020
Zhu, C., Rawat, A. S., Zaheer, M., Bhojanapalli, S., Li, D., Yu, F., and Kumar, S · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E. H., Le, Q., and Zhou, D · 2022
Later among the works it cites.
Gpt-4 can’t reason, 2023
Arkoudas, K · 2023
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens, 2023
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J · 2023
Later among the works it cites.
Interpreting clip’s image representation via text-based decomposition, 2023
Gandelsman, Y., Efros, A. A., and Steinhardt, J · 2023
Later among the works it cites.
Studying large language model generalization with influence functions
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Implicit representations of meaning in neural language models
Li, B. Z., Nye, M., and Andreas, J · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Natural language descriptions of deep visual features
Hernandez, E., Schwettmann, S., Bau, D., Bagashvili, T., Torralba, A., and Andreas, J · 2022
Cited alongside, same era.
Doubly right object recognition: A why prompt for visual rationales
Mao, C., Teotia, R., Sundar, A., Menon, S., Yang, J., Wang, X., and Vondrick, C · 2022
Cited alongside, same era.
Inspecting and editing knowledge representations in language models
Hernandez, E., Li, B. Z., and Andreas, J · 2023
Later among the works it cites.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue, 2023
OpenAI · 2023
Later among the works it cites.
Future lens: Anticipating subsequent tokens from a single hidden state, 2023
Pal, K., Sun, J., Yuan, A., Wallace, B. C., and Bau, D · 2023
Later among the works it cites.
Code llama: Open foundation models for code, 2023
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
Surís, D., Menon, S., and Vondrick, C · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting, 2023
Turpin, M., Michael, J., Perez, E., and Bowman, S. R · 2023
Later among the works it cites.
Patchscopes: A unifying framework for inspecting hidden representations of language models, 2024
Ghandeharioun, A., Caciularu, A., Pearce, A., Dixon, L., and Geva, M · 2024
Closest in time.
Linearity of relation decoding in transformer language models
Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D · 2024
Closest in time.