Towards automated circuit discovery for mechanistic interpretability
Original
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Later among the works it cites.
Visual programming: Compositional visual reasoning without training
Gupta, T. and Kembhavi, A · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing
Original
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Later among the works it cites.
Rigorously assessing natural language explanations of neurons
Original
Huang, J., Geiger, A., D’Oosterlinck, K., Wu, Z., and Potts, C · 2023
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations, 2023
Kirichenko, P., Izmailov, P., and Wilson, A. G · 2023
Later among the works it cites.
Segment anything
Original
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., and Girshick, R · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Original
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., et al · 2023
Later among the works it cites.
Spawrious: A benchmark for fine control of spurious correlation biases, 2023
Lynch, A., Dovonon, G. J.-S., Kaddour, J., and Silva, R · 2023
Later among the works it cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., and Sun, M · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools, 2023
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Later among the works it cites.
Find: A function description benchmark for evaluating interpretability methods, 2023
Schwettmann, S., Shaham, T. R., Materzynska, J., Chowdhury, N., Li, S., Andreas, J., Bau, D., and Torralba, A · 2023
Later among the works it cites.
Explaining black box text modules in natural language with language models, 2023
Singh, C., Hsu, A. R., Antonello, R., Jain, S., Huth, A. G., Yu, B., and Gao, J · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning, 2023
Surís, D., Menon, S., and Vondrick, C · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models, 2023
Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N · 2023
Later among the works it cites.
Language in a bottle: Language model guided concept bottlenecks for interpretable image classification, 2023
Yang, Y., Panagopoulou, A., Zhou, S., Jin, D., Callison-Burch, C., and Yatskar, M · 2023
Later among the works it cites.
React: Synergizing reasoning and acting in language models, 2023
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2023
Later among the works it cites.
Segment everything everywhere all at once, 2023
Original
Zou, X., Yang, J., Zhang, H., Li, F., Li, L., Wang, J., Wang, L., Gao, J., and Lee, Y. J · 2023
Later among the works it cites.
Interpreting clip’s image representation via text-based decomposition, 2024
Gandelsman, Y., Efros, A. A., and Steinhardt, J · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.
Gpt-4v(ision) is a generalist web agent, if grounded, 2024
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Closest in time.