Fetching the paper…
Reading the bibliography…
In this work, we explore LLM's internal representation space to identify attention heads that contain the most truthful and accurate information.
2018
Earlier work this paper cites.
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal, “Can a suit of armor conduct electricity? a new dataset for open book question answering,” in EMNLP , 2018
2018
Earlier work this paper cites.
A. Gokaslan, V. Cohen, E. Pavlick, and S. Tellex, “Openwebtext corpus,” http://Skylion007.github.io/OpenWebTextCorpus , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” Proceedings of the International Conference on Learning Representations (ICLR) , 2021
2021
Earlier work this paper cites.
T. Hartvigsen, S. Gabriel, H. Palangi, M. Sap, D. Ray, and E. Kamar, “Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection,” 2022
2022
Earlier work this paper cites.
A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. Bowman, “BBQ: A hand-built bias benchmark for question answering,” in Findings of the Association for Computational Linguistics: ACL 2022 , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 2086–2105. [Online]. Available: https://aclanthology.org/2022.findings-acl.165
2022
Earlier work this paper cites.
S. Lin, J. Hilton, and O. Evans, “Truthfulqa: Measuring how models mimic human falsehoods,” 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, S. Johnston, S. El-Showk, A. Jones, N. Elhage, T. Hume, A. Chen, Y. Bai, S. Bowman, S. Fort, D. Ganguli, D. Hernandez, J. Jacobson, J. Kernion, S. Kravec, L. Lovitt, K. Ndousse, C. Olsson, S. Ringer, D. Amodei, T. Brown, J. Clark, N. Joseph, B. Mann, S. McCandlish, C. Olah, and J. Kaplan, “Language models (mostly) know what they know,” 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
K.-C. Yeh, J.-A. Chi, D.-C. Lian, and S.-K. Hsieh, “Evaluating interfaced LLM bias,” in Proceedings of the 35th Conference on Computational Linguistics and Speech Processing (ROCLING 2023) , 2023, pp. 292–299
2023
Cited alongside, same era.
A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.-K. Dombrowski, S. Goel, N. Li, M. J. Byun, Z. Wang, A. Mallen, S. Basart, S. Koyejo, D. Song, M. Fredrikson, J. Z. Kolter, and D. Hendrycks, “Representation engineering: A top-down approach to ai transparency,” 2023
2023
Cited alongside, same era.
K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg, “Inference-time intervention: Eliciting truthful answers from a language model,” 2023
2023
Cited alongside, same era.
J. Hościłowicz, M. Sowański, P. Czubowski, and A. Janicki, “Can we use probing to better understand fine-tuning and knowledge distillation of the BERT nlu?” in Proceedings of the 15th International Conference on Agents and Artificial Intelligence (ICAART), Volume 3, Lisbon, Portugal, February 22-24, 2023 , A. P. Rocha, L. Steels, and H. J. van den Herik, Eds. SCITEPRESS, 2023, pp. 625–632. [Online]. Available: https://doi.org/10.5220/0011724900003393
2023
Later among the works it cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023
2023
Later among the works it cites.
2024
Closest in time.