Probing for incremental parse states in autoregressive language models
Tiwalayo Eisape, Vineet Gangireddy, Roger Levy, and Yoon Kim. 2022 · 2022
Later among the works it cites.
Negation, coordination, and quantifiers in contextualized language models
Aikaterini-Lida Kalouli, Rita Sevastjanova, Christin Beck, and Maribel Romero. 2022 · 2022
Later among the works it cites.
SHAP-Based explanation methods: A review for NLP interpretability
Edoardo Mosca, Ferenc Szigeti, Stella Tragianni, Daniel Gallagher, and Georg Groh. 2022 · 2022
Later among the works it cites.
Beware the Rationalization Trap! When Language Model Explainability Diverges from our Mental Models of Language
Rita Sevastjanova and Mennatallah El-Assady. 2022 · 2022
Later among the works it cites.
Universal stanford dependencies: A cross-linguistic typology
Marie-Catherine de Marneffe, Timothy Dozat, Natalia Silveira, Katri Haverinen, Filip Ginter, Joakim Nivre, and Christopher D Manning · 2023
Later among the works it cites.
Mistral 7b
Original
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Later among the works it cites.
Using captum to explain generative language models
Original
Vivek Miglani, Aobo Yang, Aram H Markosyan, Diego Garcia-Olano, and Narine Kokhlikyan. 2023 · 2023
Later among the works it cites.
Language models are not naysayers: an analysis of language models on negation benchmarks
Thinh Hung Truong, Timothy Baldwin, Karin Verspoor, and Trevor Cohn. 2023 · 2023
Later among the works it cites.
Explainability for large language models: A survey
Original
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2023 · 2023
Later among the works it cites.