Stance detection: A survey
Dilek Küçük and Fazli Can · 2021
Later among the works it cites.
Generative interventions for causal learning
Chengzhi Mao, Augustine Cha, Amogh Gupta, Hao Wang, Junfeng Yang, and Carl Vondrick · 2021
Later among the works it cites.
Are vqa systems rad? measuring robustness to augmented data with focused interventions
Daniel Rosenberg, Itai Gat, Amir Feder, and Roi Reichart · 2021
Later among the works it cites.
Tailor: Generating and perturbing text with semantic controls
Original
Alexis Ross, Tongshuang Wu, Hao Peng, Matthew E Peters, and Matt Gardner · 2021
Later among the works it cites.
Understanding the behaviour of contrastive loss
Feng Wang and Huaping Liu · 2021
Later among the works it cites.
Measuring association between labels and free-text rationales
Sarah Wiegreffe, Ana Marasovic, and Noah A. Smith · 2021
Later among the works it cites.
Polyjuice: Automated, general-purpose counterfactual generation
Original
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld · 2021
Later among the works it cites.
Cebab: Estimating the causal effects of real-world concepts on NLP model behavior
Eldar David Abraham, Karel D’Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu · 2022
Later among the works it cites.
PADA: example-based prompt learning for on-the-fly adaptation to unseen domains
Eyal Ben-David, Nadav Oved, and Roi Reichart · 2022
Later among the works it cites.
Docogen: Domain counterfactual generation for low resource domain adaptation
Nitay Calderon, Eyal Ben-David, Amir Feder, and Roi Reichart · 2022
Later among the works it cites.
A crash course in good and bad controls
Carlos Cinelli, Andrew Forney, and Judea Pearl · 2022
Later among the works it cites.
Framework for evaluating faithfulness of local explanations
Sanjoy Dasgupta, Nave Frost, and Michal Moshkovitz · 2022
Later among the works it cites.
Domain adaptation for stance detection towards unseen target on social media
Ruofan Deng, Li Pan, and Chloé Clavel · 2022
Later among the works it cites.
A functional information perspective on model interpretation
Itai Gat, Nitay Calderon, Roi Reichart, and Tamir Hazan · 2022
Later among the works it cites.
A survey on stance detection for mis- and disinformation identification
Momchil Hardalov, Arnav Arora, Preslav Nakov, and Isabelle Augenstein · 2022
Later among the works it cites.
Rethinking attention-model explainability through faithfulness violation test
Yibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong, Jing Li, and Shiqi Wang · 2022
Later among the works it cites.
Towards faithful model explanation in NLP: A survey
Original
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Later among the works it cites.
Evaluating explanations: How much do explanations from the teacher aid students?
Danish Pruthi, Rachit Bansal, Bhuwan Dhingra, Livio Baldini Soares, Michael Collins, Zachary C. Lipton, Graham Neubig, and William W. Cohen · 2022
Later among the works it cites.
On the sensitivity and stability of model interpretations in NLP
Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang · 2022
Later among the works it cites.
A systematic study of knowledge distillation for natural language generation with pseudo-target training
Nitay Calderon, Subhabrata Mukherjee, Roi Reichart, and Amir Kantor · 2023
Closest in time.
Causal reasoning and large language models: Opening a new frontier for causality
Original
Emre Kiciman, Robert Ness, Amit Sharma, and Chenhao Tan · 2023
Closest in time.
Logical satisfiability of counterfactuals for faithful explanations in NLI
Suzanna Sia, Anton Belyy, Amjad Almahairi, Madian Khabsa, Luke Zettlemoyer, and Lambert Mathias · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Closest in time.
Investigating opinions on public policies in digital media: Setting up a supervised machine learning tool for stance classification
Christina Viehmann, Tilman Beck, Marcus Maurer, Oliver Quiring, and Iryna Gurevych · 2023
Closest in time.
Causal proxy models for concept-based model explanations
Zhengxuan Wu, Karel D’Oosterlinck, Atticus Geiger, Amir Zur, and Christopher Potts · 2023
Closest in time.
Implicit counterfactual data augmentation for deep neural networks
Original
Xiaoling Zhou and Ou Wu · 2023
Closest in time.
On the validity of covariate adjustment for estimating causal effects
Ilya Shpitser, Tyler J. VanderWeele, and James M. Robins · 2078
Closest in time.