Fetching the paper…
Reading the bibliography…
Understanding and manipulating the causal generation mechanisms in language models is essential for controlling their behavior.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna M. Wallach, Jennifer T. Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Cem Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 1901
Earlier work this paper cites.
Correlation and causation
Sewall Wright · 1921
Earlier work this paper cites.
Psychophysical analysis
Louis L Thurstone · 1927
Earlier work this paper cites.
The probability approach in econometrics
Trygve Haavelmo · 1944
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications: a series of lectures , volume 33
Emil Julius Gumbel · 1954
Earlier work this paper cites.
Individual choice behavior
R. Duncan Luce · 1959
Earlier work this paper cites.
Conditional logit analysis of qualitative choice behavior
Daniel McFadden · 1974
Earlier work this paper cites.
Thurstone’s discriminal processes fifty years later
R. Duncan Luce · 1977
Earlier work this paper cites.
The relationship between Luce’s choice axiom, Thurstone’s theory of comparative judgment, and the double exponential distribution
John I. Yellott · 1977
Earlier work this paper cites.
Probabilistic reasoning in intelligent systems - networks of plausible inference
Judea Pearl · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
Thurstone and sensory scaling: Then and now
R. Duncan Luce · 1994
Earlier work this paper cites.
Causal mediation analysis for interpreting neural NLP: The case of gender bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber · 2004
Earlier work this paper cites.
Complete identification methods for the causal hierarchy
Ilya Shpitser and Judea Pearl · 2008
Earlier work this paper cites.
Dynamic discrete choice structural models: A survey
Victor Aguirregabiria and Pedro Mira · 2009
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
Judea Pearl · 2009
Earlier work this paper cites.
Discrete Choice Methods with Simulation
Kenneth E. Train · 2009
Earlier work this paper cites.
On the partition function and random maximum a-posteriori perturbations
Tamir Hazan and Tommi Jaakkola · 2012
Earlier work this paper cites.
Generate your counterfactuals: Towards controlled counterfactual generation for text
Nishtha Madaan, Inkit Padhi, Naveen Panwar, and Diptikalyan Saha · 2012
Earlier work this paper cites.
A* sampling
Chris J. Maddison, Daniel Tarlow, and Tom Minka · 2014
Earlier work this paper cites.
Gumbel machinery
J. Chris Maddison and Daniel Tarlow · 2014
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam T. Kalai · 2016
Earlier work this paper cites.
Perturbations, Optimization, and Statistics
Tamir Hazan, George Papandreou, and Daniel Tarlow · 2016
Earlier work this paper cites.
Causal Inference in Statistics: A Primer
J. Pearl, M. Glymour, and N.P. Jewell · 2016
Earlier work this paper cites.
Parameter estimation for generalized thurstone choice models
Milan Vojnovic and Seyoung Yun · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg · 2017
Earlier work this paper cites.
Gumbel machinery, 2017
Chris J. Maddison and Daniel Tarlow · 2017
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Earlier work this paper cites.
Elements of Causal Inference: Foundations and Learning Algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schlkopf · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Cited alongside, same era.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2018
Cited alongside, same era.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Adversarial concept erasure in kernel space
Shauli Ravfogel, Francisco Vargas, Yoav Goldberg, and Ryan Cotterell · 2022
Later among the works it cites.
Interpretability in the wild: A circuit for indirect object identification in GPT-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2022
Later among the works it cites.
A measure-theoretic characterization of tight language models
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, and Ryan Cotterell · 2023
Later among the works it cites.
Counterfactual analysis in dynamic latent-state models
Martin Haugh and Raghav Singal · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The emergence of number and syntax units in LSTM language models
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni · 2019
Cited alongside, same era.
Counterfactual off-policy evaluation with Gumbel-max structural causal models
Michael Oberst and David Sontag · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Cited alongside, same era.
Reducing sentiment bias in language models via counterfactual evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli · 2020
Cited alongside, same era.
Axioms for learning from pairwise comparisons
Ritesh Noothigattu, Dominik Peters, and Ariel D Procaccia · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Cited alongside, same era.
Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, and David Bau · 2023
Later among the works it cites.
Attribution patching: Activation patching at industrial scale
Neel Nanda · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Log-linear guardedness and its implications
Shauli Ravfogel, Yoav Goldberg, and Ryan Cotterell · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery
Aaquib Syed, Can Rager, and Arthur Conmy · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Natural language counterfactuals through representation surgery
Matan Avitan, Ryan Cotterell, Yoav Goldberg, and Shauli Ravfogel · 2024
Closest in time.
On affine homotopy between language encoders
Robin SM Chan, Reda Boumasmoud, Anej Svete, Yuxin Ren, Qipeng Guo, Zhijing Jin, Shauli Ravfogel, Mrinmaya Sachan, Bernhard Schölkopf, Mennatallah El-Assady, and Ryan Cotterell · 2024
Closest in time.
Counterfactual token generation in large language models
Ivi Chatzi, Nina Corvelo Benz, Eleni Straitouri, Stratis Tsirtsis, and Manuel Gomez-Rodriguez · 2024
Closest in time.
Formal Aspects of Language Modeling
Ryan Cotterell, Anej Svete, Clara Meister, Tianyu Liu, and Li Du · 2024
Closest in time.
When is a language process a language model?
Li Du, Holden Lee, Jason Eisner, and Ryan Cotterell · 2024
Closest in time.
Causal abstraction: A theoretical foundation for mechanistic interpretability
Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, and Thomas Icard · 2024
Closest in time.
Model editing harms general abilities of large language models: Regularization to the rescue
Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, and Nanyun Peng · 2024
Closest in time.
A geometric notion of causal probing
Clément Guerner, Anej Svete, Tianyu Liu, Alexander Warstadt, and Ryan Cotterell · 2024
Closest in time.
Model editing at scale leads to gradual and catastrophic forgetting
Akshat Gupta, Anurag Rao, and Gopala Anumanchipalli · 2024
Closest in time.
AtP*: An efficient and scalable method for localizing LLM behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Inference-time intervention: eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2024
Closest in time.
Aaron Mueller · 2024
Closest in time.
Aaron Mueller, Jannik Brinkmann, Millicent Li, Samuel Marks, Koyena Pal, Nikhil Prakash, Can Rager, Aruna Sankaranarayanan, Arnab Sen Sharma, Jiuding Sun, Eric Todd, David Bau, and Yonatan Belinkov · 2024
Closest in time.
Why does new knowledge create messy ripple effects in LLMs?
Jiaxin Qin, Zixuan Zhang, Chi Han, Manling Li, Pengfei Yu, and Heng Ji · 2024
Closest in time.
Multi-property steering of large language models with dynamic activation composition
Daniel Scalena, Gabriele Sarti, and Malvina Nissim · 2024
Closest in time.
Representation surgery: Theory and practice of affine steering
Shashwat Singh, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni, Ryan Cotterell, and Ponnurangam Kumaraguru · 2024
Closest in time.