Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have shown impressive performance in natural language tasks, but their outputs can exhibit undesirable attributes or biases.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Causality: models, reasoning and inference . Vol. 29
Judea Pearl. 2000 · 2000
Earlier work this paper cites.
Causal inference using potential outcomes: Design, modeling, decisions
Donald B Rubin. 2005 · 2005
Earlier work this paper cites.
Linguistic regularities in continuous space word representations. In Proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: Human language technologies . 746–751
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013 · 2013
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation. In International conference on machine learning . PMLR, 1180–1189
Yaroslav Ganin and Victor Lempitsky. 2015 · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Advances in neural information processing systems . 4349–4357
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
The Book of Why
Judea Pearl and Dana Mackenzie. 2018 · 2018
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019 · 2019
Earlier work this paper cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A Smith, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
Causal inference meets machine learning. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining . 3527–3528
Peng Cui, Zheyan Shen, Sheng Li, Liuyi Yao, Yaliang Li, Zhixuan Chu, and Jing Gao. 2020 · 2020
Earlier work this paper cites.
Rongzhou Bao, Jiayi Wang, and Hai Zhao. 2021 · 2021
Earlier work this paper cites.
Graph infomax adversarial learning for treatment effect estimation with networked observational data. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining . 176–184
Zhixuan Chu, Stephen L Rathbun, and Sheng Li. 2021 · 2021
Earlier work this paper cites.
Bold: Dataset and metrics for measuring biases in open-ended language generation. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency . 862–872
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Earlier work this paper cites.
An empirical survey of the effectiveness of debiasing techniques for pre-trained language models
Nicholas Meade, Elinor Poole-Dayan, and Siva Reddy. 2021 · 2021
Earlier work this paper cites.
A survey on causal inference
Liuyi Yao, Zhixuan Chu, Sheng Li, Yaliang Li, Jing Gao, and Aidong Zhang. 2021 · 2021
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022 · 2022
Earlier work this paper cites.
Can prompt probe pretrained language models? understanding the invisible risks from a causal view
Boxi Cao, Hongyu Lin, Xianpei Han, Fangchao Liu, and Le Sun. 2022 · 2022
Earlier work this paper cites.
Multi-task adversarial learning for treatment effect estimation in basket trials. In Conference on health, inference, and learning . PMLR, 79–91
Zhixuan Chu, Stephen L Rathbun, and Sheng Li. 2022 · 2022
Earlier work this paper cites.
Word embeddings via causal inference: Gender bias reducing and semantic information preserving. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 11864–11872
Lei Ding, Dengdeng Yu, Jinhan Xie, Wenxing Guo, Shenggang Hu, Meichen Liu, Linglong Kong, Hongsheng Dai, Yanchun Bao, and Bei Jiang. 2022 · 2022
Earlier work this paper cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Earlier work this paper cites.
VADER-Sentiment-Analysis
C.J. Hutto. 2022 · 2022
Earlier work this paper cites.
Incorporating Causal Analysis into Diversified and Logical Response Generation. In Proceedings of the 29th International Conference on Computational Linguistics . 378–388
Jiayi Liu, Wei Wei, Zhixuan Chu, Xing Gao, Ji Zhang, Tan Yan, and Yulin Kang. 2022 · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Extracting latent steering vectors from pretrained language models
Nishant Subramani, Nivedita Suresh, and Matthew E Peters. 2022 · 2022
Cited alongside, same era.
Rock: Causal inference principles for reasoning about commonsense causality. In International Conference on Machine Learning . PMLR, 26750–26771
Jiayao Zhang, Hongming Zhang, Weijie Su, and Dan Roth. 2022 · 2022
Cited alongside, same era.
The internal state of an llm knows when its lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Answering causal questions with augmented llms
Nick Pawlowski, James Vaughan, Joel Jennings, and Cheng Zhang. 2023 · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
Jailbreaking Black Box Large Language Models in Twenty Queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Data-centric financial large language models
Zhixuan Chu, Huaiyu Guo, Xinyuan Zhou, Yijia Wang, Fei Yu, Hong Chen, Wanqing Xu, Xin Lu, Qing Cui, Longfei Li, et al · 2023
Cited alongside, same era.
Causal Effect Estimation: Recent Progress, Challenges, and Opportunities
Zhixuan Chu and Sheng Li. 2023 · 2023
Cited alongside, same era.
Is chatgpt a good causal reasoner? a comprehensive evaluation
Jinglong Gao, Xiao Ding, Bing Qin, and Ting Liu. 2023 · 2023
Cited alongside, same era.
Language models represent space and time
Wes Gurnee and Max Tegmark. 2023 · 2023
Cited alongside, same era.
Activation Addition: Steering Language Models Without Optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023 · 2023
Later among the works it cites.
Haoran Wang and Kai Shu. 2023 · 2023
Later among the works it cites.
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Zeming Wei, Yifei Wang, and Yisen Wang. 2023 · 2023
Later among the works it cites.
Unveiling the implicit toxicity in large language models
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang. 2023 · 2023
Later among the works it cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher D Manning, Christopher Potts, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023b · 2023
Later among the works it cites.
INSIDE: LLMs’ Internal States Retain the Power of Hallucination Detection
Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. 2024b · 2024
Closest in time.
Learning a structural causal model for intuition reasoning in conversation
Hang Chen, Bingyu Liao, Jing Luo, Wenjing Zhu, and Xinyu Yang. 2024a · 2024
Closest in time.
Causal Interventional Prediction System for Robust and Explainable Effect Forecasting
Zhixuan Chu, Hui Ding, Guang Zeng, Shiyu Wang, and Yiming Li. 2024a · 2024
Closest in time.
Llm-guided multi-view hypergraph learning for human-centric explainable recommendation
Zhixuan Chu, Yan Wang, Qing Cui, Longfei Li, Wenqing Chen, Sheng Li, Zhan Qin, and Kui Ren. 2024b · 2024
Closest in time.
Zhixuan Chu, Yan Wang, Feng Zhu, Lu Yu, Longfei Li, and Jinjie Gu. 2024c · 2024
Closest in time.
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
Zhixuan Chu, Lei Zhang, Yichen Sun, Siqiao Xue, Zhibo Wang, Zhan Qin, and Kui Ren. 2024d · 2024
Closest in time.
Intelligent Agents with LLM-based Process Automation. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 5018–5027
Yanchu Guan, Dong Wang, Zhixuan Chu, Shiyu Wang, Feiyue Ni, Ruihua Song, and Chenyi Zhuang. 2024 · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2024 · 2024
Closest in time.
Lei Liu, Xiaoyan Yang, Junchi Lei, Xiaoyang Liu, Yue Shen, Zhiqiang Zhang, Peng Wei, Jinjie Gu, Zhixuan Chu, Zhan Qin, et al · 2024
Closest in time.
Large Language Models for Data Annotation: A Survey
Zhen Tan, Alimohammad Beigi, Song Wang, Ruocheng Guo, Amrita Bhattacharjee, Bohan Jiang, Mansooreh Karami, Jundong Li, Lu Cheng, and Huan Liu. 2024 · 2024
Closest in time.
Deception and Manipulation in Generative AI
Christian Tarsney. 2024 · 2024
Closest in time.
Guangya Wan, Yuqi Wu, Mengxuan Hu, Zhixuan Chu, and Sheng Li. 2024 · 2024
Closest in time.
A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and More
Zhichao Wang, Bin Bi, Shiva Kumar Pentyala, Kiran Ramnath, Sougata Chaudhuri, Shubham Mehrotra, Xiang-Bo Mao, Sitaram Asur, et al · 2024
Closest in time.
Explainability for large language models: A survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024 · 2024
Closest in time.