Fetching the paper…
Reading the bibliography…
As LLMs become increasingly prevalent across various applications, it is critical to establish safety guardrails to moderate input/output content of LLMs.
A logical approach to factoring belief networks
Adnan Darwiche · 2002
Earlier work this paper cites.
A differential approach to inference in bayesian networks
Adnan Darwiche · 2003
Earlier work this paper cites.
Markov logic networks
Matthew Richardson and Pedro Domingos · 2006
Earlier work this paper cites.
A tutorial on spectral clustering
Ulrike Von Luxburg · 2007
Earlier work this paper cites.
Ensemble machine learning: methods and applications
Cha Zhang and Yunqian Ma · 2012
Earlier work this paper cites.
Probabilistic sentential decision diagrams
Doga Kisa, Guy Van den Broeck, Arthur Choi, and Adnan Darwiche · 2014
Earlier work this paper cites.
Learning sum-product networks with direct and indirect variable interactions
Amirmohammad Rooshenas and Daniel Lowd · 2014
Earlier work this paper cites.
Logic tensor networks: Deep learning and logical reasoning from data and knowledge
Luciano Serafini and Artur d’Avila Garcez · 2016
Earlier work this paper cites.
On relaxing determinism in arithmetic circuits
Arthur Choi and Adnan Darwiche · 2017
Earlier work this paper cites.
Deepproblog: Neural probabilistic logic programming
Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, and Luc De Raedt · 2018
Earlier work this paper cites.
Concept learning with energy-based models
Igor Mordatch · 2018
Earlier work this paper cites.
Honghua Dong, Jiayuan Mao, Tian Lin, Chong Wang, Lihong Li, and Denny Zhou · 2019
Earlier work this paper cites.
Challenges in automated debiasing for toxic language detection
Xuhui Zhou · 2020
Earlier work this paper cites.
An energy-based model for neuro-symbolic reasoning on knowledge graphs
Dominik Dold and Josep Soler Garrido · 2021
Earlier work this paper cites.
Knowledge-enhanced machine learning pipeline against diverse adversarial attacks
Nezihe Merve Gürel, Xiangyu Qi, Luka Rimanic, Ce Zhang, and Bo Li · 2021
Earlier work this paper cites.
Bert-beta: A proactive probabilistic approach to text moderation
Fei Tan, Yifan Hu, Kevin Yen, and Changwei Hu · 2021
Earlier work this paper cites.
Logic tensor networks
Samy Badreddine, Artur d’Avila Garcez, Luciano Serafini, and Michael Spranger · 2022
Earlier work this paper cites.
Tractable boolean and arithmetic circuits
P Hitzler and MK Sarker · 2022
Earlier work this paper cites.
A new generation of perspective api: Efficient multilingual character-level transformers
Alyssa Lees, Vinh Q Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Towards data-and knowledge-driven artificial intelligence: A survey on neuro-symbolic computing
Wenguan Wang, Yi Yang, and Fei Wu · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Zeroc: A neuro-symbolic model for zero-shot concept recognition and acquisition at inference time
Tailin Wu, Megan Tjandrasuwita, Zhengxuan Wu, Xuelin Yang, Kevin Liu, Rok Sosic, and Jure Leskovec · 2022
Cited alongside, same era.
Improving certified robustness via statistical learning with logical reasoning
Zhuolin Yang, Zhikuan Zhao, Boxin Wang, Jiawei Zhang, Linyi Li, Hengzhi Pei, Bojan Karlaš, Ji Liu, Heng Guo, Ce Zhang, and Bo Li · 2022
Cited alongside, same era.
Like a good nearest neighbor: Practical content moderation with sentence transformers
Luke Bates and Iryna Gurevych · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al · 2024
Closest in time.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su · 2024
Closest in time.
The eu artificial intelligence act
European Commission · 2024
Closest in time.
Aegis: Online adaptive ai content safety moderation with ensemble of llm experts
Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Cited alongside, same era.
Huan Ma, Changqing Zhang, Huazhu Fu, Peilin Zhao, and Bingzhe Wu · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2023
Cited alongside, same era.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails
Traian Rebedea, Razvan Dinu, Makesh Sreedhar, Christopher Parisien, and Jonathan Cohen · 2023
Cited alongside, same era.
Hummer: Towards limited competitive preference dataset
Li Jiang, Yusen Wu, Junwu Xiong, Jingqing Ruan, Yichuan Ding, Qingpei Guo, Zujie Wen, Jun Zhou, and Xiaotie Deng · 2024
Closest in time.
Colep: Certifiably robust learning-reasoning conformal prediction via probabilistic circuits
Mintong Kang, Nezihe Merve Gürel, Linyi Li, and Bo Li · 2024
Closest in time.
Watch your language: Investigating content moderation with large language models
Deepak Kumar, Yousef AbuHashem, and Zakir Durumeric · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao · 2024
Closest in time.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al · 2024
Closest in time.
Meta ais terms of service, 2024
Meta · 2024
Closest in time.
Openai usage policies (current), 2024
OpenAI · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Navigating the overkill in large language models
Chenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang, Xuanjing Huang, Xun Zhao, and Dahua Lin · 2024
Closest in time.
A strongreject for empty jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
Closest in time.
Rigorllm: Resilient guardrails for large language models against undesired content
Zhuowen Yuan, Zidi Xiong, Yi Zeng, Ning Yu, Ruoxi Jia, Dawn Song, and Bo Li · 2024
Closest in time.
Gpt-4v (ision) is a generalist web agent, if grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su · 2024
Closest in time.
Prompt-driven llm safeguarding via directed representation optimization
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.