Fetching the paper…
Reading the bibliography…
Recent advancements in large language model (LLM) agents have significantly accelerated scientific discovery automation, yet concurrently raised critical ethical and safety concerns.
Perceptions of chemical safety in laboratories
Walid Al-Zyoud, Alshaimaa M Qunies, Ayana UC Walters, and Nigel K Jalsa. 2019 · 2019
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, and 1 others. 2022 · 2022
Earlier work this paper cites.
Saibench: Benchmarking ai for science
Yatao Li and Jianfeng Zhan. 2022 · 2022
Earlier work this paper cites.
Backdoor learning: A survey
Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022 · 2022
Earlier work this paper cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022 · 2022
Earlier work this paper cites.
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Earlier work this paper cites.
Word embeddings are steers for language models
Chi Han, Jialiang Xu, Manling Li, Yi Fung, Chenkai Sun, Nan Jiang, Tarek Abdelzaher, and Heng Ji. 2023 · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and 1 others. 2023 · 2023
Earlier work this paper cites.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju. 2023 · 2023
Earlier work this paper cites.
Deepinception: Hypnotize large language model to be jailbreaker
Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. 2023 · 2023
Earlier work this paper cites.
Human values in multiagent systems
Nardine Osman and Mark d’Inverno. 2023 · 2023
Earlier work this paper cites.
A defense strategy for false data injection attacks in multi-agent systems
Lucheng Sun, Tiejun Wu, and Ya Zhang. 2023 · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, and 1 others. 2023 · 2023
Earlier work this paper cites.
Evil geniuses: Delving into the safety of llm-based agents
Yu Tian, Xiao Yang, Jingyuan Zhang, Yinpeng Dong, and Hang Su. 2023 · 2023
Earlier work this paper cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Earlier work this paper cites.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2023 · 2023
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Earlier work this paper cites.
Agentharm: A benchmark for measuring harmfulness of llm agents
Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrikson, and 1 others. 2024 · 2024
Earlier work this paper cites.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. 2024 · 2024
Earlier work this paper cites.
Iteralign: Iterative constitutional alignment of large language models
Xiusi Chen, Hongzhi Wen, Sreyashi Nag, Chen Luo, Qingyu Yin, Ruirui Li, Zheng Li, and Wei Wang. 2024a · 2024
Earlier work this paper cites.
Exploring large language model based intelligent agents: Definitions, methods, and prospects
Yuheng Cheng, Ceyao Zhang, Zhengwen Zhang, Xiangrui Meng, Sirui Hong, Wenhao Li, Zihao Wang, Zekai Wang, Feng Yin, Junhua Zhao, and 1 others. 2024 · 2024
Earlier work this paper cites.
Safety-aware fine-tuning of large language models
Hyeong Kyu Choi, Xuefeng Du, and Yixuan Li. 2024 · 2024
Earlier work this paper cites.
Agentdojo: A dynamic environment to evaluate attacks and defenses for llm agents
Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. 2024 · 2024
Cited alongside, same era.
Ai agents under threat: A survey of key security challenges and future pathways
Zehang Deng, Yongjian Guo, Changzhou Han, Wanlun Ma, Junwu Xiong, Sheng Wen, and Yang Xiang. 2024 · 2024
Cited alongside, same era.
Safeguarding large language models: A survey
Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, and 1 others. 2024 · 2024
Cited alongside, same era.
How far are we from AGI: Are LLMs all we need?
Tao Feng, Chuanyang Jin, Jingyu Liu, Kunlun Zhu, Haoqin Tu, Zirui Cheng, Guanyu Lin, and Jiaxuan You. 2024 · 2024
Cited alongside, same era.
Large language model based multi-agents: A survey of progress and challenges
Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang. 2024 · 2024
Later among the works it cites.
Safeagentbench: A benchmark for safe task planning of embodied llm agents
Sheng Yin, Xianghe Pang, Yuanzhuo Ding, Menglan Chen, Yutong Bi, Yichen Xiong, Wenhao Huang, Zhen Xiang, Jing Shao, and Siheng Chen. 2024 · 2024
Later among the works it cites.
Researchtown: Simulator of human research community
Haofei Yu, Zhaochen Hong, Zirui Cheng, Kunlun Zhu, Keyang Xuan, Jinwei Yao, Tao Feng, and Jiaxuan You. 2024 · 2024
Later among the works it cites.
R-judge: Benchmarking safety risk awareness for llm agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, and 1 others. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. 2024 · 2024
Cited alongside, same era.
Yifeng He, Ethan Wang, Yuyang Rong, Zifei Cheng, and Hao Chen. 2024 · 2024
Cited alongside, same era.
On the resilience of multi-agent systems with malicious agents
Jen-tse Huang, Jiaxu Zhou, Tailin Jin, Xuhui Zhou, Zixi Chen, Wenxuan Wang, Youliang Yuan, Maarten Sap, and Michael R Lyu. 2024 · 2024
Cited alongside, same era.
Flooding spread of manipulated knowledge in llm-based multi-agent communities
Tianjie Ju, Yiting Wang, Xinbei Ma, Pengzhou Cheng, Haodong Zhao, Yulong Wang, Lifeng Liu, Jian Xie, Zhuosheng Zhang, and Gongshen Liu. 2024 · 2024
Cited alongside, same era.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. 2024 · 2024
Cited alongside, same era.
Prompt infection: Llm-to-llm prompt injection within multi-agent systems
Donghyun Lee and Mo Tiwari. 2024 · 2024
Cited alongside, same era.
Rethinking jailbreaking through the lens of representation engineering
Tianlong Li, Shihan Dou, Wenhao Liu, Muling Wu, Changze Lv, Rui Zheng, Xiaoqing Zheng, and Xuanjing Huang. 2024 · 2024
Cited alongside, same era.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, and 1 others. 2024 · 2024
Cited alongside, same era.
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. 2024 · 2024
Later among the works it cites.
Haicosystem: An ecosystem for sandboxing safety risks in human-ai interactions
Xuhui Zhou, Hyunwoo Kim, Faeze Brahman, Liwei Jiang, Hao Zhu, Ximing Lu, Frank Xu, Bill Yuchen Lin, Yejin Choi, Niloofar Mireshghallah, Ronan Le Bras, and Maarten Sap. 2024 · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, J Zico Kolter, Matt Fredrikson, and Dan Hendrycks. 2024 · 2024
Later among the works it cites.
Gemini 2.5 pro
Google DeepMind. 2025 · 2025
Closest in time.
A practical memory injection attack against llm agents
Shen Dong, Shaochen Xu, Pengfei He, Yige Li, Jiliang Tang, Tianming Liu, Hui Liu, and Zhen Xiang. 2025 · 2025
Closest in time.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, and 1 others. 2025 · 2025
Closest in time.
Internal activation as the polar star for steering unsafe llm behavior
Peixuan Han, Cheng Qian, Xiusi Chen, Yuji Zhang, Denghui Zhang, and Heng Ji. 2025 · 2025
Closest in time.
Red-teaming llm multi-agent systems via communication attacks
Pengfei He, Yupin Lin, Shen Dong, Han Xu, Yue Xing, and Hui Liu. 2025 · 2025
Closest in time.
Stronger universal and transferable attacks by suppressing refusals
David Huang, Avidan Shah, Alexandre Araujo, David Wagner, and Chawin Sitawarin. 2025 · 2025
Closest in time.
Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, Yuheng Cheng, Suyuchen Wang, Xiaoqiang Wang, Yuyu Luo, Haibo Jin, Peiyan Zhang, Ollie Liu, Jiaqi Chen, Huan Zhang, and 28 others. 2025 · 2025
Closest in time.
Llm4sr: A survey on large language models for scientific research
Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du. 2025 · 2025
Closest in time.
Junyuan Mao, Fanci Meng, Yifan Duan, Miao Yu, Xiaojun Jia, Junfeng Fang, Yuxuan Liang, Kun Wang, and Qingsong Wen. 2025 · 2025
Closest in time.
Openai gpt-4.5 system card
OpenAI. 2025 · 2025
Closest in time.
Ai idea bench 2025: Ai research idea generation benchmark
Yansheng Qiu, Haoquan Zhang, Zhaopan Xu, Ming Li, Diping Song, Zheng Wang, and Kaipeng Zhang. 2025 · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum. 2025 · 2025
Closest in time.
G-safeguard: A topology-guided security lens and treatment on llm-based multi-agent systems
Shilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan, Fanci Meng, Chongye Guo, Kun Wang, and Yang Wang. 2025 · 2025
Closest in time.
Tinyscientist: A lightweight framework for building research agents
Haofei Yu, Keyang Xuan, Fenghai Li, Zijie Lei, and Jiaxuan You. 2025 · 2025
Closest in time.
Dolphin: Closed-loop open-ended auto-research through thinking, practice, and feedback
Jiakang Yuan, Xiangchao Yan, Botian Shi, Tao Chen, Wanli Ouyang, Bo Zhang, Lei Bai, Yu Qiao, and Bowen Zhou. 2025 · 2025
Closest in time.
Scientific large language models: A survey on biological & chemical domains
Qiang Zhang, Keyan Ding, Tianwen Lv, Xinda Wang, Qingyu Yin, Yiwen Zhang, Jing Yu, Yuhao Wang, Xiaotong Li, Zhuoyi Xiang, and 1 others. 2025 · 2025
Closest in time.