Fetching the paper…
Reading the bibliography…
As language models (LMs) are used to build autonomous agents in real environments, ensuring their adversarial robustness becomes a critical challenge.
Evasion attacks against machine learning at test time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Šrndić, Pavel Laskov, Giorgio Giacinto, and Fabio Roli · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David A. Wagner · 2016
Earlier work this paper cites.
Adversarial examples in the physical world
Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio · 2016
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang · 2017
Earlier work this paper cites.
Adversarial machine learning at scale
Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio · 2017
Earlier work this paper cites.
Certified defenses against adversarial examples
Aditi Raghunathan, Jacob Steinhardt, and Percy Liang · 2018
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy M. Cohen, Elan Rosenfeld, and J. Zico Kolter · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Ouyang Long, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2021
Earlier work this paper cites.
Video pretraining (VPT): learning to act by watching unlabeled online videos
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, and et al · 2022
Earlier work this paper cites.
Do as I can, not as I say: Grounding language in robotic affordances
Brian Ichter, Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, and et al · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Earlier work this paper cites.
WebShop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan · 2022
Earlier work this paper cites.
Image hijacks: Adversarial images can control generative models at runtime
Luke Bailey, Euan Ong, Stuart Russell, and Scott Emmons · 2023
Earlier work this paper cites.
RT-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, and et al · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Mind2Web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samual Stevens, Boshi Wang, Huan Sun, and Yu Su · 2023
Cited alongside, same era.
How robust is Google’s Bard to adversarial image attacks?
Yinpeng Dong, Huanran Chen, Jiawei Chen, Zhengwei Fang, Xiao Yang, Yichi Zhang, Yu Tian, Hang Su, and Jun Zhu · 2023
Cited alongside, same era.
Misusing tools in large language models with visual adversarial examples
Xiaohan Fu, Zihan Wang, Shuheng Li, Rajesh K. Gupta, Niloofar Mireshghallah, Taylor Berg-Kirkpatrick, and Earlence Fernandes · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Gemini Team Google · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
WorkArena: How capable are web agents at solving common knowledge work tasks?
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, et al · 2024
Closest in time.
Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast
Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Ye Wang, Jing Jiang, and Min Lin · 2024
Closest in time.
Defending against indirect prompt injection attacks with spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman · 2024
Closest in time.
SWE-Bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan · 2024
Closest in time.
Manipulating large language models to increase product visibility
Aounon Kumar and Himabindu Lakkaraju · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Automatically auditing large language models via discrete optimization
Erik Jones, Anca D. Dragan, Aditi Raghunathan, and Jacob Steinhardt · 2023
Cited alongside, same era.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen McAleer · 2023
Cited alongside, same era.
GPT-4 technical report
OpenAI · 2023
Cited alongside, same era.
Android in the wild: A large-scale dataset for Android device control
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy P. Lillicrap · 2023
Cited alongside, same era.
On the adversarial robustness of multi-modal foundation models
Christian Schlarmann and Matthias Hein · 2023
Cited alongside, same era.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Erfan Shayegani, Yue Dong, and Nael B. Abu-Ghazaleh · 2023
Cited alongside, same era.
Images are Achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models
Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen · 2024
Closest in time.
EIA: Environmental injection attack on generalist web agents for privacy leakage
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun · 2024
Closest in time.
A trembling house of cards? mapping adversarial attacks against language agents
Lingbo Mo, Zeyi Liao, Boyuan Zheng, Yu Su, Chaowei Xiao, and Huan Sun · 2024
Closest in time.
Adversarial search engine optimization for large language models
Fredrik Nestaas, Edoardo Debenedetti, and Florian Simon Tramèr · 2024
Closest in time.
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2024
Closest in time.
Autonomous evaluation and refinement of digital agents
Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr · 2024
Closest in time.
Identifying the risks of LM agents with an LM-emulated sandbox
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J Maddison, and Tatsunori Hashimoto · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2024
Closest in time.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar H. Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
Closest in time.
OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu · 2024
Closest in time.
UFO: A UI-focused agent for windows OS interaction
Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, and Qi Zhang · 2024
Closest in time.
GPT-4V(ision) is a generalist web agent, if grounded
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su · 2024
Closest in time.
WebArena: A realistic web environment for building autonomous agents
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al · 2024
Closest in time.
PoisonedRAG: Knowledge poisoning attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2024
Closest in time.