Fetching the paper…
Reading the bibliography…
An Artificial Intelligence (AI) agent is a software entity that autonomously performs tasks or makes decisions based on pre-defined objectives and data inputs.
A robust layered control system for a mobile robot
Rodney Brooks. 1986 · 1986
Earlier work this paper cites.
An intelligent system for document retrieval in distributed office environments
Uttam Mukhopadhyay, Larry M Stephens, Michael N Huhns, and Ronald D Bonnell. 1986 · 1986
Earlier work this paper cites.
Intelligent agents: Theory and practice
Michael Wooldridge and Nicholas R Jennings. 1995 · 1995
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Weights and Biases
Chris Van Pelt Lukas Biewald. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Security and privacy for innovative automotive applications: A survey
Jerry den Hartog, Nicola Zannone, et al · 2018
Earlier work this paper cites.
Accelerating the machine learning lifecycle with MLflow
Matei Zaharia, Andrew Chen, Aaron Davidson, Ali Ghodsi, Sue Ann Hong, Andy Konwinski, Siddharth Murching, Tomas Nykodym, Paul Ogilvie, Mani Parkhe, et al · 2018
Earlier work this paper cites.
Identifying and Reducing Gender Bias in Word-Level Language Models. In Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop
Shikha Bordia and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
A framework for understanding unintended consequences of machine learning
Harini Suresh and John V Guttag. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Februus: Input purification defense against trojan attacks on deep neural network systems. In ACSAC . 897–912
Bao Gia Doan, Ehsan Abbasnejad, and Damith C Ranasinghe. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Findings of EMNLP
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Poly-encoders: Architectures and pre-training strategies for fast and accurate multi-sentence scoring. arXiv. In ICLR
S Humeau, K Shuster, M Lachaux, and J Weston. 2020 · 2020
Earlier work this paper cites.
Universal litmus patterns: Revealing backdoor attacks in cnns. In CVPR . 301–310
Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash, and Heiko Hoffmann. 2020 · 2020
Earlier work this paper cites.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks. In NeurIPS
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
{ \{ TextShield } \} : Robust text classification based on multimodal embedding and neural machine translation. In USENIX Security
Jinfeng Li, Tianyu Du, Shouling Ji, Rong Zhang, Quan Lu, Min Yang, and Ting Wang. 2020 · 2020
Earlier work this paper cites.
Spatiotemporal attacks for embodied agents. In ECCV
Aishan Liu, Tairan Huang, Xianglong Liu, Yitao Xu, Yuqing Ma, Xinyun Chen, Stephen J Maybank, and Dacheng Tao. 2020 · 2020
Earlier work this paper cites.
A survey of IoT security based on a layered architecture of sensing and data analysis
Hichem Mrabet, Sana Belguith, Adeeb Alhomoud, and Abderrazak Jemai. 2020 · 2020
Earlier work this paper cites.
Information leakage in embedding models. In ACM SIGSAC conference on computer and communications security . 377–390
Congzheng Song and Ananth Raghunathan. 2020 · 2020
Earlier work this paper cites.
Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding. In EMNLP
Nouha Dziri, Andrea Madotto, Osmar Zaïane, and Avishek Joey Bose. 2021 · 2021
Earlier work this paper cites.
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume . Association for Computational Linguistics
Gautier Izacard and Edouard Grave. 2021 · 2021
Earlier work this paper cites.
Probing toxic content in large pre-trained language models. In International Joint Conference on Natural Language Processing
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit-Yan Yeung. 2021 · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training Gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Earlier work this paper cites.
Retrieval Augmentation Reduces Hallucination in Conversation. In Findings of EMNLP
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021 · 2021
Earlier work this paper cites.
Challenges in Detoxifying Language Models. In Findings of EMNLP
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021 · 2021
Earlier work this paper cites.
On the utility of dreaming: A general model for how learning in artificial agents can benefit from data hallucination
David Windridge, Henrik Svensson, and Serge Thill. 2021 · 2021
Earlier work this paper cites.
Language models as agent models
Jacob Andreas. 2022 · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Fine-grained controllable text generation using non-residual prompting. In Annual Meeting of the Association for Computational Linguistics
Fredrik Carlsson, Joey Öhman, Fangyu Liu, Severine Verlinden, Joakim Nivre, and Magnus Sahlgren. 2022 · 2022
Earlier work this paper cites.
Quarantine: Sparsity can uncover the trojan attack trigger for free. In CVPR
Tianlong Chen, Zhenyu Zhang, Yihua Zhang, Shiyu Chang, Sijia Liu, and Zhangyang Wang. 2022 · 2022
Earlier work this paper cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
Meta Fundamental AI Research Diplomacy Team (FAIR)†, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, et al · 2022
Earlier work this paper cites.
Factuality enhanced language models for open-ended text generation
Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale N Fung, Mohammad Shoeybi, and Bryan Catanzaro. 2022 · 2022
Earlier work this paper cites.
Quark: Controllable text generation with reinforced unlearning
Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
Toxicity detection with generative prompt-based inference
Yau-Shian Wang and Yingshan Chang. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Conversational health agents: A personalized llm-powered agent framework
Mahyar Abbasian, Iman Azimi, Amir M Rahmani, and Ramesh Jain. 2023 · 2023
Earlier work this paper cites.
GPT-4 Technical Report
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
GPTCache: An open-source semantic cache for LLM applications enabling faster answers and cost savings. In Workshop for Natural Language Processing Open Source Software . 212–218
Fu Bang. 2023 · 2023
Earlier work this paper cites.
A Literature Review of Human–AI Synergy in Decision Making: From the Perspective of Affordance Actualization Theory
Ying Bao, Wankun Gong, and Kaiwen Yang. 2023 · 2023
Earlier work this paper cites.
Language model unalignment: Parametric red-teaming to expose hidden harms and biases
Rishabh Bhardwaj and Soujanya Poria. 2023 · 2023
Earlier work this paper cites.
Grounding large language models in interactive environments with online reinforcement learning. In ICML
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer. 2023 · 2023
Earlier work this paper cites.
Detection and Defense Against Prominent Attacks on Preconditioned LLM-Integrated Virtual Assistants. In IEEE Asia-Pacific Conference on Computer Science and Data Engineering (CSDE) . IEEE, 1–5
Chun Fai Chan, Daniel Wankit Yip, and Aysan Esmradi. 2023 · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023 · 2023
Earlier work this paper cites.
Gamegpt: Multi-agent collaborative framework for game development
Dake Chen, Hanbin Wang, Yunhao Huo, Yuzhao Li, and Haoyang Zhang. 2023b · 2023
Earlier work this paper cites.
Scalable multi-robot collaboration with large language models: Centralized or decentralized systems?
Yongchao Chen, Jacob Arkin, Yang Zhang, Nicholas Roy, and Chuchu Fan. 2023a · 2023
Earlier work this paper cites.
Formally specifying the high-level behavior of LLM-based agents
Maxwell Crouse, Ibrahim Abdelaziz, Kinjal Basu, Soham Dan, Sadhana Kumaravel, Achille Fokoue, Pavan Kapanipathi, and Luis Lastras. 2023 · 2023
Earlier work this paper cites.
How to jailbreak chatgpt. https://watcher.guru/news/how-to-jailbreak-chatgpt
Lavina Daryanani. 2023 · 2023
Earlier work this paper cites.
ChatGPT and the rise of large language models: the new AI-driven infodemic threat in public health
Luigi De Angelis, Francesco Baglivo, Guglielmo Arzilli, Gaetano Pierpaolo Privitera, Paolo Ferragina, Alberto Eugenio Tozzi, and Caterina Rizzo. 2023 · 2023
Earlier work this paper cites.
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2023 · 2023
Earlier work this paper cites.
Toxicity in chatgpt: Analyzing persona-assigned language models. In Findings of EMNLP
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023 · 2023
Earlier work this paper cites.
Generative AI: Practical Steps to Reduce Hallucination and Improve Performance of Systems Built with Large Language Models. In Designing with ML: How to Build Usable Machine Learning Applications. Self-published on designingwithml.com
V. Dibia. 2023 · 2023
Earlier work this paper cites.
Unleashing cheapfakes through trojan plugins of large language models
Tian Dong, Guoxing Chen, Shaofeng Li, Minhui Xue, Rayne Holland, Yan Meng, Zhen Liu, and Haojin Zhu. 2023 · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2023a · 2023
Earlier work this paper cites.
ChatGPT plugins: Data exfiltration via images & cross plugin request forgery
Embrace The Red. 2023 · 2023
Earlier work this paper cites.
A comprehensive survey of attack techniques, implementation, and mitigation strategies in large language models. In International Conference on Ubiquitous Security
Aysan Esmradi, Daniel Wankit Yip, and Chun Fai Chan. 2023 · 2023
Earlier work this paper cites.
Dissecting American Fuzzy Lop: A FuzzBench Evaluation
Andrea Fioraldi, Alessandro Mantovani, Dominik Maier, and Davide Balzarotti. 2023 · 2023
Earlier work this paper cites.
Bias and fairness in large language models: A survey
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2023 · 2023
Earlier work this paper cites.
AI pitfalls and what not to do: mitigating bias in AI
Judy Wawira Gichoya, Kaesha Thomas, Leo Anthony Celi, Nabile Safdar, Imon Banerjee, John D Banja, Laleh Seyyed-Kalantari, Hari Trivedi, and Saptarshi Purkayastha. 2023 · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In ACM Workshop on Artificial Intelligence and Security . 79–90
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Earlier work this paper cites.
Application of Large Language Models to DDoS Attack Detection. In International Conference on Security and Privacy in Cyber-Physical Systems and Smart Vehicles
Michael Guastalla, Yiyi Li, Arvin Hekmati, and Bhaskar Krishnamachari. 2023 · 2023
Earlier work this paper cites.
From chatgpt to threatgpt: Impact of generative ai in cybersecurity and privacy
Maanak Gupta, CharanKumar Akiri, Kshitiz Aryal, Eli Parker, and Lopamudra Praharaj. 2023 · 2023
Earlier work this paper cites.
Chatllm network: More brains, more intelligence
Rui Hao, Linmei Hu, Weijian Qi, Qingliu Wu, Yirui Zhang, and Liqiang Nie. 2023 · 2023
Earlier work this paper cites.
Methods for measuring, updating, and visualizing factual beliefs in language models. In Conference of the European Chapter of the Association for Computational Linguistics
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. 2023 · 2023
Earlier work this paper cites.
Memory Matters: The Need to Improve Long-Term Memory in LLM-Agents. In Proceedings of the AAAI Symposium Series
Kostas Hatalis, Despina Christou, Joshua Myers, Steven Jones, Keith Lambert, Adam Amos-Binks, Zohreh Dannenhauer, and Dustin Dannenhauer. 2023 · 2023
Earlier work this paper cites.
MISGENDERED: Limits of Large Language Models in Understanding Pronouns. In Annual Meeting of the Association for Computational Linguistics
Tamanna Hossain, Sunipa Dev, and Sameer Singh. 2023 · 2023
Cited alongside, same era.
UK cybersecurity agency warns of chatbot ‘prompt injection’ attacks — theguardian.com
https://www.theguardian.com/profile/hibaq farah. 2024 · 2023
Cited alongside, same era.
Enabling Intelligent Interactions between an Agent and an LLM: A Reinforcement Learning Approach
Bin Hu, Chenyang Zhao, Pu Zhang, Zihao Zhou, Yuanhang Yang, Zenglin Xu, and Bin Liu. 2023 · 2023
Cited alongside, same era.
A survey of safety and trustworthiness of large language models through the lens of verification and validation
Xiaowei Huang, Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, et al · 2023
Cited alongside, same era.
Boosting LLM Reasoning: Push the Limits of Few-shot Learning with Reinforced In-Context Pruning
Xijie Huang, Li Lyna Zhang, Kwang-Ting Cheng, and Mao Yang. 2023b · 2023
Automatic Evaluation of Attribution by Large Language Models. In EMNLP
Xiang Yue, Boshi Wang, Ziru Chen, Kai Zhang, Yu Su, and Huan Sun. 2023 · 2023
Later among the works it cites.
Rladapter: Bridging large language models to reinforcement learning in open worlds
Wanpeng Zhang and Zongqing Lu. 2023 · 2023
Later among the works it cites.
Siren’s song in the AI ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2023
Later among the works it cites.
Defending large language models against jailbreaking attacks through goal prioritization
Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023c · 2023
Later among the works it cites.
Competeai: Understanding the competition behaviors in large language model-based agents
Qinlin Zhao, Jindong Wang, Yixuan Zhang, Yiqiao Jin, Kaijie Zhu, Hao Chen, and Xing Xie. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trustgpt: A benchmark for trustworthy and responsible large language models
Yue Huang, Qihui Zhang, Lichao Sun, et al · 2023
Cited alongside, same era.
Preventing generation of verbatim memorization in language models gives a false sense of privacy. In International Natural Language Generation Conference
Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A Choquette-Choo, and Nicholas Carlini. 2023 · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
Identifying and Mitigating Vulnerabilities in LLM-Integrated Applications. In ICLR
Fengqing Jiang, Zhangchen Xu, Luyao Niu, Boxin Wang, Jinyuan Jia, Bo Li, and Radha Poovendran. 2023 · 2023
Cited alongside, same era.
Large language models struggle to learn long-tail knowledge. In ICML
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023 · 2023
Cited alongside, same era.
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023 · 2023
Cited alongside, same era.
AvalonBench: Evaluating LLMs Playing the Game of Avalon. In NeurIPS Workshop
Jonathan Light, Min Cai, Sheng Shen, and Ziniu Hu. 2023 · 2023
Cited alongside, same era.
Investigating the prompt leakage effect and black-box defenses for multi-turn LLM interactions
Divyansh Agarwal, Alexander R Fabbri, Philippe Laban, Shafiq Joty, Caiming Xiong, and Chien-Sheng Wu. 2024 · 2024
Closest in time.
The Future of Cognitive Strategy-enhanced Persuasive Dialogue Agents: New Perspectives and Trends
Mengqi Chen, Bin Guo, Hao Wang, Haoyu Li, Qian Zhao, Jingqi Liu, Yasan Ding, Yan Pan, and Zhiwen Yu. 2024a · 2024
Closest in time.
Benchmarking large language models on controllable generation under diversified instructions
Yihan Chen, Benfeng Xu, Quan Wang, Yi Liu, and Zhendong Mao. 2024b · 2024
Closest in time.
Combating Adversarial Attacks with Multi-Agent Debate
Steffi Chern, Zhen Fan, and Andy Liu. 2024 · 2024
Closest in time.
Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
Stav Cohen, Ron Bitton, and Ben Nassi. 2024 · 2024
Closest in time.
Risk taxonomy, mitigation, and assessment benchmarks of large language model systems
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al · 2024
Closest in time.
LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. 2024 · 2024
Closest in time.
Determinants of LLM-assisted Decision-Making
Eva Eigner and Thorsten Händler. 2024 · 2024
Closest in time.
Coercing LLMs to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein. 2024 · 2024
Closest in time.
Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning. In IEEE/ACM International Conference on Software Engineering
Mingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang, Ge Li, Zhi Jin, Xiaoguang Mao, and Xiangke Liao. 2024 · 2024
Closest in time.
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Ye Wang, Jing Jiang, and Min Lin. 2024 · 2024
Closest in time.
Build AI powered applications with confidence
Guardrails AI. 2024 · 2024
Closest in time.
Defending Against Indirect Prompt Injection Attacks With Spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. 2024 · 2024
Closest in time.
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In ICLR
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024 · 2024
Closest in time.
" My agent understands me better": Integrating Dynamic Human-like Memory Recall and Consolidation in LLM-Based Agents. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–7
Yuki Hou, Haruki Tamoto, and Homei Miyashita. 2024 · 2024
Closest in time.
TrustAgent: Towards Safe and Trustworthy LLM-based Agents through Agent Constitution
Wenyue Hua, Xianjun Yang, Zelong Li, Cheng Wei, and Yongfeng Zhang. 2024 · 2024
Closest in time.
Testing and Understanding Erroneous Planning in LLM Agents through Synthesized User Inputs
Zhenlan Ji, Daoyuan Wu, Pingchuan Ma, Zongjie Li, and Shuai Wang. 2024 · 2024
Closest in time.
C-RAG: Certified Generation Risks for Retrieval-Augmented Language Models
Mintong Kang, Nezihe Merve Gürel, Ning Yu, Dawn Song, and Bo Li. 2024 · 2024
Closest in time.
Guide Your Agent with Adaptive Multimodal Rewards
Changyeon Kim, Younggyo Seo, Hao Liu, Lisa Lee, Jinwoo Shin, Honglak Lee, and Kimin Lee. 2024 · 2024
Closest in time.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2024 · 2024
Closest in time.
Vocabulary Attack to Hijack Large Language Model Applications
Patrick Levi and Christoph P Neumann. 2024 · 2024
Closest in time.
Personal llm agents: Insights and survey about the capability, efficiency and security
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al · 2024
Closest in time.
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
Zelong Li, Wenyue Hua, Hao Wang, He Zhu, and Yongfeng Zhang. 2024a · 2024
Closest in time.
Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks
Bill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi, and Xiang Ren. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2024
Closest in time.
The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
Tula Masterman, Sandi Besen, Mason Sawtell, and Alex Chao. 2024 · 2024
Closest in time.
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback. In Findings of EACL
Nikhil Mehta, Milagro Teruel, Xin Deng, Sergio Figueroa Sanz, Ahmed Awadallah, and Julia Kiseleva. 2024 · 2024
Closest in time.
AIOS: LLM Agent Operating System
Kai Mei, Zelong Li, Shuyuan Xu, Ruosong Ye, Yingqiang Ge, and Yongfeng Zhang. 2024 · 2024
Closest in time.
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
Lingbo Mo, Zeyi Liao, Boyuan Zheng, Yu Su, Chaowei Xiao, and Huan Sun. 2024 · 2024
Closest in time.
Generating Benchmarks for Factuality Evaluation of Language Models. In Conference of the European Chapter of the Association for Computational Linguistics
Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, and Yoav Shoham. 2024 · 2024
Closest in time.
A Technological Perspective on Misuse of Available AI
Lukas Pöhler, Valentin Schrader, Alexander Ladwein, and Florian von Keller. 2024 · 2024
Closest in time.
Large language model evaluation via multi ai agents: Preliminary results
Zeeshan Rasheed, Muhammad Waseem, Kari Systä, and Pekka Abrahamsson. 2024 · 2024
Closest in time.
MAMBA: an Effective World Model Approach for Meta-Reinforcement Learning. In ICLR
Zohar Rimon, Tom Jurgenson, Orr Krupnik, Gilad Adler, and Aviv Tamar. 2024 · 2024
Closest in time.
Identifying the Risks of LM Agents with an LM-Emulated Sandbox. In ICLR
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto. 2024 · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2024 · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024 · 2024
Closest in time.
Mariogpt: Open-ended text2level generation through large language models
Shyam Sudhakaran, Miguel González-Duque, Matthias Freiberger, Claire Glanois, Elias Najarro, and Sebastian Risi. 2024 · 2024
Closest in time.
Do large language models show decision heuristics similar to humans? A case study using GPT-3.5
Gaurav Suri, Lily R Slater, Ali Ziaee, and Morgan Nguyen. 2024 · 2024
Closest in time.
True Knowledge Comes from Practice: Aligning Large Language Models with Embodied Environments via Reinforcement Learning. In ICLR
Weihao Tan, Wentao Zhang, Shanqi Liu, Longtao Zheng, Xinrun Wang, and Bo An. 2024 · 2024
Closest in time.
Prioritizing Safeguarding Over Autonomy: Risks of LLM Agents for Science
Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan, Yichi Zhang, Wangchunshu Zhou, Meng Qu, Yilun Zhao, Jian Tang, Zhuosheng Zhang, et al · 2024
Closest in time.
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game. In ICLR
Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell. 2024 · 2024
Closest in time.
Bootstrapping llm-based task-oriented dialogue agents via self-talk
Dennis Ulmer, Elman Mansimov, Kaixiang Lin, Justin Sun, Xibin Gao, and Yi Zhang. 2024 · 2024
Closest in time.
Introducing v0. 5 of the AI Safety Benchmark from MLCommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M Ahmed, Victor Akinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Hamelin, et al · 2024
Closest in time.
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024 · 2024
Closest in time.
Describe, explain, plan and select: interactive planning with LLMs enables open-world multi-task agents
Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Shawn Ma, and Yitao Liang. 2024 · 2024
Closest in time.
Long-form factuality in large language models
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, et al · 2024
Closest in time.
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
Roy Weiss, Daniel Ayzenshteyn, Guy Amit, and Yisroel Mirsky. 2024 · 2024
Closest in time.
WIPI: A New Web Threat for LLM-Driven Web Agents
Fangzhou Wu, Shutong Wu, Yulong Cao, and Chaowei Xiao. 2024a · 2024
Closest in time.
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick McDaniel, and Chaowei Xiao. 2024b · 2024
Closest in time.
Designing Heterogeneous LLM Agents for Financial Sentiment Analysis
Frank Xing. 2024 · 2024
Closest in time.
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun. 2024a · 2024
Closest in time.
PRSA: Prompt Reverse Stealing Attacks against Large Language Models
Yong Yang, Xuhong Zhang, Yi Jiang, Xi Chen, Haoyu Wang, Shouling Ji, and Zonghui Wang. 2024b · 2024
Closest in time.
Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models. In ICASSP
Dongyu Yao, Jianshu Zhang, Ian G Harris, and Marcel Carlsson. 2024 · 2024
Closest in time.
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
Shenglai Zeng, Jiankun Zhang, Pengfei He, Yue Xing, Yiding Liu, Han Xu, Jie Ren, Shuaiqiang Wang, Dawei Yin, Yi Chang, et al · 2024
Closest in time.
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu. 2024a · 2024
Closest in time.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024 · 2024
Closest in time.
Privacyasst: Safeguarding user privacy in tool-using large language model agents
Xinyu Zhang, Huiyu Xu, Zhongjie Ba, Zhibo Wang, Yuan Hong, Jian Liu, Zhan Qin, and Kui Ren. 2024d · 2024
Closest in time.
Effective Prompt Extraction from Language Models
Yiming Zhang, Nicholas Carlini, and Daphne Ippolito. 2024b · 2024
Closest in time.
Intention analysis prompting makes large language models a good jailbreak defender
Yuqi Zhang, Liang Ding, Lefei Zhang, and Dacheng Tao. 2024c · 2024
Closest in time.
A Survey on the Memory Mechanism of Large Language Model based Agents
Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. 2024a · 2024
Closest in time.
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models. In ICLR
Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao. 2024 · 2024
Closest in time.
PoisonedRAG: Knowledge Poisoning Attacks to Retrieval-Augmented Generation of Large Language Models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2024 · 2024
Closest in time.