Fetching the paper…
Reading the bibliography…
The model context protocol (MCP) has been widely adapted as an open standard enabling the seamless integration of generative AI agents.
Towards a common enumeration of vulnerabilities
David E Mann and Steven M Christey · 1999
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Prevention of phishing attacks using ai-based cybersecurity awareness training
Meraj Farheen Ansari, Pawan Kumar Sharma, and Bibhu Dash · 2022
Earlier work this paper cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, et al · 2023
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Earlier work this paper cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Earlier work this paper cites.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al · 2023
Earlier work this paper cites.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2023
Earlier work this paper cites.
Introducing Llama 3.1: Our most capable models to date
AI@Meta · 2024
Earlier work this paper cites.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Balcells Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda · 2024
Earlier work this paper cites.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, et al · 2024
Cited alongside, same era.
Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents
Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr · 2024
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, et al · 2024
Cited alongside, same era.
Redcode: Risky code execution and generation benchmark for code agents
Chengquan Guo, Xun Liu, Chulin Xie, Andy Zhou, Yi Zeng, Zinan Lin, Dawn Song, and Bo Li · 2024
Cited alongside, same era.
Towards efficient exact optimization of language model alignment
Haozhe Ji, Cheng Lu, Yilin Niu, Pei Ke, Hongning Wang, Jun Zhu, Jie Tang, and Minlie Huang · 2024
Llamafirewall: An open source guardrail system for building secure ai agents
Sahana Chennabasappa, Cyrus Nikolaidis, Daniel Song, David Molnar, Stephanie Ding, Shengye Wan, Spencer Whitman, Lauren Deason, Nicholas Doucette, Abraham Montilla, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Anchored preference optimization and contrastive revisions: Addressing underspecification in alignment
Karel D’Oosterlinck, Winnie Xu, Chris Develder, Thomas Demeester, Amanpreet Singh, Christopher Potts, Douwe Kiela, and Shikib Mehri · 2025
Closest in time.
Create chatbots that speak different languages with Gemini, Gemma, Translation LLM, and Model Context Protocol
Google · 2025
Closest in time.
MCP Toolbox for Databases: Simplify AI Agent Access to Enterprise Data
Google · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Binary classifier optimization for large language model alignment
Seungjae Jung, Gunsoo Han, Daniel Wontae Nam, and Kyoung-Woon On · 2024
Cited alongside, same era.
Gemma 2: Improving open language models at a practical size
Team Gemma@Google · 2024
Cited alongside, same era.
Team Qwen@Alibaba · 2024
Cited alongside, same era.
Rag llms are not safer: A safety analysis of retrieval-augmented generation for large language models
Bang An, Shiyue Zhang, and Mark Dredze · 2025
Cited alongside, same era.
Filesystem MCP Server - Node.js server implementing Model Context Protocol (MCP) for filesystem operations
Anthropic · 2025
Cited alongside, same era.
Introducing the Model Context Protocol
Anthropic · 2025
Cited alongside, same era.
MCP Quickstart For Claude Desktop Users
Anthropic · 2025
Cited alongside, same era.
Closest in time.
Mcp guardian: A security-first layer for safeguarding mcp-based ai system
Sonu Kumar, Anubhav Girdhar, Ritesh Patil, and Divyansh Tripathi · 2025
Closest in time.
MCP Security Notification: Tool Poisoning Attacks
Invariant Labs · 2025
Closest in time.
Introducing Model Context Protocol (MCP) in Copilot Studio
Microsoft · 2025
Closest in time.
OpenAI Agents SDK - Model context protocol
OpenAI · 2025
Closest in time.
Model Card for distilroberta-base-rejection-v1
ProtectAI · 2025
Closest in time.
Mcp safety audit: Llms with the model context protocol allow major security exploits
Brandon Radosevich and John Halloran · 2025
Closest in time.
How to use Anthropic MCP Server with open LLMs, OpenAI or Google Gemini
Philipp Schmid · 2025
Closest in time.
Stripe Agent Toolkit
Stripe · 2025
Closest in time.
Surgical, cheap, and flexible: Mitigating false refusal in language models via single vector ablation
Xinpeng Wang, Chengzhi Hu, Paul Röttger, and Barbara Plank · 2025
Closest in time.