Fetching the paper…
Reading the bibliography…
The remarkable success of Large Language Models (LLMs) has illuminated a promising pathway toward achieving Artificial General Intelligence for both academic and industrial communities, owing to their unprecedented performance across various applications.
N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tramer, “Membership inference attacks from first principles,” in 2022 IEEE symposium on security and privacy (SP) . IEEE, 2022, pp. 1897–1914
1914
Earlier work this paper cites.
M. R. Genesereth and S. P. Ketchpel, “The kqml protocol: A specification of language and communication,” in Proceedings of the Third International Conference on Information and Knowledge Management (CIKM) . ACM, 1993, pp. 1–10
1993
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of artificial intelligence research , vol. 4, pp. 237–285, 1996
1996
Earlier work this paper cites.
D. S. Milojicic, M. Breugst, I. Busse, J. Campbell, S. Covaci, B. Friedman, K. Kosaka, D. B. Lange, K. Ono, M. Oshima, C. Tham, S. Virdhagriswaran, and J. White, “Masif: The omg mobile agent system interoperability facility,” in Proceedings of the Second International Workshop on Mobile Agents , ser. MA ’98. Berlin, Heidelberg: Springer-Verlag, 1998, p. 50–67
1998
Earlier work this paper cites.
F. for Intelligent Physical Agents, “Fipa communicative act library specification,” https://www.fipa.org/specs/fipa00037/SC00037J.html , 2000
2000
Earlier work this paper cites.
F. Curbera, M. Duftler, R. Khalaf, W. Nagy, N. Mukhi, and S. Weerawarana, “Web services: Why and how,” IBM Systems Journal , vol. 41, no. 2, pp. 170–177, 2002
2002
Earlier work this paper cites.
L. Panait and S. Luke, “Cooperative multi-agent learning: The state of the art,” Autonomous agents and multi-agent systems , vol. 11, pp. 387–434, 2005
2005
Earlier work this paper cites.
G. Hohpe and B. Woolf, Enterprise Integration Patterns: Designing, Building, and Deploying Messaging Solutions , ser. Addison-Wesley Signature Series (Fowler). Addison-Wesley Professional, 2006
2006
Earlier work this paper cites.
M. Macháček and O. Bojar, “Results of the wmt14 metrics shared task,” in Proceedings of the Ninth Workshop on Statistical Machine Translation , 2014, pp. 293–301
2014
Earlier work this paper cites.
X. Zhang, J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom, “Teaching machines to read and comprehend,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
R. Shokri, M. Stronati, C. Song, and V. Shmatikov, “Membership inference attacks against machine learning models,” in 2017 IEEE symposium on security and privacy (SP) . IEEE, 2017, pp. 3–18
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Caliskan, J. J. Bryson, and A. Narayanan, “Semantics derived automatically from language corpora contain human-like biases,” Science , vol. 356, no. 6334, pp. 183–186, 2017
2017
Earlier work this paper cites.
Y. Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274 , 2017
2017
Earlier work this paper cites.
B. Tran, J. Li, and A. Madry, “Spectral signatures in backdoor attacks,” in Advances in Neural Information Processing Systems . Curran Associates, Inc., 2018
2018
Earlier work this paper cites.
D. Acemoglu and P. Restrepo, “Artificial intelligence, automation, and work,” in The economics of artificial intelligence: An agenda . University of Chicago Press, 2018, pp. 197–236
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. Carlini, C. Liu, Ú. Erlingsson, J. Kos, and D. Song, “The secret sharer: Evaluating and testing unintended memorization in neural networks,” in 28th USENIX security symposium (USENIX security 19) , 2019, pp. 267–284
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
OECD, “OECD Principles on Artificial Intelligence,” https://oecd.ai/en/ai-principles , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
X. Zhang, X. Zhu, and L. Lessard, “Online data poisoning attacks,” in Learning for Dynamics and Control . PMLR, 2020, pp. 201–210
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Li, A. K. Sahu, A. Talwalkar, and V. Smith, “Federated learning: Challenges, methods, and future directions,” IEEE signal processing magazine , vol. 37, no. 3, pp. 50–60, 2020
2020
Earlier work this paper cites.
L. Li, Y. Fan, M. Tse, and K.-Y. Lin, “A review of applications in federated learning,” Computers & Industrial Engineering , vol. 149, p. 106854, 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Zhuang and D. Hadfield-Menell, “Consequences of misaligned ai,” Advances in Neural Information Processing Systems , vol. 33, pp. 15 763–15 773, 2020
2020
Earlier work this paper cites.
A. Mannes, “Governance, risk, and artificial intelligence,” Ai Magazine , vol. 41, no. 1, pp. 61–69, 2020
2020
Earlier work this paper cites.
Q. P. Nguyen, B. K. H. Low, and P. Jaillet, “Variational bayesian unlearning,” Advances in Neural Information Processing Systems , vol. 33, pp. 16 025–16 036, 2020
2020
Earlier work this paper cites.
Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and S. Y. Philip, “A comprehensive survey on graph neural networks,” IEEE transactions on neural networks and learning systems , vol. 32, no. 1, pp. 4–24, 2020
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in neural information processing systems , vol. 33, pp. 9459–9474, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson et al. , “Extracting training data from large language models,” in 30th USENIX security symposium (USENIX Security 21) , 2021, pp. 2633–2650
2021
Earlier work this paper cites.
C. Zhang, Y. Xie, H. Bai, B. Yu, W. Li, and Y. Gao, “A survey on federated learning,” Knowledge-Based Systems , vol. 216, p. 106775, 2021
2021
Earlier work this paper cites.
G. Sun, Y. Cong, J. Dong, Q. Wang, L. Lyu, and J. Liu, “Data poisoning attacks on federated machine learning,” IEEE Internet of Things Journal , vol. 9, no. 13, pp. 11 365–11 375, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang, “Challenges in detoxifying language models,” in Findings of the Association for Computational Linguistics: EMNLP 2021 , 2021, pp. 2447–2469
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. Everitt, M. Hutter, R. Kumar, and V. Krakovna, “Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective,” Synthese , vol. 198, no. Suppl 27, pp. 6435–6467, 2021
2021
Earlier work this paper cites.
UNESCO, “Recommendation on the Ethics of Artificial Intelligence,” https://unesdoc.unesco.org/ark:/48223/pf0000381137 , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Zanella-Beguelin, S. Tople, A. Paverd, and B. Köpf, “Grey-box extraction of natural language models,” in International Conference on Machine Learning . PMLR, 2021, pp. 12 278–12 286
2021
Earlier work this paper cites.
R. Panchendrarajan and S. Bhoi, “Dataset reconstruction attack against language models,” 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
B. Krause, A. D. Gotmare, B. McCann, N. S. Keskar, S. Joty, R. Socher, and N. F. Rajani, “Gedi: Generative discriminator guided sequence generation,” in Findings of the Association for Computational Linguistics: EMNLP 2021 , 2021, pp. 4929–4952
2021
Earlier work this paper cites.
A. Liu, M. Sap, X. Lu, S. Swayamdipta, C. Bhagavatula, N. A. Smith, and Y. Choi, “Dexperts: Decoding-time controlled text generation with experts and anti-experts,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics , 2021, pp. 6691–6706
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Tu, T. Wang, J. Wang, S. Manivasagam, M. Ren, and R. Urtasun, “Adversarial attacks on multi-agent communication,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 7768–7777
2021
Earlier work this paper cites.
J. Blumenkamp and A. Prorok, “The emergence of adversarial communication in multi-agent reinforcement learning,” in Conference on Robot Learning . PMLR, 2021, pp. 1394–1414
2021
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
M. Goldblum, D. Tsipras, C. Xie, X. Chen, A. Schwarzschild, D. Song, A. Mądry, B. Li, and T. Goldstein, “Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 2, pp. 1563–1580, 2022
2022
Earlier work this paper cites.
N. Kandpal, E. Wallace, and C. Raffel, “Deduplicating training data mitigates privacy risks in language models,” in International Conference on Machine Learning . PMLR, 2022, pp. 10 697–10 707
2022
Earlier work this paper cites.
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang, “Quantifying memorization across neural language models,” in The Eleventh International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Zhang, L. Lyu, W. Wang, L. Sun, and X. Sun, “How to inject backdoors with better consistency: Logit anchoring on clean data,” in International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Hu, Z. Salcic, L. Sun, G. Dobbie, P. S. Yu, and X. Zhang, “Membership inference attacks on machine learning: A survey,” ACM Computing Surveys (CSUR) , vol. 54, no. 11s, pp. 1–37, 2022
2022
Earlier work this paper cites.
J. Ye, A. Maddi, S. K. Murakonda, V. Bindschaedler, and R. Shokri, “Enhanced membership inference attacks against machine learning models,” in Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , 2022, pp. 3093–3106
2022
Earlier work this paper cites.
L. Lyu, H. Yu, X. Ma, C. Chen, L. Sun, J. Zhao, Q. Yang, and P. S. Yu, “Privacy and robustness in federated learning: Attacks and defenses,” IEEE transactions on neural networks and learning systems , vol. 35, no. 7, pp. 8726–8746, 2022
2022
Earlier work this paper cites.
Z. Zhang, A. Panda, L. Song, Y. Yang, M. Mahoney, P. Mittal, R. Kannan, and J. Gonzalez, “Neurotoxin: Durable backdoors in federated learning,” in International Conference on Machine Learning . PMLR, 2022, pp. 26 429–26 446
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. F. A. R. D. T. (FAIR)†, A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu et al. , “Human-level play in the game of diplomacy by combining language models with strategic reasoning,” Science , vol. 378, no. 6624, pp. 1067–1074, 2022
2022
Earlier work this paper cites.
J. Skalse, N. Howe, D. Krasheninnikov, and D. Krueger, “Defining and characterizing reward gaming,” Advances in Neural Information Processing Systems , vol. 35, pp. 9460–9471, 2022
2022
Earlier work this paper cites.
M. M. Maas, “Aligning ai regulation to sociotechnical change,” in The Oxford Handbook of AI Governance , 2022
2022
Earlier work this paper cites.
F. Urbina, F. Lentzos, C. Invernizzi, and S. Ekins, “Dual use of artificial-intelligence-powered drug discovery,” Nature machine intelligence , vol. 4, no. 3, pp. 189–191, 2022
2022
Earlier work this paper cites.
E. Mostaque, “Democratizing ai, stable diffusion & generative models,” https://exchange.scale.com/public/videos/emad-mostaque-stability-ai-stable-diffusion-open-source , 2022
2022
Earlier work this paper cites.
Z. Zhang, Y. Zhou, X. Zhao, T. Che, and L. Lyu, “Prompt certified machine unlearning with randomized gradient smoothing and quantization,” Advances in Neural Information Processing Systems , vol. 35, pp. 13 433–13 455, 2022
2022
Earlier work this paper cites.
K. Meng, D. Bau, A. Andonian, and Y. Belinkov, “Locating and editing factual associations in gpt,” Advances in Neural Information Processing Systems , vol. 35, pp. 17 359–17 372, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. Mitchell, C. Lin, A. Bosselut, C. D. Manning, and C. Finn, “Memory-based model editing at scale,” in International Conference on Machine Learning . PMLR, 2022, pp. 15 817–15 831
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Thudi, H. Jia, I. Shumailov, and N. Papernot, “On the necessity of auditable algorithmic definitions for machine unlearning,” in 31st USENIX security symposium (USENIX Security 22) , 2022, pp. 4007–4022
2022
Earlier work this paper cites.
A. Thudi, G. Deza, V. Chandrasekaran, and N. Papernot, “Unrolling sgd: Understanding factors influencing machine unlearning,” in 2022 IEEE 7th European Symposium on Security and Privacy (EuroS&P) . IEEE, 2022, pp. 303–319
2022
Earlier work this paper cites.
B. Liu, Q. Liu, and P. Stone, “Continual learning and private unlearning,” in Conference on Lifelong Learning Agents . PMLR, 2022, pp. 243–254
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
X. He, C. Chen, L. Lyu, and Q. Xu, “Extracted bert model leaks more information than you think!” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022 . Association for Computational Linguistics, 2022, pp. 1530–1537
2022
Earlier work this paper cites.
X. He, Q. Xu, L. Lyu, F. Wu, and C. Wang, “Protecting intellectual property of language generation apis with lexical watermark,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 10, 2022, pp. 10 758–10 766
2022
Earlier work this paper cites.
X. He, Q. Xu, Y. Zeng, L. Lyu, F. Wu, J. Li, and R. Jia, “Cater: Intellectual property protection on text generation apis via conditional watermarks,” Advances in Neural Information Processing Systems , vol. 35, pp. 5431–5445, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
F. Faal, K. Schmitt, and J. Y. Yu, “Reward modeling for mitigating toxicity in transformer-based language models,” Applied Intelligence , vol. 53, no. 7, p. 8421–8435, 2022
2022
Earlier work this paper cites.
Y. Zhang, J. Wang, and J. Sang, “Counterfactually measuring and eliminating social bias in vision-language pre-training models,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 4996–5004
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” in NeurIPS , 2022
2022
Earlier work this paper cites.
M. Hao, H. Li, H. Chen, P. Xing, G. Xu, and T. Zhang, “Iron: Private inference on transformers,” Advances in neural information processing systems , vol. 35, pp. 15 718–15 731, 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. U. Hadi, R. Qureshi, A. Shah, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili et al. , “A survey on large language models: Applications, challenges, limitations, and practical usage,” Authorea Preprints , vol. 3, 2023
2023
Earlier work this paper cites.
S. McLean, G. J. Read, J. Thompson, C. Baber, N. A. Stanton, and P. M. Salmon, “The risks associated with artificial general intelligence: A systematic review,” Journal of Experimental & Theoretical Artificial Intelligence , vol. 35, no. 5, pp. 649–663, 2023
2023
Earlier work this paper cites.
J. Ruan, Y. Chen, B. Zhang, Z. Xu, T. Bao, H. Mao, Z. Li, X. Zeng, R. Zhao et al. , “Tptu: Task planning and tool usage of large language model-based ai agents,” in NeurIPS 2023 Foundation Models for Decision Making Workshop , 2023
2023
Earlier work this paper cites.
V. Sorin, E. Klang, M. Sklair-Levy, I. Cohen, D. B. Zippel, N. Balint Lahat, E. Konen, and Y. Barash, “Large language model (chatgpt) as a support tool for breast tumor board,” NPJ Breast Cancer , vol. 9, no. 1, p. 44, 2023
2023
Earlier work this paper cites.
R. Yang, L. Song, Y. Li, S. Zhao, Y. Ge, X. Li, and Y. Shan, “Gpt4tools: Teaching large language model to use tools via self-instruction,” Advances in Neural Information Processing Systems , vol. 36, pp. 71 995–72 007, 2023
2023
Earlier work this paper cites.
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” Advances in Neural Information Processing Systems , vol. 36, pp. 68 539–68 551, 2023
2023
Earlier work this paper cites.
W. Wang, L. Dong, H. Cheng, X. Liu, X. Yan, J. Gao, and F. Wei, “Augmenting language models with long-term memory,” Advances in Neural Information Processing Systems , vol. 36, pp. 74 530–74 543, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Majumdar, K. Yadav, S. Arnaud, J. Ma, C. Chen, S. Silwal, A. Jain, V.-P. Berges, T. Wu, J. Vakil et al. , “Where are we in the search for an artificial visual cortex for embodied intelligence?” Advances in Neural Information Processing Systems , vol. 36, pp. 655–677, 2023
2023
Earlier work this paper cites.
H. Wang, J. Li, H. Wu, E. Hovy, and Y. Sun, “Pre-trained language models and their applications,” Engineering , vol. 25, pp. 51–65, 2023
2023
Earlier work this paper cites.
N. Lukas, A. Salem, R. Sim, S. Tople, L. Wutschitz, and S. Zanella-Béguelin, “Analyzing leakage of personally identifiable information in language models,” in 2023 IEEE Symposium on Security and Privacy (SP) . IEEE, 2023, pp. 346–363
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. U. Hadi, R. Qureshi, A. Shah, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili et al. , “Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects,” Authorea Preprints , vol. 1, pp. 1–26, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Sun, Y. Chen, G. Tao, C. Fang, X. Zhang, Q. Zhang, and B. Luo, “Backdooring neural code search,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics . Toronto, Canada: Association for Computational Linguistics, July 9-14 2023, pp. 9692–9708
2023
Earlier work this paper cites.
M. Pan, Y. Zeng, L. Lyu, X. Lin, and R. Jia, “ { \{ ASSET } \} : Robust backdoor data detection across a multiplicity of deep learning paradigms,” in 32nd USENIX Security Symposium (USENIX Security 23) , 2023, pp. 2725–2742
2023
Earlier work this paper cites.
X. Sun, X. Li, Y. Meng, X. Ao, L. Lyu, J. Li, and T. Zhang, “Defending against backdoor attacks in natural language generation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 4, 2023, pp. 5257–5265
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Gupta, C. Akiri, K. Aryal, E. Parker, and L. Praharaj, “From chatgpt to threatgpt: Impact of generative ai in cybersecurity and privacy,” IEEE Access , vol. 11, pp. 80 218–80 245, 2023
2023
Earlier work this paper cites.
S. Kim, S. Yun, H. Lee, M. Gubri, S. Yoon, and S. J. Oh, “Propile: C,” Advances in Neural Information Processing Systems , vol. 36, pp. 20 750–20 762, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Shu, J. Wang, C. Zhu, J. Geiping, C. Xiao, and T. Goldstein, “On the exploitability of instruction tuning,” Advances in Neural Information Processing Systems , vol. 36, pp. 61 836–61 856, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
N. Ding, Y. Qin, G. Yang, F. Wei, Z. Yang, Y. Su, S. Hu, Y. Chen, C.-M. Chan, W. Chen et al. , “Parameter-efficient fine-tuning of large-scale pre-trained language models,” Nature Machine Intelligence , vol. 5, no. 3, pp. 220–235, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Wan, E. Wallace, S. Shen, and D. Klein, “Poisoning language models during instruction tuning,” in International Conference on Machine Learning . PMLR, 2023, pp. 35 413–35 425
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Lee, S. Phatale, H. Mansoor, K. R. Lu, T. Mesnard, J. Ferret, C. Bishop, E. Hall, V. Carbune, and A. Rastogi, “Rlaif: Scaling reinforcement learning from human feedback with ai feedback,” 2023
2023
Earlier work this paper cites.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in Neural Information Processing Systems , vol. 36, pp. 53 728–53 741, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang, “Beavertails: Towards improved safety alignment of llm via a human-preference dataset,” Advances in Neural Information Processing Systems , vol. 36, pp. 24 678–24 704, 2023
2023
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” in The Twelfth International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Yu, Y. Zhuang, J. Zhang, Y. Meng, A. J. Ratner, R. Krishna, J. Shen, and C. Zhang, “Large language model as attributed training data generator: A tale of diversity and bias,” Advances in Neural Information Processing Systems , vol. 36, pp. 55 734–55 784, 2023
2023
Earlier work this paper cites.
Y. Chen, Q. Fu, Y. Yuan, Z. Wen, G. Fan, D. Liu, D. Zhang, Z. Li, and Y. Xiao, “Hallucination detection: Robustly discerning reliable answers in large language models,” in Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , 2023, pp. 245–255
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
S. Prabhumoye, M. Patwary, M. Shoeybi, and B. Catanzaro, “Adding instructions during pretraining: Effective way of controlling toxicity in language models,” in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , 2023, pp. 2636–2651
2023
Earlier work this paper cites.
T. Markov, C. Zhang, S. Agarwal, F. E. Nekoul, T. Lee, S. Adler, A. Jiang, and L. Weng, “A holistic approach to undesired content detection in the real world,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 12, 2023, pp. 15 009–15 018
2023
Earlier work this paper cites.
J. Dai, X. Pan, R. Sun, J. Ji, X. Xu, M. Liu, Y. Wang, and Y. Yang, “Safe rlhf: Safe reinforcement learning from human feedback,” in The Twelfth International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
R. Tang, J. Yuan, Y. Li, Z. Liu, R. Chen, and X. Hu, “Setting the trap: Capturing and defeating backdoor threats in plms through honeypots,” NeurIPS , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Tan, S. Shi, X. Qiu, C. Qu, Z. Qi, Y. Xu, and Y. Qi, “Self-criticism: Aligning large language models with their understanding of helpfulness, honesty, and harmlessness,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track , M. Wang and I. Zitouni, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 650–662. [Online]. Available: https://aclanthology.org/2023.emnlp-industry.62/
2023
Earlier work this paper cites.
P. Hacker, A. Engel, and M. Mauer, “Regulating chatgpt and other large generative ai models,” in Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency . Association for Computing Machinery, 2023
2023
Earlier work this paper cites.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al. , “Judging llm-as-a-judge with mt-bench and chatbot arena,” Advances in Neural Information Processing Systems , vol. 36, pp. 46 595–46 623, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Li, T. Zhang, Y. Dubois, R. Taori, I. Gulrajani, C. Guestrin, P. Liang, and T. B. Hashimoto, “Alpacaeval: An automatic evaluator of instruction-following models,” 2023
2023
Earlier work this paper cites.
OpenAI, “Moderation api,” https://platform.openai.com/docs/guides/moderation/overview , 2023
2023
Earlier work this paper cites.
H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa, “Llama guard: Llm-based input-output safeguard for human-ai conversations,” CoRR , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer et al. , “Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.” in NeurIPS , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Conover, R. Staats, A. Rane, G. Shani, K. Katz, A. Powell, A. Ross, A. Maas, and A. Zhang, “Databricks-dolly: Introducing dolly-15k, democratizing the magic of instruction following,” https://github.com/databrickslabs/dolly , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
OpenAI, “Gpt-4 technical report,” ArXiv , vol. abs/2303.08774, 2023
2023
Earlier work this paper cites.
F. Ward, F. Toni, F. Belardinelli, and T. Everitt, “Honesty is the best policy: defining and mitigating ai deception,” Advances in neural information processing systems , vol. 36, pp. 2313–2341, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
L. Schulz, N. Alon, J. Rosenschein, and P. Dayan, “Emergent deception and skepticism via theory of mind,” in First Workshop on Theory of Mind in Communicating Agents , 2023
2023
Earlier work this paper cites.
A. Pan, J. S. Chan, A. Zou, N. Li, S. Basart, T. Woodside, H. Zhang, S. Emmons, and D. Hendrycks, “Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark,” in International conference on machine learning . PMLR, 2023, pp. 26 837–26 867
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
L. Gao, J. Schulman, and J. Hilton, “Scaling laws for reward model overoptimization,” in International Conference on Machine Learning . PMLR, 2023, pp. 10 835–10 866
2023
Earlier work this paper cites.
E. Perez, S. Ringer, K. Lukosiute, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath et al. , “Discovering language model behaviors with model-written evaluations,” in Findings of the Association for Computational Linguistics: ACL 2023 , 2023, pp. 13 387–13 434
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Tallberg, E. Erman, M. Furendal, J. Geith, M. Klamberg, and M. Lundgren, “The global governance of artificial intelligence: Next steps for empirical and normative research,” International Studies Review , vol. 25, no. 3, p. viad040, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Meta, “Meta and Microsoft introduce the next generation of Llama,” https://ai.meta.com/blog/llama-2 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
P. Chavez, “An ai challenge: Balancing open and closed systems,” https://cepa.org/article/an-ai-challenge-balancing-open-and-closed-systems , 2023
2023
Earlier work this paper cites.
T. Che, Y. Zhou, Z. Zhang, L. Lyu, J. Liu, D. Yan, D. Dou, and J. Huan, “Fast federated machine unlearning with nonlinear functional theory,” in International conference on machine learning . PMLR, 2023, pp. 4241–4268
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Gandikota, J. Materzynska, J. Fiotto-Kaufman, and D. Bau, “Erasing concepts from diffusion models,” 2023 IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 2426–2436, 2023
2023
Earlier work this paper cites.
E. Zhang, K. Wang, X. Xu, Z. Wang, and H. Shi, “Forget-me-not: Learning to forget in text-to-image diffusion models,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , pp. 1755–1764, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Peng, J. Yi, F. Wu, S. Wu, B. B. Zhu, L. Lyu, B. Jiao, T. Xu, G. Sun, and X. Xie, “Are you copying my model? protecting the copyright of large language models for eaas via backdoor watermark,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 7653–7668
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security , 2023, pp. 79–90
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” Advances in Neural Information Processing Systems , vol. 36, pp. 80 079–80 110, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Diera, N. Lell, A. Garifullina, and A. Scherp, “Memorization of named entities in fine-tuned bert models,” in International Cross-Domain Conference for Machine Learning and Knowledge Extraction . Springer, 2023, pp. 258–279
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
L. Yan, Z. Zhang, G. Tao, K. Zhang, X. Chen, G. Shen, and X. Zhang, “Parafuzz: An interpretability-driven technique for detecting poisoned samples in nlp,” Advances in Neural Information Processing Systems , vol. 36, pp. 66 755–66 767, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Yi, Y. Xie, B. Zhu, K. Hines, E. Kiciman, G. Sun, X. Xie, and F. Wu, “Benchmarking and defending against indirect prompt injection attacks on large language models,” CoRR , 2023
2023
Earlier work this paper cites.
T. Rebedea, R. Dinu, M. N. Sreedhar, C. Parisien, and J. Cohen, “Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , 2023, pp. 431–445
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang et al. , “Self-refine: Iterative refinement with self-feedback,” Advances in Neural Information Processing Systems , vol. 36, pp. 46 534–46 594, 2023
2023
Earlier work this paper cites.
D. Jiang, X. Ren, and B. Y. Lin, “Llm-blender: Ensembling large language models with pairwise ranking and generative fusion,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 14 165–14 178
2023
Earlier work this paper cites.
M. Cao, M. Fatemi, J. C. Cheung, and S. Shabanian, “Systematic rectification of language models via dead-end analysis,” in The Eleventh International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
W. Wang, J.-T. Huang, W. Wu, J. Zhang, Y. Huang, S. Li, P. He, and M. R. Lyu, “Mttm: Metamorphic testing for textual content moderation software,” 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) , pp. 2387–2399, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:256826966
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Zhang, J. Yang, P. Ke, F. Mi, H. Wang, and M. Huang, “Defending large language models against jailbreaking attacks through goal prioritization,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , 2023, pp. 8865–8887
2023
Earlier work this paper cites.
Y. Xie, J. Yi, J. Shao, J. Curl, L. Lyu, Q. Chen, X. Xie, and F. Wu, “Defending chatgpt against jailbreak attack via self-reminders,” Nature Machine Intelligence , vol. 5, no. 12, pp. 1486–1496, 2023
2023
Earlier work this paper cites.
S. Utpala, S. Hooker, and P.-Y. Chen, “Locally differentially private document generation using zero shot prompting,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , 2023, pp. 8442–8457
2023
Earlier work this paper cites.
H. Duan, A. Dziedzic, N. Papernot, and F. Boenisch, “Flocks of stochastic parrots: Differentially private prompt learning for large language models,” Advances in Neural Information Processing Systems , vol. 36, pp. 76 852–76 871, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Borzunov, M. Ryabinin, A. Chumachenko, D. Baranchuk, T. Dettmers, Y. Belkada, P. Samygin, and C. A. Raffel, “Distributed inference and fine-tuning of large language models over the internet,” Advances in neural information processing systems , vol. 36, pp. 12 312–12 331, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Zhu, J. Wang, J. Zhou, Z. Wang, H. Chen, Y. Wang, L. Yang, W. Ye, Y. Zhang, N. Gong et al. , “Promptrobust: Towards evaluating the robustness of large language models on adversarial prompts,” in Proceedings of the 1st ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis , 2023, pp. 57–68
2023
Earlier work this paper cites.
G. Dong, J. Zhao, T. Hui, D. Guo, W. Wang, B. Feng, Y. Qiu, Z. Gongque, K. He, Z. Wang et al. , “Revisit input perturbation problems for llms: A unified robustness evaluation framework for noisy slot filling task,” in CCF International Conference on Natural Language Processing and Chinese Computing . Springer, 2023, pp. 682–694
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
L. Gustafson, C. Rolland, N. Ravi, Q. Duval, A. Adcock, C.-Y. Fu, M. Hall, and C. Ross, “Facet: Fairness in computer vision evaluation benchmark,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 20 370–20 382
2023
Earlier work this paper cites.
A. Seth, M. Hemani, and C. Agarwal, “Dear: Debiasing vision-language models with additive residuals,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6820–6829
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,” Advances in Neural Information Processing Systems , vol. 36, pp. 38 154–38 180, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “Camel: Communicative agents for" mind" exploration of large language model society,” Advances in Neural Information Processing Systems , vol. 36, pp. 51 991–52 008, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Chen, G. Wang, H. Guo, Y. Wang, and Q. Yan, “Understanding multi-turn toxic behaviors in open-domain chatbots,” in Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , 2023, pp. 282–296
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
N. Rahman and E. Santacana, “Beyond fair use: Legal risk evaluation for training llms on copyrighted text,” in ICML Workshop on Generative AI and Law , 2023
2023
Earlier work this paper cites.
J. Kirchenbauer, J. Geiping, Y. Wen, J. Katz, I. Miers, and T. Goldstein, “A watermark for large language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 17 061–17 084
2023
Earlier work this paper cites.
Y. Wan, W. Wang, P. He, J. Gu, H. Bai, and M. R. Lyu, “Biasasker: Measuring the bias in conversational ai system,” Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:258833296
2023
Earlier work this paper cites.
Cyberspace Administration of China, “Interim measures for the management of generative artificial intelligence services,” 2023, accessed: 2025-03-07. [Online]. Available: https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al. , “A survey on evaluation of large language models,” ACM transactions on intelligent systems and technology , vol. 15, no. 3, pp. 1–45, 2024
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
S. Sonko, A. O. Adewusi, O. C. Obi, S. Onwusinkwue, and A. Atadoga, “A critical review towards artificial general intelligence: Challenges, ethical considerations, and the path forward,” World Journal of Advanced Research and Reviews , vol. 21, no. 3, pp. 1262–1268, 2024
2024
Earlier work this paper cites.
W. Zhong, L. Guo, Q. Gao, H. Ye, and Y. Wang, “Memorybank: Enhancing large language models with long-term memory,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp. 19 724–19 731
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin et al. , “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, p. 186345, 2024
2024
Earlier work this paper cites.
Y. Yan and J. Lee, “Georeasoner: Reasoning on geospatially grounded context for natural language understanding,” in Proceedings of the 33rd ACM International Conference on Information and Knowledge Management , 2024, pp. 4163–4167
2024
Earlier work this paper cites.
M. Zhou, H. Dong, H. Song, N. Zheng, W.-H. Chen, and H. Wang, “Embodied intelligence-based perception, decision-making, and control for autonomous operations of rail transportation,” IEEE Transactions on Intelligent Vehicles , 2024
2024
Earlier work this paper cites.
C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He et al. , “A comprehensive survey on pretrained foundation models: A history from bert to chatgpt,” International Journal of Machine Learning and Cybernetics , pp. 1–65, 2024
2024
Earlier work this paper cites.
H. R. Kirk, B. Vidgen, P. Röttger, and S. A. Hale, “The benefits, risks and bounds of personalizing the alignment of large language models to individuals,” Nature Machine Intelligence , vol. 6, no. 4, pp. 383–392, 2024
2024
Earlier work this paper cites.
Z. Zhou, H. Yu, X. Zhang, R. Xu, F. Huang, and Y. Li, “How alignment and jailbreak work: Explain llm safety through intermediate hidden states,” in Findings of the Association for Computational Linguistics: EMNLP 2024 , 2024, pp. 2461–2488
2024
Earlier work this paper cites.
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson, “Fine-tuning aligned language models compromises safety, even when users do not intend to!” in ICLR , 2024. [Online]. Available: https://openreview.net/forum?id=hTEGyKf0dZ
2024
Earlier work this paper cites.
D. Halawi, A. Wei, E. Wallace, T. T. Wang, N. Haghtalab, and J. Steinhardt, “Covert malicious finetuning: Challenges in safeguarding LLM adaptation,” in Proceedings of the 41st International Conference on Machine Learning . PMLR, 2024, pp. 17 298–17 312
2024
Earlier work this paper cites.
W. Hawkins, B. Mittelstadt, and C. Russell, “The effect of fine-tuning on language model toxicity,” in Neurips Safe Generative AI Workshop 2024 , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
A. Zhou, B. Li, and H. Wang, “Robust prompt optimization for defending language models against jailbreaking attacks,” in Advances in Neural Information Processing Systems , vol. 37. Curran Associates, Inc., 2024, pp. 40 184–40 211
2024
Later among the works it cites.
Y. Mo, Y. Wang, Z. Wei, and Y. Wang, “Fight back against jailbreaking via prompt adversarial tuning,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Z. Liang, Y. Xu, Y. Hong, P. Shang, Q. Wang, Q. Fu, and K. Liu, “A survey of multimodel large language models,” in Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering , 2024, pp. 405–409
2024
Cited alongside, same era.
H. Zhao, H. Chen, F. Yang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, and M. Du, “Explainability for large language models: A survey,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 2, pp. 1–38, 2024
2024
Cited alongside, same era.
M. A. K. Raiaan, M. S. H. Mukta, K. Fatema, N. M. Fahad, S. Sakib, M. M. J. Mim, J. Ahmad, M. E. Ali, and S. Azam, “A review on large language models: Architectures, applications, taxonomies, open issues and challenges,” IEEE access , vol. 12, pp. 26 839–26 874, 2024
2024
Cited alongside, same era.
K. S. Kalyan, “A survey of gpt-3 family large language models including chatgpt and gpt-4,” Natural Language Processing Journal , vol. 6, p. 100048, 2024
2024
Cited alongside, same era.
Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang, “A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,” High-Confidence Computing , p. 100211, 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. He, S. Zannettou, Y. Shen, and Y. Zhang, “You only prompt once: On the capabilities of prompt learning on large language models to tackle toxic content,” in 2024 IEEE Symposium on Security and Privacy (SP) . IEEE, 2024, pp. 770–787
2024
Later among the works it cites.
2024
Later among the works it cites.
R. Xu, Z. Qi, and W. Xu, “Preemptive answer “attacks” on chain-of-thought reasoning,” in Findings of the Association for Computational Linguistics ACL 2024 , 2024, pp. 14 708–14 726
2024
Later among the works it cites.
C. Zheng, F. Yin, H. Zhou, F. Meng, J. Zhou, K.-W. Chang, M. Huang, and N. Peng, “On prompt-driven safeguarding for large language models,” in Proceedings of the 41st International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 235, 21–27 Jul 2024, pp. 61 593–61 613
2024
Later among the works it cites.
Y. Wang, X. Liu, Y. Li, M. Chen, and C. Xiao, “Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting,” in European Conference on Computer Vision . Springer, 2024, pp. 77–94
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Wu, Y. Gao, B. Zhu, Z. Zhou, X. Sun, S. Yang, J.-G. Lou, Z. Ding, and L. Yang, “Strago: Harnessing strategic guidance for prompt optimization,” in Findings of the Association for Computational Linguistics: EMNLP 2024 , 2024, pp. 10 043–10 061
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “ { \{ DistServe } \} : Disaggregating prefill and decoding for goodput-optimized large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , 2024, pp. 193–210
2024
Later among the works it cites.
H. Sun, Z. Chen, X. Yang, Y. Tian, and B. Chen, “Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding,” in First Conference on Language Modeling , 2024
2024
Later among the works it cites.
T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao, “Medusa: Simple LLM inference acceleration framework with multiple decoding heads,” in Proceedings of the 41st International Conference on Machine Learning , vol. 235. PMLR, 2024, pp. 5209–5235
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
R. Svirschevski, A. May, Z. Chen, B. Chen, Z. Jia, and M. Ryabinin, “Specexec: Massively parallel speculative decoding for interactive llm inference on consumer devices,” Advances in Neural Information Processing Systems , vol. 37, pp. 16 342–16 368, 2024
2024
Later among the works it cites.
P. Wang, D. Zhang, L. Li, C. Tan, X. Wang, M. Zhang, K. Ren, B. Jiang, and X. Qiu, “Inferaligner: Inference-time alignment for harmlessness through cross-model guidance,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024, pp. 10 460–10 479
2024
Later among the works it cites.
X. Wang, D. Wu, Z. Ji, Z. Li, P. Ma, S. Wang, Y. Li, Y. Liu, N. Liu, and J. Rahmel, “Selfdefend: Llms can defend themselves against jailbreaking in a practical manner,” CoRR , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Ghosh, P. Varshney, M. N. Sreedhar, A. Padmakumar, T. Rebedea, J. R. Varghese, and C. Parisien, “Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails,” in Neurips Safe Generative AI Workshop 2024 , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Jin, A. Zhou, J. Menke, and H. Wang, “Jailbreaking large language models against moderation guardrails via cipher characters,” Advances in Neural Information Processing Systems , vol. 37, pp. 59 408–59 435, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
D. Zhu, D. Chen, X. Wu, J. Geng, Z. Li, J. Grossklags, and L. Ma, “Privauditor: Benchmarking data protection vulnerabilities in llm adaptation techniques,” Advances in Neural Information Processing Systems , vol. 37, pp. 9668–9689, 2024
2024
Later among the works it cites.
L. Rossi, B. Marek, V. Hanke, X. Wang, M. Backes, A. Dziedzic, and F. Boenisch, “Auditing empirical privacy protection of private llm adaptations,” in Neurips Safe Generative AI Workshop 2024
2024
Later among the works it cites.
T. Singh, H. Aditya, V. K. Madisetti, and A. Bahga, “Whispered tuning: Data privacy preservation in fine-tuning llms through differential privacy,” Journal of Software Engineering and Applications , vol. 17, no. 1, pp. 1–22, 2024
2024
Later among the works it cites.
O. Cartwright, H. Dunbar, and T. Radcliffe, “Evaluating privacy compliance in commercial large language models-chatgpt, claude, and gemini,” 2024
2024
Later among the works it cites.
Y. Song, R. Liu, S. Chen, Q. Ren, Y. Zhang, and Y. Yu, “Securesql: Evaluating data leakage of large language models as natural language interfaces to databases,” in Findings of the Association for Computational Linguistics: EMNLP 2024 , 2024, pp. 5975–5990
2024
Later among the works it cites.
X. Liu, Y. Zhu, J. Gu, Y. Lan, C. Yang, and Y. Qiao, “Mm-safetybench: A benchmark for safety evaluation of multimodal large language models,” in European Conference on Computer Vision . Springer, 2024, pp. 386–403
2024
Later among the works it cites.
W. Luo, S. Ma, X. Liu, X. Guo, and C. Xiao, “Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks,” arXiv e-prints , pp. arXiv–2404, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob et al. , “Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 375–14 385
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Zhang, W. Shao, H. Liu, Y. Ma, P. Luo, Y. Qiao, and K. Zhang, “Avibench: Towards evaluating the robustness of large vision-language model on adversarial visual-instructions,” arXiv e-prints , pp. arXiv–2403, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
E. Slyman, S. Lee, S. Cohen, and K. Kafle, “Fairdedup: Detecting and mitigating vision-language fairness disparities in semantic dataset deduplication,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 905–13 916
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Wang, A. Zhang, N. Duy Tai, J. Sun, T.-S. Chua et al. , “Ali-agent: Assessing llms’ alignment with human values via agent-based evaluation,” Advances in Neural Information Processing Systems , vol. 37, pp. 99 040–99 088, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Yang, C. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press, “Swe-agent: Agent-computer interfaces enable automated software engineering,” Advances in Neural Information Processing Systems , vol. 37, pp. 50 528–50 652, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. S. Kannan, V. L. Venkatesh, and B.-C. Min, “Smart-llm: Smart multi-agent robot task planning using large language models,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2024, pp. 12 140–12 147
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Yang, X. Bi, Y. Lin, S. Chen, J. Zhou, and X. Sun, “Watch out for your agents! investigating backdoor threats to llm-based agents,” Advances in Neural Information Processing Systems , vol. 37, pp. 100 938–100 964, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Zhang, H. Xu, Z. Ba, Z. Wang, Y. Hong, J. Liu, Z. Qin, and K. Ren, “Privacyasst: Safeguarding user privacy in tool-using large language model agents,” IEEE Transactions on Dependable and Secure Computing , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
K. N. Jeptoo and C. Sun, “Enhancing fake news detection with large language models through multi-agent debates,” in CCF International Conference on Natural Language Processing and Chinese Computing . Springer, 2024, pp. 474–486
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Yang, S. S. Raman, A. Shah, and S. Tellex, “Plug in the safety chip: Enforcing constraints for llm-driven robot agents,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 14 435–14 442
2024
Later among the works it cites.
J. Zhang, C. Xu, and B. Li, “Chatscene: Knowledge-enabled safety-critical scenario generation for autonomous vehicles,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 15 459–15 469
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Sun, N. Salami Pargoo, P. Jin, and J. Ortiz, “Optimizing autonomous driving for safety: A human-centric approach with llm-enhanced rlhf,” in Companion of the 2024 on ACM International Joint Conference on Pervasive and Ubiquitous Computing , 2024, pp. 76–80
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Deng, Y. Guo, C. Han, W. Ma, J. Xiong, S. Wen, and Y. Xiang, “Ai agents under threat: A survey of key security challenges and future pathways,” ACM Computing Surveys , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
B. Chen, G. Li, X. Lin, Z. Wang, and J. Li, “Blockagents: Towards byzantine-robust llm-based multi-agent coordination via blockchain,” in Proceedings of the ACM Turing Award Celebration Conference-China 2024 , 2024, pp. 187–192
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Yuan, F. Chen, Z. Zhang, and Y. Yu, “Communication-robust multi-agent learning by adaptable auxiliary multi-agent adversary generation,” Frontiers of Computer Science , vol. 18, no. 6, p. 186331, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Yu, Y. Qiu, Q. Yao, Y. Shen, X. Zhang, and J. Wang, “Robust communicative multi-agent reinforcement learning with active defense,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 16, 2024, pp. 17 575–17 582
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Guo, X. Liu, C. Xie, A. Zhou, Y. Zeng, Z. Lin, D. Song, and B. Li, “Redcode: Risky code execution and generation benchmark for code agents,” Advances in Neural Information Processing Systems , vol. 37, pp. 106 190–106 236, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Zhu, B. Wu, Z. Zhang, and B. Wu, “Riskawarebench: Towards evaluating physical risk awareness for high-level planning of llm-based embodied agents,” arXiv e-prints , pp. arXiv–2408, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Wu, S. Zhao, Q. Huang, K. Huang, M. Yasunaga, K. Cao, V. Ioannidis, K. Subbian, J. Leskovec, and J. Y. Zou, “Avatar: Optimizing llm agents for tool usage via contrastive reasoning,” Advances in Neural Information Processing Systems , vol. 37, pp. 25 981–26 010, 2024
2024
Later among the works it cites.
Z. Shen, “Llm with tools: A survey,” arXiv preprint arXiv:2409.18807 , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
G. Sriramanan, S. Bharti, V. S. Sadasivan, S. Saha, P. Kattakinda, and S. Feizi, “Llm-check: Investigating detection of hallucinations in large language models,” Advances in Neural Information Processing Systems , vol. 37, pp. 34 188–34 216, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Liu, Y. Yao, J.-F. Ton, X. Zhang, R. Guo, H. Cheng, Y. Klochkov, M. F. Taufiq, and H. Li, “Trustworthy llms: a survey and guideline for evaluating large language models’ alignment,” 2024
2024
Later among the works it cites.
G. Feretzakis and V. S. Verykios, “Trustworthy ai: Securing sensitive data in large language models,” AI , vol. 5, no. 4, pp. 2773–2800, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
X. Zou, Y. Yan, X. Hao, Y. Hu, H. Wen, E. Liu, J. Zhang, Y. Li, T. Li, Y. Zheng et al. , “Deep learning for cross-domain data fusion in urban computing: Taxonomy, advances, and outlook,” Information Fusion , vol. 113, p. 102606, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou et al. , “The rise and potential of large language model based agents: A survey,” Science China Information Sciences , vol. 68, no. 2, p. 121101, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Chen, W. Sun, C. Fang, Z. Chen, Y. Ge, T. Han, Q. Zhang, Y. Liu, Z. Chen, and B. Xu, “Security of language models for code: A systematic literature review,” ACM Transactions on Software Engineering and Methodology , vol. 1, no. 1, pp. 1–66, 2025
2025
Closest in time.
W. Qu, Y. Zhou, Y. Wu, T. Xiao, B. Yuan, Y. Li, and J. Zhang, “Prompt inversion attack against collaborative inference of large language models,” in IEEE S&P , 2025
2025
Closest in time.
J. Wu, S. Yang, R. Zhan, Y. Yuan, L. S. Chao, and D. F. Wong, “A survey on llm-generated text detection: Necessity, methods, and future directions,” Computational Linguistics , pp. 1–66, 2025
2025
Closest in time.
W. Sun, Y. Chen, C. Fang, Y. Feng, Y. Xiao, A. Guo, Q. Zhang, Y. Liu, B. Xu, and Z. Chen, “Eliminating backdoors in neural code models for secure code understanding,” in Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering . Trondheim, Norway: ACM, Mon 23 - Fri 27 June 2025, pp. 1–23
2025
Closest in time.
X. Qi, A. Panda, K. Lyu, X. Ma, S. Roy, A. Beirami, P. Mittal, and P. Henderson, “Safety alignment should be made more than just a few tokens deep,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=6Mxhg9PtDE
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. C. Das, M. H. Amini, and Y. Wu, “Security and privacy challenges of large language models: A survey,” ACM Computing Surveys , vol. 57, no. 6, pp. 1–39, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
W. Sun, Y. Chen, M. Yuan, C. Fan, Z. Chen, C. Wang, Y. Liu, B. Xu, and Z. Chen, “Show me your code! kill code poisoning: A lightweight method based on code naturalness,” in Proceedings of the IEEE/ACM 47th International Conference on Software Engineering . Ottawa, Ontario, Canada: IEEE Computer Society, Sun 27 April - Sat 3 May 2025, pp. 1–13
2025
Closest in time.
Q. Zhang, H. Qiu, D. Wang, Y. Li, T. Zhang, W. Zhu, H. Weng, L. Yan, and C. Zhang, “A benchmark for semantic sensitive information in llms outputs,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.
Y. He, B. Li, L. Liu, Z. Ba, W. Dong, Y. Li, Z. Qin, K. Ren, and C. Chen, “Towards label-only membership inference attack against pre-trained large language models,” in USENIX Security , 2025
2025
Closest in time.
J. Ren, K. Chen, C. Chen, V. Sehwag, Y. Xing, J. Tang, and L. Lyu, “Self-comparison for dataset-level membership inference in large (vision-) language model,” in Proceedings of the ACM on Web Conference 2025 , 2025, pp. 910–920
2025
Closest in time.
2025
Closest in time.
S. Li, F. Liu, L. Cui, J. Lu, Q. Xiao, X. Yang, P. Liu, K. Sun, Z. Ma, and X. Wang, “Safe planner: Empowering safety awareness in large pre-trained models for robot task planning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 14, 2025, pp. 14 619–14 627
2025
Closest in time.
2025
Closest in time.
N. Chakraborty, M. Ornik, and K. Driggs-Campbell, “Hallucination detection in foundation models for decision-making: A flexible definition and review of the state of the art,” ACM Computing Surveys , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
H. Shen, P.-Y. Chen, P. Das, and T. Chen, “SEAL: Safety-enhanced aligned LLM fine-tuning via bilevel data selection,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=VHguhvcoM5
2025
Closest in time.
2025
Closest in time.
F. Barez, T. Fu, A. Prabhu, S. Casper, A. Sanyal, A. Bibi, A. O’Gara, R. Kirk, B. Bucknall, T. Fist, L. Ong, P. Torr, K. Lam, R. Trager, D. Krueger, S. Mindermann, J. Hernández-Orallo, M. Geva, and Y. Gal, “Open problems in machine unlearning for AI safety,” CoRR , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
R. Ye, J. Chai, X. Liu, Y. Yang, Y. Wang, and S. Chen, “Emerging safety attack and defense in federated instruction tuning of large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=sYNWqQYJhz
2025
Closest in time.
J. Li and J.-E. Kim, “Safety alignment shouldn’t be complicated,” 2025. [Online]. Available: https://openreview.net/forum?id=9H91juqfgb
2025
Closest in time.
S. Li, L. Yao, L. Zhang, and Y. Li, “Safety layers in aligned large language models: The key to LLM security,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=kUH1yPMAn7
2025
Closest in time.
Z. Zhou, H. Yu, X. Zhang, R. Xu, F. Huang, K. Wang, Y. Liu, J. Fang, and Y. Li, “On the role of attention heads in large language model safety,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=h0Ak8A5yqw
2025
Closest in time.
M. Li, W. M. Si, M. Backes, Y. Zhang, and Y. Wang, “SaloRA: Safety-alignment preserved low-rank adaptation,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=GOoVzE9nSj
2025
Closest in time.
F. Eiras, A. Petrov, P. Torr, M. P. Kumar, and A. Bibi, “Do as i do (safely): Mitigating task-specific fine-tuning risks in large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=lXE5lB6ppV
2025
Closest in time.
B. Yi, T. Huang, S. Chen, T. Li, Z. Liu, Z. Chu, and Y. Li, “Probe before you talk: Towards black-box defense against backdoor unalignment for large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=EbxYDBhE3S
2025
Closest in time.
2025
Closest in time.
M. Zhu, Y. Weng, L. Yang, Y. Wei, N. Zhang, and Y. Zhang, “Locking down the finetuned LLMs safety,” 2025. [Online]. Available: https://openreview.net/forum?id=YGoFl5KKFc
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Q. Liu, C. Shang, L. Liu, N. Pappas, J. Ma, N. A. John, S. Doss, L. Marquez, M. Ballesteros, and Y. Benajiba, “Unraveling and mitigating safety alignment degradation of vision-language models,” 2025. [Online]. Available: https://openreview.net/forum?id=EEWpE9cR27
2025
Closest in time.
NIST, “Artificial intelligence risk management framework: Generative artificial intelligence profile (initial public draft),” 2024, accessed: 2025-05-29. [Online]. Available: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-1.ipd.pdf
2025
Closest in time.
R. Tamirisa, B. Bharathi, L. Phan, A. Zhou, A. Gatti, T. Suresh, M. Lin, J. Wang, R. Wang, R. Arel, A. Zou, D. Song, B. Li, D. Hendrycks, and M. Mazeika, “Tamper-Resistant Safeguards for Open-Weight LLMs,” Feb. 2025
2025
Closest in time.
2025
Closest in time.
H. Zhang, J. Huang, K. Mei, Y. Yao, Z. Wang, C. Zhan, H. Wang, and Y. Zhang, “Agent security bench (ASB): Formalizing and benchmarking attacks and defenses in LLM-based agents,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=V4y0CpX4hK
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
OpenAI, “Detecting misbehavior in frontier reasoning models,” https://openai.com/index/chain-of-thought-monitoring/ , Mar. 2025, accessed: 2025-05-14
2025
Closest in time.
V. Krakovna, J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg, “Specification gaming: the flip side of ai ingenuity,” 2020, accessed: 2025-03-30. [Online]. Available: https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
2025
Closest in time.
S. Liu, Y. Yao, J. Jia, S. Casper, N. Baracaldo, P. Hase, Y. Yao, C. Y. Liu, X. Xu, H. Li et al. , “Rethinking machine unlearning for large language models,” Nature Machine Intelligence , pp. 1–14, 2025
2025
Closest in time.
Y. Yao, X. Xu, and Y. Liu, “Large language model unlearning,” Advances in Neural Information Processing Systems , vol. 37, pp. 105 425–105 475, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
A. Blanco-Justicia, J. Domingo-Ferrer, N. M. Jebreel, B. Manzanares-Salor, and D. Sánchez, “Unlearning in large language models: We are not there yet,” Computer , vol. 58, no. 1, pp. 97–100, 2025
2025
Closest in time.
2025
Closest in time.
A. Blanco-Justicia, N. Jebreel, B. Manzanares-Salor, D. Sánchez, J. Domingo-Ferrer, G. Collell, and K. Eeik Tan, “Digital forgetting in large language models: A survey of unlearning methods,” Artificial Intelligence Review , vol. 58, no. 3, p. 90, 2025
2025
Closest in time.
N. Li, C. Zhou, Y. Gao, H. Chen, Z. Zhang, B. Kuang, and A. Fu, “Machine unlearning: Taxonomy, metrics, applications, challenges, and prospects,” IEEE Transactions on Neural Networks and Learning Systems , 2025
2025
Closest in time.
K. Zhao, M. Kurmanji, G.-O. Bărbulescu, E. Triantafillou, and P. Triantafillou, “What makes unlearning hard and what to do about it,” Advances in Neural Information Processing Systems , vol. 37, pp. 12 293–12 333, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Q. Zhang, H. Qiu, D. Wang, Y. Li, T. Zhang, W. Zhu, H. Weng, L. Yan, and C. Zhang, “Large scale knowledge washing,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
D. Sanyal and M. Mandal, “Alu: Agentic llm unlearning,” arXiv preprint arXiv:2502.00406 , 2025
2025
Closest in time.
2025
Closest in time.
H. Liu, P. Xiong, T. Zhu, and S. Y. Philip, “A survey on machine unlearning: Techniques and new emerged privacy risks,” Journal of Information Security and Applications , vol. 90, p. 104010, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Zhang and Z. Wei, “Boosting jailbreak attack with momentum,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2025, pp. 1–5
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Hu, G. Wu, S. Mitra, R. Zhang, T. Sun, H. Huang, and V. Swaminathan, “Token-level adversarial prompt detection based on perplexity measures and contextual information,” in ICLR 2025 Workshop on Building Trust in Language Models and Applications , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Zhang, L. Ding, L. Zhang, and D. Tao, “Intention analysis makes LLMs a good jailbreak defender,” in Proceedings of the 31st International Conference on Computational Linguistics , 2025, pp. 2947–2968
2025
Closest in time.
X. Song, S. Duan, and G. Liu, “Alis: Aligned llm instruction security strategy for unsafe input prompt,” in Proceedings of the 31st International Conference on Computational Linguistics , 2025, pp. 9124–9146
2025
Closest in time.
OpenAI, “Improving model safety behavior with rule-based rewards,” https://openai.com/index/improving-model-safety-behavior-with-rule-based-rewards/ , 2025, accessed: 2025-03-24
2025
Closest in time.
Y. Zhang, L. Ding, L. Zhang, and D. Tao, “Intention analysis makes llms a good jailbreak defender,” in Proceedings of the 31st International Conference on Computational Linguistics , 2025, pp. 2947–2968
2025
Closest in time.
S. Slocum and D. Hadfield-Menell, “Inverse prompt engineering for task-specific LLM safety,” 2025. [Online]. Available: https://openreview.net/forum?id=3MDmM0rMPQ
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Tang, T. Fan, and C. Huang, “Autoagent: A fully-automated and zero-code framework for llm agents,” arXiv e-prints , pp. arXiv–2502, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Chen, Z. Xiang, C. Xiao, D. Song, and B. Li, “Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases,” Advances in Neural Information Processing Systems , vol. 37, pp. 130 185–130 213, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
H. Chase, “Langchain: Build applications with llms through composability,” https://github.com/langchain-ai/langchain , 2022, accessed: Apr. 2025
2025
Closest in time.
J. Wu et al. , “Llamaindex: Connecting llms to your knowledge,” https://github.com/jerryjliu/llama_index , 2023, accessed: Apr. 2025
2025
Closest in time.
OpenAI, “Function calling in openai models,” https://platform.openai.com/docs/guides/functions , 2023, accessed: Apr. 2025
2025
Closest in time.
Anthropic, “Model context protocol,” 2024, accessed: 2025-04-19. [Online]. Available: https://www.anthropic.com/news/model-context-protocol
2025
Closest in time.
Google, “A2a: Agent2agent protocol,” 2025, accessed: 2025-04-21. [Online]. Available: https://github.com/google/A2A
2025
Closest in time.
G. Chang, “Anp: Agent network protocol,” 2024, accessed: 2025-04-21. [Online]. Available: https://www.agent-network-protocol.com/
2025
Closest in time.
WildCardAI, “agents.json specification,” https://github.com/wild-card-ai/agents-json , 2025, accessed: 2025-04-22
2025
Closest in time.
NEAR, “Aitp: Agent interaction & transaction protocol,” 2025, accessed: 2025-04-22. [Online]. Available: https://aitp.dev/
2025
Closest in time.
L. F. Al and L. Data, “Acp: Agent communication protocol,” 2025, accessed: 2025-04-22. [Online]. Available: https://github.com/orgs/i-am-bee/discussions/284
2025
Closest in time.
G. Cisco, Langchain, “Acp: Agent connect protocol,” 2025, accessed: 2025-04-22. [Online]. Available: https://spec.acp.agntcy.org/
2025
Closest in time.
Eclipse, “Language model operating system (lmos),” https://eclipse.dev/lmos/ , 2025, accessed: 2025-04-22
2025
Closest in time.
AlEngineerFoundation, “Agent protocol,” https://agentprotocol.ai/ , 2025, accessed: 2025-04-22
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
A. Liu, Y. Zhou, X. Liu, T. Zhang, S. Liang, J. Wang, Y. Pu, T. Li, J. Zhang, W. Zhou et al. , “Compromising llm driven embodied agents with contextual backdoor attacks,” IEEE Transactions on Information Forensics and Security , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Shao, Y. Li, H. Yao, Y. He, Z. Qin, and K. Ren, “Explanation as a watermark: Towards harmless and multi-bit model ownership verification via watermarking feature attribution,” in NDSS , 2025
2025
Closest in time.
European Union, “Artificial intelligence act,” 2024, accessed: 2025-03-07. [Online]. Available: https://artificialintelligenceact.eu/
2025
Closest in time.
The White House, “Safe, secure, and trustworthy development and use of artificial intelligence,” 2023, accessed: 2025-03-07
2025
Closest in time.
J. Zhao, S. Wang, Y. Zhao, X. Hou, K. Wang, P. Gao, Y. Zhang, C. Wei, and H. Wang, “Models are codes: Towards measuring malicious code poisoning attacks on pre-trained model hubs,” in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering , 2024, pp. 2087–2098
2098
Closest in time.