Fetching the paper…
Reading the bibliography…
The advent of advanced AI underscores the urgent need for comprehensive safety evaluations, necessitating collaboration across communities (i.e., AI, software engineering, and governance).
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Guide to the Software Engineering Body of Knowledge (SWEBOK): Version 3.0
IEEE Computer Society. 2014 · 2014
Earlier work this paper cites.
Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the Protection of Natural Persons with regard to the Processing of Personal Data and on the Free Movement of Such Data, and Repealing Directive 95/46/EC (General Data Protection Regulation)
2016 · 2016
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
ISO/IEC 22989:2022, Information technology – Artificial intelligence – Artificial intelligence concepts and terminology
2022 · 2022
Earlier work this paper cites.
The Bletchley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023
2023 · 2023
Earlier work this paper cites.
Managing ai risks in an era of rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, et al · 2023
Earlier work this paper cites.
Fairness Testing: A Comprehensive Survey and Analysis of Trends
Zhenpeng Chen, Jie M Zhang, Max Hort, Mark Harman, and Federica Sarro. 2023 · 2023
Earlier work this paper cites.
Responsible ai pattern catalogue: A collection of best practices for ai governance and engineering
Qinghua Lu, Liming Zhu, Xiwei Xu, Jon Whittle, Didar Zowghi, and Aurelie Jacquet. 2023b · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023 · 2023
Cited alongside, same era.
AI Risk Management Framework (AI RMF 1.0)
US National Institute of Standards and Technology (NIST). 2023 · 2023
Cited alongside, same era.
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence
The White House. 2023 · 2023
Cited alongside, same era.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
AI Safety Institute Approach to Evaluations
AI Safety Institute. 2024 · 2024
Closest in time.
Evaluating LLM Systems: Metrics, Challenges, and Best Practices
Jane Huang, Kirk Li, and Daniel Yehdego. 2024 · 2024
Closest in time.
A Causal Framework for AI Regulation and Auditing
Lee Sharkey, Clíodhna Ní Ghuidhir, Dan Braun, Jérémy Scheurer, Mikita Balesni, Lucius Bushnaq, Charlotte Stix, and Marius Hobbhahn. 2024 · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Closest in time.
Introducing v0. 5 of the ai safety benchmark from mlcommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M Ahmed, Victor Akinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Hamelin, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sociotechnical safety evaluation of generative ai systems
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, et al · 2023
Cited alongside, same era.
Navigating privacy and copyright challenges across the data lifecycle of generative ai
Dawen Zhang, Boming Xia, Yue Liu, Xiwei Xu, Thong Hoang, Zhenchang Xing, Mark Staples, Qinghua Lu, and Liming Zhu. 2023 · 2023
Cited alongside, same era.
Don’t Make Your LLM an Evaluation Benchmark Cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. 2023 · 2023
Cited alongside, same era.
Responsible AI: Best Practices for Creating Trustworthy AI Systems
Qinghua Lu, Liming Zhu, Jon Whittle, and Xiwei Xu. 2023a
Cited in the paper.
Towards a Responsible AI Metrics Catalogue: A Collection of Metrics for AI Accountability. In 3rd International Conference on AI Engineering–Software Engineering for AI (CAIN ’24)
Boming Xia, Qinghua Lu, Liming Zhu, Sung Une Lee, Yue Liu, and Zhenchang Xing. 2024 · 2024
Closest in time.
The Shift from Models to Compound AI Systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024 · 2024
Closest in time.
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
Zhexin Zhang, Yida Lu, Jingyuan Ma, Di Zhang, Rui Li, Pei Ke, Hao Sun, Lei Sha, Zhifang Sui, Hongning Wang, et al · 2024
Closest in time.