Fetching the paper…
Reading the bibliography…
AI progress is creating a growing range of risks and opportunities, but it is often unclear how they should be navigated.
Towards efficient data valuation based on the shapley value
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gurel, Bo Li, Ce Zhang, Dawn Song, and Costas Spanos · 1902
Earlier work this paper cites.
Blameworthiness in Multi-Agent settings
Meir Friedenberg and Joseph Y Halpern · 1903
Earlier work this paper cites.
Global AI ethics: A review of the social impacts and ethical implications of artificial intelligence
Alexa Hagerty and Igor Rubinov · 1907
Earlier work this paper cites.
CrypTFlow: Secure TensorFlow inference
Nishant Kumar, Mayank Rathee, Nishanth Chandran, Divya Gupta, Aseem Rastogi, and Rahul Sharma · 1909
Earlier work this paper cites.
Advances and open problems in federated learning
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, Rafael G L D’Oliveira, Hubert Eichner, Salim El Rouayheb, David Evans, Josh Gardner, Zachary Garrett, Adrià Gascón, Badih Ghazi, Phillip B Gibbons, Marco Gruteser, Zaid Harchaoui, Chaoyang He, Lie He, Zhouyuan Huo, Ben Hutchinson, Justin Hsu, Martin Jaggi, Tara Javidi, Gauri Joshi, Mikhail Khodak, Jakub Konečný, Aleksandra Korolova, Farinaz Koushanfar, Sanmi Koyejo, Tancrède Lepoint, Yang Liu, Prateek Mittal, Mehryar Mohri, Richard Nock, Ayfer Özgür, Rasmus Pagh, Mariana Raykova, Hang Qi, Daniel Ramage, Ramesh Raskar, Dawn Song, Weikang Song, Sebastian U Stich, Ziteng Sun, Ananda Theertha Suresh, Florian Tramèr, Praneeth Vepakomma, Jianyu Wang, Li Xiong, Zheng Xu, Qiang Yang, Felix X Yu, Han Yu, and Sen Zhao · 1912
Earlier work this paper cites.
A sociotechnical view of algorithmic fairness
Mateusz Dolata, Stefan Feuerriegel, and Gerhard Schwabe · 1917
Earlier work this paper cites.
Generative AI models should include detection mechanisms as a condition for public release
Alistair Knott, Dino Pedreschi, Raja Chatila, Tapabrata Chakraborti, Susan Leavy, Ricardo Baeza-Yates, David Eyers, Andrew Trotman, Paul D Teal, Przemyslaw Biecek, Stuart Russell, and Yoshua Bengio · 1957
Earlier work this paper cites.
Protocols for secure computations
Andrew C Yao · 1982
Earlier work this paper cites.
How to generate and exchange secrets
Andrew Chi-Chih Yao · 1986
Earlier work this paper cites.
Zero knowledge proofs of identity
Uriel Fiege, Amos Fiat, and Adi Shamir · 1987
Earlier work this paper cites.
Typologies and taxonomies: An introduction to classification techniques , volume 12
Kenneth D Bailey · 1994
Earlier work this paper cites.
Secure multi-party computation
Oded Goldreich · 1998
Earlier work this paper cites.
Proofs of work and bread pudding Protocols(Extended abstract)
Markus Jakobsson and Ari Juels · 1999
Earlier work this paper cites.
Qualitative methods: what are they and why use them?
Shoshanna Sofaer · 1999
Earlier work this paper cites.
Qualitative research & evaluation methods
Michael Quinn Patton · 2002
Earlier work this paper cites.
Typologies, taxonomies, and the benefits of policy classification
Kevin B Smith · 2002
Earlier work this paper cites.
Quantifying construct validity: Two simple measures
Drew Westen and Robert Rosenthal · 2003
Earlier work this paper cites.
Visible Thinking: Unlocking Causal Mapping for Practical Business Results
Colin Eden, Fran Ackermann, Charles B Finn, and John M Bryson · 2004
Earlier work this paper cites.
FALCON: Honest-Majority maliciously secure framework for private deep learning
Sameer Wagh, Shruti Tople, Fabrice Benhamouda, Eyal Kushilevitz, Prateek Mittal, and Tal Rabin · 2004
Earlier work this paper cites.
A review of audio fingerprinting
Pedro Cano, Eloi Batlle, Ton Kalker, and Jaap Haitsma · 2005
Earlier work this paper cites.
On construct validity: Issues of method and measurement
Gregory T Smith · 2005
Earlier work this paper cites.
Universal adversarial attacks with natural triggers for text classification
Liwei Song, Xinwei Yu, Hsuan-Tung Peng, and Karthik Narasimhan · 2005
Earlier work this paper cites.
Verifying Treaty Compliance
Rudolf Avenhaus, Nicholas Kyriakopoulos, Michel Richard, and Gotthard Stein (eds.) · 2006
Earlier work this paper cites.
Influence functions in deep learning are fragile
Samyadeep Basu, Philip Pope, and Soheil Feizi · 2006
Earlier work this paper cites.
Qualitative data analysis for health services research: developing taxonomy, themes, and theory
Elizabeth H Bradley, Leslie A Curry, and Kelly J Devers · 2007
Earlier work this paper cites.
SplitNN-driven vertical partitioning
Iker Ceballos, Vivek Sharma, Eduardo Mugica, Abhishek Singh, Alberto Roman, Praneeth Vepakomma, and Ramesh Raskar · 2008
Earlier work this paper cites.
A fully homomorphic encryption scheme
Craig Gentry · 2009
Earlier work this paper cites.
Construct validity: Advances in theory and methodology
Milton E Strauss and Gregory T Smith · 2009
Earlier work this paper cites.
The de-democratization of AI: Deep learning and the compute divide in artificial intelligence research
Nur Ahmed and Muntasir Wahed · 2010
Earlier work this paper cites.
Concealed data poisoning attacks on NLP models
Eric Wallace, Tony Z Zhao, Shi Feng, and Sameer Singh · 2010
Earlier work this paper cites.
Dominant resource fairness: Fair allocation of multiple resource types
Ali Ghodsi, Matei Zaharia, Benjamin Hindman, Andy Konwinski, Scott Shenker, and Ion Stoica · 2011
Earlier work this paper cites.
IEEE draft guide: Adoption of the project management institute (PMI) standard: A guide to the project management body of knowledge (PMBOK guide)-2008 (4th edition), June 2011
IEEE · 2011
Earlier work this paper cites.
Preventing repeated real world AI failures by cataloging incidents: The AI incident database
Sean McGregor · 2011
Earlier work this paper cites.
The normativity of copying in copyright law
Shyamkrishna Balganesh · 2012
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2012
Earlier work this paper cites.
A method for taxonomy development and its application in information systems
Robert C Nickerson, Upkar Varshney, and Jan Muntermann · 2013
Earlier work this paper cites.
Systemic risk elicitation: Using causal maps to engage stakeholders and build a comprehensive view of risks
Fran Ackermann, Susan Howick, John Quigley, Lesley Walls, and Tom Houghton · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
Multi-Resource fair allocation in heterogeneous cloud computing systems
Wei Wang, Ben Liang, and Baochun Li · 2014
Earlier work this paper cites.
Crypto-Nets: Neural networks over encrypted data
Pengtao Xie, Misha Bilenko, Tom Finley, Ran Gilad-Bachrach, Kristin Lauter, and Michael Naehrig · 2014
Earlier work this paper cites.
A game theory approach to fair and efficient resource allocation in cloud computing
Xin Xu and Huiqun Yu · 2014
Earlier work this paper cites.
Trusted execution environment: What it is, and what it is not
Mohamed Sabt, Mohammed Achemlal, and Abdelmadjid Bouabdallah · 2015
Earlier work this paper cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Alexandra Chouldechova · 2016
Earlier work this paper cites.
EU general data protection regulation (GDPR): Regulation (EU) 2016/679 of the european parliament and of the council of 27 april 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing directive 95/46/EC (general data protection regulation)
European Commission · 2016
Earlier work this paper cites.
Stealing machine learning models via prediction APIs
Florian Tramèr, Fan Zhang, Ari Juels, Michael K Reiter, and Thomas Ristenpart · 2016
Earlier work this paper cites.
Regulation of emerging risks
Matthew T Wansley · 2016
Earlier work this paper cites.
Nudging robots: Innovative solutions to regulate artificial intelligence
M Guihot, A F Matthew, and N P Suzor · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Certified defenses for data poisoning attacks
Jacob Steinhardt, Pang Wei Koh, and Percy Liang · 2017
Earlier work this paper cites.
Unsupervised deep video hashing with balanced rotation
Wu, Li Liu, Yuchen Guo, Guiguang Ding, J Han, Jialie Shen, and Ling Shao · 2017
Earlier work this paper cites.
A survey on homomorphic encryption schemes: Theory and implementation
Abbas Acar, Hidayet Aksu, A Selcuk Uluagac, and Mauro Conti · 2018
Earlier work this paper cites.
AI governance: A research agenda
Allan Dafoe · 2018
Earlier work this paper cites.
A Pragmatic Introduction to Secure Multi-Party Computation
David Evans, Vladimir Kolesnikov, and Mike Rosulek · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2018
Earlier work this paper cites.
B-TREPID: Batteryless tamper-resistant envelope with a PUF and integrity detection
Vincent Immler, Johannes Obermaier, Martin König, Matthias Hiller, and Georg Sig · 2018
Earlier work this paper cites.
A fair resource allocation approach in cloud computing environments
Maha Jebalia, Asma Ben Letaïfa, Mohamed Hamdi, and Sami Tabbane · 2018
Earlier work this paper cites.
Inherent Trade-Offs in algorithmic fairness
Jon Kleinberg · 2018
Earlier work this paper cites.
Toward scalable verification for Safety-Critical deep networks
Lindsey Kuper, Guy Katz, Justin Gottschlich, Kyle Julian, Clark Barrett, and Mykel Kochenderfer · 2018
Earlier work this paper cites.
Scaling shared model governance via model splitting
Miljan Martic, Jan Leike, Andrew Trask, Matteo Hessel, Shane Legg, and Pushmeet Kohli · 2018
Earlier work this paper cites.
The past, present, and future of physical security enclosures: From Battery-Backed monitoring to PUF-Based inherent security and beyond
Johannes Obermaier and Vincent Immler · 2018
Earlier work this paper cites.
Knockoff nets: Stealing functionality of Black-Box models
Tribhuvanesh Orekondy, B Schiele, and Mario Fritz · 2018
Earlier work this paper cites.
Chameleon: A hybrid secure computation framework for machine learning applications
M Sadegh Riazi, Christian Weinert, Oleksandr Tkachenko, Ebrahim M Songhori, Thomas Schneider, and Farinaz Koushanfar · 2018
Earlier work this paper cites.
A generic framework for privacy preserving deep learning
Theo Ryffel, Andrew Trask, Morten Dahl, Bobby Wagner, Jason Mancuso, Daniel Rueckert, and Jonathan Passerat-Palmbach · 2018
Earlier work this paper cites.
Classification of global catastrophic risks connected with artificial intelligence
Alexey Turchin and David Denkenberger · 2018
Earlier work this paper cites.
Split learning for health: Distributed deep learning without sharing raw patient data
Praneeth Vepakomma, Otkrist Gupta, Tristan Swedish, and Ramesh Raskar · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Earlier work this paper cites.
DeepAttest: an end-to-end attestation framework for deep neural networks
Huili Chen, Cheng Fu, Bita Darvish Rouhani, Jishen Zhao, and Farinaz Koushanfar · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
Data Shapley: Equitable Valuation of Data for Machine Learning
Amirata Ghorbani and James Zou · 2019
Earlier work this paper cites.
Recent trends in applying TPM to cloud computing
Shohreh Hosseinzadeh, Bernardo Sequeiros, Pedro R M Inácio, and Ville Leppänen · 2019
Earlier work this paper cites.
Secure physical enclosures from covers with Tamper-Resistance
Vincent Immler, Johannes Obermaier, Kuan Kuan Ng, Fei Xiang Ke, Jinyu Lee, Yak Peng Lim, Wei Koon Oh, Keng Hoong Wee, and Georg Sigl · 2019
Earlier work this paper cites.
The marabou framework for verification and analysis of deep neural networks
Guy Katz, Derek A Huang, Duligur Ibeling, Kyle Julian, Christopher Lazarus, Rachel Lim, Parth Shah, Shantanu Thakoor, Haoze Wu, Aleksandar Zeljić, David L Dill, Mykel J Kochenderfer, and Clark Barrett · 2019
Earlier work this paper cites.
Rethink government with AI
Helen Margetts and Cosmina Dorobantu · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
Towards deep neural network training on encrypted data
Karthik Nandakumar, Nalini Ratha, Sharath Pankanti, and Shai Halevi · 2019
Earlier work this paper cites.
Oecd principles on artificial intelligence
OECD · 2019
Earlier work this paper cites.
Energy and Policy Considerations for Deep Learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Earlier work this paper cites.
SecureNN: 3-party secure computation for neural network training
Sameer Wagh, Divya Gupta, and Nishanth Chandran · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé, III, and Hanna Wallach · 2020
Earlier work this paper cites.
Toward trustworthy AI development: Mechanisms for supporting verifiable claims
Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, Tegan Maharaj, Pang Wei Koh, Sara Hooker, Jade Leung, Andrew Trask, Emma Bluemke, Jonathan Lebensold, Cullen O’Keefe, Mark Koren, Théo Ryffel, J B Rubinovitz, Tamay Besiroglu, Federica Carugati, Jack Clark, Peter Eckersley, Sarah de Haas, Maritza Johnson, Ben Laurie, Alex Ingerman, Igor Krawczuk, Amanda Askell, Rosario Cammarota, Andrew Lohn, David Krueger, Charlotte Stix, Peter Henderson, Logan Graham, Carina Prunkl, Bianca Martin, Elizabeth Seger, Noa Zilberman, Seán Ó hÉigeartaigh, Frens Kroeger, Girish Sastry, Rebecca Kagan, Adrian Weller, Brian Tse, Elizabeth Barnes, Allan Dafoe, Paul Scharre, Ariel Herbert-Voss, Martijn Rasser, Shagun Sodhani, Carrick Flynn, Thomas Krendl Gilbert, Lisa Dyer, Saif Khan, Yoshua Bengio, and Markus Anderljung · 2020
Earlier work this paper cites.
Cryptanalytic extraction of neural network models
Nicholas Carlini, Matthew Jagielski, and Ilya Mironov · 2020
Earlier work this paper cites.
Hard choices in artificial intelligence: Addressing normative uncertainty through sociotechnical commitments (AIES ’20)
Roel I J Dobbe, Thomas Krendl Gilbert, and Yonatan Mintz · 2020
Earlier work this paper cites.
Physical unclonable functions
Yansong Gao, Said F Al-Sarawi, and Derek Abbott · 2020
Earlier work this paper cites.
Towards verification of neural networks for small unmanned aircraft collision avoidance
Ahmed Irfan, Kyle D Julian, Haoze Wu, Clark Barrett, Mykel J Kochenderfer, Baoluo Meng, and James Lopez · 2020
Earlier work this paper cites.
High accuracy and high fidelity extraction of neural networks
Matthew Jagielski, Nicholas Carlini, David Berthelot, Alex Kurakin, and Nicolas Papernot · 2020
Earlier work this paper cites.
Hacking AI: A primer for policymakers on machine learning cybersecurity
Andrew Lohn · 2020
Earlier work this paper cites.
Manufacturing consent: The modern pandemic of technosolutionism
Katina Michael, Roba Abbas, Rafael A Calvo, George Roussos, Eusebio Scornavacca, and Samuel Fosso Wamba · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan, IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Leaky DNN: Stealing Deep-Learning model secret with GPU Context-Switching Side-Channel
Junyi Wei, Yicheng Zhang, Zhe Zhou, Zhou Li, and Mohammad Abdullah Al Faruque · 2020
Earlier work this paper cites.
Multimodal datasets: Misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Earlier work this paper cites.
Machine unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Earlier work this paper cites.
What will it take to fix benchmarking in natural language understanding?
Samuel R Bowman and George Dahl · 2021
Earlier work this paper cites.
Poisoning the unlabeled dataset of Semi-Supervised learning
Nicholas Carlini · 2021
Earlier work this paper cites.
Poisoning and backdooring contrastive learning
Nicholas Carlini and Andreas Terzis · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel · 2021
Earlier work this paper cites.
The limits of global inclusion in AI development
Alan Chan, Chinasa T Okolo, Zachary Terner, and Angelina Wang · 2021
Earlier work this paper cites.
Data quality measures and efficient evaluation algorithms for Large-Scale High-Dimensional data
Hyeongmin Cho and Sangkyun Lee · 2021
Earlier work this paper cites.
AI, governance and ethics
Angela Daly, Thilo Hagendorff, Li Hui, Monique Mann, Vidushi Marda, Ben Wagner, Wayne Wei Wang, Hans-W Micklitz, Oreste Pollicino, Amnon Reichman, Andrea Simoncini, Giovanni Sartor, and Giovanni De Gregorio · 2021
Earlier work this paper cites.
Mostafa Dehghani, Yi Tay, Alexey A Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals · 2021
Earlier work this paper cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Earlier work this paper cites.
Interactive proofs for verifying machine learning
Shafi Goldwasser, Guy N Rothblum, Jonathan Shafer, and Amir Yehudayoff · 2021
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Building a national AI research resource: A blueprint for the national research cloud
Daniel E Ho, Jennifer King, Russell C Wald, and Christopher Wan · 2021
Earlier work this paper cites.
ISO/IEC 22237-1:2021: Information technology — data centre facilities and infrastructures — part 1: General concepts
International Organization for Standardization · 2021
Earlier work this paper cites.
Proof-of-Learning: Definitions and practice
Hengrui Jia, Mohammad Yaghini, Christopher A Choquette-Choo, Natalie Dullerud, Anvith Thudi, Varun Chandrasekaran, and Nicolas Papernot · 2021
Earlier work this paper cites.
CrypTen: Secure Multi-Party Computation Meets Machine Learning
Brian Knott, Shobha Venkataraman, Awni Hannun, Shubho Sengupta, and Mark Ibrahim · 2021
Earlier work this paper cites.
Platform research access in article 31 of the digital services act: Sword without a shield?
Paddy Leerssen · 2021
Earlier work this paper cites.
What’s in the box? an analysis of undesirable content in the Common Crawl corpus
Alexandra Luccioni and Joseph Viviano · 2021
Earlier work this paper cites.
Scientific credibility of machine translation research: A Meta-Evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino · 2021
Earlier work this paper cites.
Algorithmic impact assessments and accountability: The co-construction of impacts
Jacob Metcalf, Emanuel Moss, Elizabeth Anne Watkins, Ranjit Singh, and Madeleine Clare Elish · 2021
Earlier work this paper cites.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2021
Earlier work this paper cites.
Assembling accountability: Algorithmic impact assessment for the public interest
Emanuel Moss, Elizabeth Anne Watkins, Ranjit Singh, Madeleine Clare Elish, and Jacob Metcalf · 2021
Earlier work this paper cites.
AI and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily Denton, Emily M Bender, Alex Hanna, and Amandalynne Paullada · 2021
Earlier work this paper cites.
You autocomplete me: Poisoning vulnerabilities in neural code completion
Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov · 2021
Earlier work this paper cites.
An institutional view of algorithmic impact assessments
Andrew D Selbst · 2021
Earlier work this paper cites.
CryptGPU: Fast Privacy-Preserving machine learning on the GPU
Tan, Knott, Tian, and Wu · 2021
Earlier work this paper cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Earlier work this paper cites.
Why and how governments should monitor AI development
Jess Whittlestone and Jack Clark · 2021
Earlier work this paper cites.
Detecting AI trojans using meta neural analysis
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A Gunter, and Bo Li · 2021
Earlier work this paper cites.
Common regulatory capacity for AI
Mhairi Aitken, David Leslie, Florian Ostmann, Jacob Pratt, Helen Margetts, and Cosmina Dorobantu · 2022
Earlier work this paper cites.
Towards tracing factual knowledge in language models back to the training data
Ekin Akyürek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu · 2022
Earlier work this paper cites.
Adversarial example detection for DNN models: a review and experimental comparison
Ahmed Aldahdooh, Wassim Hamidouche, Sid Ahmed Fezza, and Olivier Déforges · 2022
Earlier work this paper cites.
Choking off china’s access to the future of AI
Gregory C Allen · 2022
Earlier work this paper cites.
Improved export controls enforcement technology needed for U.S. national security
Gregory C Allen, Emily Benson, and William Alan Reinsch · 2022
Earlier work this paper cites.
Compute funds and pre-trained models, April 2022
Markus Anderljung, Lennart Heim, and Toby Shevlane · 2022
Earlier work this paper cites.
Reconstructing training data with informed adversaries
Borja Balle, Giovanni Cherubin, and Jamie Hayes · 2022
Earlier work this paper cites.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr · 2022
Earlier work this paper cites.
The Oxford Handbook of AI Governance
Justin B Bullock, Yu-Che Chen, Johannes Himmelreich, Valerie M Hudson, Anton Korinek, Matthew M Young, and Baobao Zhang (eds.) · 2022
Earlier work this paper cites.
Implementation of additional export controls: Certain advanced computing and semiconductor manufacturing items; supercomputer and semiconductor end use; entity list modification, October 2022b
Bureau of Industry and Security · 2022
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Earlier work this paper cites.
The privacy onion effect: Memorization is relative
Nicholas Carlini, Matthew Jagielski, Chiyuan Zhang, Nicolas Papernot, Andreas Terzis, and Florian Tramer · 2022
Earlier work this paper cites.
RLPrompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark · 2022
Earlier work this paper cites.
Trusted execution environments: Applications and organizational challenges
Tim Geppert, Stefan Deml, David Sturzenegger, and Nico Ebert · 2022
Earlier work this paper cites.
Do machine learning platforms provide out-of-the-box reproducibility?
Odd Erik Gundersen, Saeid Shamsaliei, and Richard Juul Isdahl · 2022
Earlier work this paper cites.
Secure multiparty computations in floating-point arithmetic
Chuan Guo, Awni Hannun, Brian Knott, Laurens van der Maaten, Mark Tygert, and Ruiyu Zhu · 2022
Earlier work this paper cites.
Evaluation gaps in machine learning practice
Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran · 2022
Earlier work this paper cites.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Earlier work this paper cites.
Intel converged security and management engine (intel CSME) security
Intel · 2022
Earlier work this paper cites.
Deduplicating training data mitigates privacy risks in language models
Nikhil Kandpal, Eric Wallace, and Colin Raffel · 2022
Earlier work this paper cites.
Caliptra: A datacenter system on a chip (SOC) root of trust (RoT)
Bryan Kelly, Andrés Lagar-Cavilla, Jeff Andersen, Prabhu Jayana, Piotr Kwidzinski, Rob Strong, John Traver, Louis Ferraro, Ishwar Agarwal, Anjana Parthasarathy, Bharat Pillilli, Vishal Soni, Marius Schilder, Sudhir Mathane, Nathan Nadarajah, and Kor Nielsen · 2022
Earlier work this paper cites.
Quality at a glance: An audit of Web-Crawled multilingual datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F P Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi · 2022
Earlier work this paper cites.
Gradient-Based constrained sampling from language models
Sachin Kumar, Biswajit Paria, and Yulia Tsvetkov · 2022
Earlier work this paper cites.
Assessing environmental impacts of nanoscale semi-conductor manufacturing from the life cycle assessment perspective
Tsai-Chi Kuo, Chien-Yun Kuo, and Liang-Wei Chen · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
A framework for deprecating datasets: Standardizing documentation, identification, and communication
Alexandra Sasha Luccioni, Frances Corry, Hamsini Sridharan, Mike Ananny, Jason Schultz, and Kate Crawford · 2022
Earlier work this paper cites.
Rethinking AI for good governance
Helen Margetts · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran · 2022
Earlier work this paper cites.
A survey of machine unlearning
Thanh Tam Nguyen, Thanh Trung Huynh, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen · 2022
Earlier work this paper cites.
Diffusion models for adversarial purification
Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, and Animashree Anandkumar · 2022
Earlier work this paper cites.
The carbon footprint of machine learning training will plateau, then shrink
David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David R So, Maud Texier, and Jeff Dean · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
Data cards: Purposeful and transparent dataset documentation for responsible AI
Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson · 2022
Earlier work this paper cites.
The fallacy of ai functionality
Inioluwa Deborah Raji, I Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst · 2022
Earlier work this paper cites.
Outsider oversight: Designing a third party audit ecosystem for AI governance
Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, and Daniel Ho · 2022
Earlier work this paper cites.
Adversarial concept erasure in kernel space
Shauli Ravfogel, Francisco Vargas, Yoav Goldberg, and Ryan Cotterell · 2022
Earlier work this paper cites.
The chip manufacturing industry: Environmental impacts and eco-efficiency analysis
Marcello Ruberti · 2022
Earlier work this paper cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Earlier work this paper cites.
Compute trends across three eras of machine learning
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos · 2022
Earlier work this paper cites.
Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction
Renee Shelby, Shalaleh Rismani, Kathryn Henne, Ajung Moon, Negar Rostamzadeh, Paul Nicholas, N’mah Yilla, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk · 2022
Earlier work this paper cites.
Toward human readable prompt tuning: Kubrick’s the shining is a good movie, and a good prompt too?
Weijia Shi, Xiaochuang Han, Hila Gonen, Ari Holtzman, Yulia Tsvetkov, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Metadata archaeology: Unearthing data subsets by leveraging training dynamics
Shoaib Ahmed Siddiqui, Nitarshan Rajkumar, Tegan Maharaj, David Krueger, and Sara Hooker · 2022
Earlier work this paper cites.
Participation is not a design fix for machine learning
Mona Sloane, Emanuel Moss, Olaitan Awomolo, and Laura Forlano · 2022
Earlier work this paper cites.
Blueprint for an AI bill of rights: Making automated systems work for the american people
The White House Office of Science and Technology Policy · 2022
Earlier work this paper cites.
Will we run out of data? Limits of LLM scaling based on human-generated data
Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho · 2022
Earlier work this paper cites.
Chain-of-Thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2022
Cited alongside, same era.
“adversarial examples” for Proof-of-Learning
Rui Zhang, Jian Liu, Yuan Ding, Zhibo Wang, Qingbiao Wu, and Kui Ren · 2022
Cited alongside, same era.
Keeping an eye on AI: Approaches to government monitoring of the AI landscape
Ada Lovelace Institute · 2023
Cited alongside, same era.
Interim report: Governing AI for humanity
Advisory Body on Artificial Intelligence · 2023
Cited alongside, same era.
BIS innovation hub expands suptech and regtech research to include monetary policy tech, February 2021
Bank for International Settlements · 2024
Closest in time.
Adrien Basdevant, Camille François, Victor Storchan, Kevin Bankston, Ayah Bdeir, Brian Behlendorf, Merouane Debbah, Sayash Kapoor, Yann LeCun, Mark Surman, Helen King-Turvey, Nathan Lambert, Stefano Maffulli, Nik Marda, Govind Shivkumar, and Justine Tunney · 2024
Closest in time.
International scientific report on the safety of advanced AI
Yohsua Bengio, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Danielle Goldfarb, Hoda Heidari, Leila Khalatbari, Shayne Longpre, Vasilios Mavroudis, Mantas Mazeika, Kwan Yee Ng, Chinasa T Okolo, Deborah Raji, Theodora Skeadas, Florian Tramèr, and Soren Mindermann · 2024
Closest in time.
About OpenSAFELY, 2024
Bennett Institute for Applied Data Science · 2024
Closest in time.
AI-generated vs human-authored texts: A multidimensional comparison
Tony Berber Sardinha · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz · 2023
Cited alongside, same era.
Self-Consuming generative models go MAD
Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard Baraniuk · 2023
Cited alongside, same era.
Nuclear arms control verification and lessons for AI treaties
Mauricio Baker · 2023
Cited alongside, same era.
Active reward learning from multiple teachers
Peter Barnett, Rachel Freedman, Justin Svegliato, and Stuart Russell · 2023
Cited alongside, same era.
AI Risk-Management standards profile for General-Purpose AI systems (GPAIS) and foundation models
Anthony Barrett, Jessica Newman, Brandie Nonnecke, Dan Hendrycks, Evan R Murphy, and Krystal Jackson · 2023
Cited alongside, same era.
LEACE: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2023
Cited alongside, same era.
Into the LAIONs den: Investigating hate in multimodal datasets
Abeba Birhane, Vinay Prabhu, Sang Han, Vishnu Naresh Boddeti, and Alexandra Sasha Luccioni · 2023
Cited alongside, same era.
Closest in time.
Mechanistic interpretability for AI safety – a review
Leonard Bereska and Efstratios Gavves · 2024
Closest in time.
Societal adaptation to advanced AI
Jamie Bernardi, Gabriel Mukobi, Hilary Greaves, Lennart Heim, and Markus Anderljung · 2024
Closest in time.
The compute divide in machine learning: A threat to academic contribution and scrutiny?
Tamay Besiroglu, Sage Andrus Bergerson, Amelia Michael, Lennart Heim, Xueyun Luo, and Neil Thompson · 2024
Closest in time.
Lessons from the trenches on reproducible evaluation of language models
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal, Wilson Y Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A Wang, Genta Indra Winata, François Yvon, and Andy Zou · 2024
Closest in time.
Does your data spark joy? Performance gains from domain upsampling at the end of training
Cody Blakeney, Mansheej Paul, Brett W Larsen, Sean Owen, and Jonathan Frankle · 2024
Closest in time.
Applying sociotechnical approaches to AI governance in practice
Miranda Bogen and Amy Winecoff · 2024
Closest in time.
Foundation model transparency reports
Rishi Bommasani, Kevin Klyman, Shayne Longpre, Betty Xiong, Sayash Kapoor, Nestor Maslej, Arvind Narayanan, and Percy Liang · 2024
Closest in time.
Location verification for AI chips
Asher Brass and Onni Aarne · 2024
Closest in time.
Commerce control List—Supplement no. 1 to part 774—category 4, 2024
Bureau of Industry and Security · 2024
Closest in time.
Overview, 2022
C2PA · 2024
Closest in time.
Stealing part of a production language model
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, David Rolnick, and Florian Tramèr · 2024
Closest in time.
Fairness in machine learning: A survey
Simon Caton and Christian Haas · 2024
Closest in time.
Evaluating predictions of model behaviour, 2024
Alan Chan · 2024
Closest in time.
Visibility into AI agents
Alan Chan, Carson Ezell, Max Kaufmann, Kevin Wei, Lewis Hammond, Herbie Bradley, Emma Bluemke, Nitarshan Rajkumar, David Krueger, Noam Kolt, Lennart Heim, and Markus Anderljung · 2024
Closest in time.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S Yu, Qiang Yang, and Xing Xie · 2024
Closest in time.
JailbreakBench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, Hamed Hassani, and Eric Wong · 2024
Closest in time.
Explainer: A sociotechnical approach to AI policy
Brian J Chen and Jacob Metcalf · 2024
Closest in time.
What is your data worth to GPT? LLM-Scale data valuation with influence functions
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, and Eric Xing · 2024
Closest in time.
Taking additional steps to address the national emergency with respect to significant malicious Cyber-Enabled activities, 01 2024
Commerce Department · 2024
Closest in time.
Restoring trust and transparency in the age of AI, 2024
Content Authenticity Initiative · 2024
Closest in time.
REGULATION OF THE EUROPEAN PARLIAMENT AND OF THE COUNCIL laying down harmonised rules on artificial intelligence and amending regulations (EC) no 300/2008, (EU) no 167/2013, (EU) no 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (artificial intelligence act), 2024
Council of the European Union · 2024
Closest in time.
Secure by design, 2023
Cybersecurity and Infrastructure Security Agency · 2024
Closest in time.
Private deep learning with MPC: A simple tutorial from scratch, 2017
Morten Dahl · 2024
Closest in time.
Towards guaranteed safe AI: A framework for ensuring robust and reliable AI systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, Alessandro Abate, Joe Halpern, Clark Barrett, Ding Zhao, Tan Zhi-Xuan, Jeannette Wing, and Joshua Tenenbaum · 2024
Closest in time.
SOPHON: Non-Fine-Tunable learning to restrain task transferability for pre-trained models
Jiangyi Deng, Shengyuan Pang, Yanjiao Chen, Liangming Xia, Yijie Bai, Haiqin Weng, and Wenyuan Xu · 2024
Closest in time.
Advanced AI evaluations at AISI: May update
Department for Science, Innovation & Technology and AI Safety Institute · 2024
Closest in time.
Toward sociotechnical AI: Mapping vulnerabilities for machine learning in context
Roel Dobbe and Anouk Wolters · 2024
Closest in time.
DiPaCo: Distributed path composition
Arthur Douillard, Qixuan Feng, Andrei A Rusu, Adhiguna Kuncoro, Yani Donchev, Rachita Chhaparia, Ionel Gog, Marc’aurelio Ranzato, Jiajun Shen, and Arthur Szlam · 2024
Closest in time.
Introducing the frontier safety framework, 2024
Anca Dragan, Helen King, and Allan Dafoe · 2024
Closest in time.
Agent AI: Surveying the horizons of multimodal interaction
Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, Bidipta Sarkar, Rohan Taori, Yusuke Noda, Demetri Terzopoulos, Yejin Choi, Katsushi Ikeuchi, Hoi Vo, Li Fei-Fei, and Jianfeng Gao · 2024
Closest in time.
Challenging the limits of benchmarking AI, 2023
Dynabench · 2024
Closest in time.
What’s in my big data?
Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hannaneh Hajishirzi, Noah A Smith, and Jesse Dodge · 2024
Closest in time.
GPTs are GPTs: Labor market impact potential of LLMs
Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock · 2024
Closest in time.
FAQs: DSA data access for researchers, 2023
EU Joint Research Centre · 2024
Closest in time.
Commission welcomes G7 leaders’ agreement on guiding principles and a code of conduct on artificial intelligence, October 2023
European Commission · 2024
Closest in time.
LLM agents can autonomously hack websites
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang · 2024
Closest in time.
Issue brief: Measuring training compute, 2024
Frontier Model Forum · 2024
Closest in time.
The ethics of advanced AI assistants
Iason Gabriel, Arianna Manzini, Geoff Keeling, Lisa Anne Hendricks, Verena Rieser, Hasan Iqbal, Nenad Tomašev, Ira Ktena, Zachary Kenton, Mikel Rodriguez, Seliem El-Sayed, Sasha Brown, Canfer Akbulut, Andrew Trask, Edward Hughes, A Stevie Bergman, Renee Shelby, Nahema Marchal, Conor Griffin, Juan Mateos-Garcia, Laura Weidinger, Winnie Street, Benjamin Lange, Alex Ingerman, Alison Lentz, Reed Enger, Andrew Barakat, Victoria Krakovna, John Oliver Siy, Zeb Kurth-Nelson, Amanda McCroskery, Vijay Bolina, Harry Law, Murray Shanahan, Lize Alberts, Borja Balle, Sarah de Haas, Yetunde Ibitoye, Allan Dafoe, Beth Goldberg, Sébastien Krier, Alexander Reese, Sims Witherspoon, Will Hawkins, Maribeth Rauh, Don Wallace, Matija Franklin, Josh A Goldstein, Joel Lehman, Michael Klenk, Shannon Vallor, Courtney Biles, Meredith Ringel Morris, Helen King, Blaise Agüera y Arcas, William Isaac, and James Manyika · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Closest in time.
Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, Daniel A Roberts, Diyi Yang, David L Donoho, and Sanmi Koyejo · 2024
Closest in time.
A false sense of safety: Unsafe information leakage in ’safe’ AI responses
David Glukhov, Ziwen Han, Ilia Shumailov, Vardan Papyan, and Nicolas Papernot · 2024
Closest in time.
Under manipulations, are some ai models harder to audit?
Augustin Godinot, Erwan Le Merrer, Gilles Trédan, Camilla Penzo, and Franois Taïani · 2024
Closest in time.
Shashwat Goel, Ameya Prabhu, Philip Torr, Ponnurangam Kumaraguru, and Amartya Sanyal · 2024
Closest in time.
Remote attestation of disaggregated machines, 2024
Google Cloud · 2024
Closest in time.
AI safety institute: Overview, November 2023
GOV.UK · 2024
Closest in time.
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation, 2024
Declan Grabb, Max Lamparth, and Nina Vasan · 2024
Closest in time.
AI regulation has its own alignment problem: The technical and institutional feasibility of disclosure, registration, licensing, and auditing
Neel Guha, Christie M Lawrence, Lindsey A Gailmard, Kit T Rodolfa, Faiz Surani, Rishi Bommasani, Inioluwa Deborah Raji, Mariano-Florentino Cuéllar, Colleen Honigsberg, Percy Liang, and Daniel E Ho · 2024
Closest in time.
Announcing azure confidential VMs with NVIDIA H100 tensor core GPUs in preview, 2023
Krishnaprasad Hande · 2024
Closest in time.
Synthetic data in AI: Challenges, applications, and ethical implications
Shuang Hao, Wenfeng Han, Tao Jiang, Yiping Li, Haonan Wu, Chunlin Zhong, Zhangjun Zhou, and He Tang · 2024
Closest in time.
Mapping the chip smuggling pipeline and improving export control compliance
Barath Harithas · 2024
Closest in time.
A trusted AI compute cluster for AI verification and evaluation, 2024
Lennart Heim · 2024
Closest in time.
Governing through the cloud: The intermediary role of compute providers in AI regulation
Lennart Heim, Tim Fist, Janet Egan, Sihao Huang, Stephen Zekany, Robert Trager, Michael Osborne, and Noa Zilberman · 2024
Closest in time.
Algorithmic progress in language models
Anson Ho, Tamay Besiroglu, Ege Erdil, David Owen, Robi Rahman, Zifan Carl Guo, David Atkinson, Neil Thompson, and Jaime Sevilla · 2024
Closest in time.
Curiosity-driven red-teaming for large language models
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang, Aldo Pareja, James Glass, Akash Srivastava, and Pulkit Agrawal · 2024
Closest in time.
On the limitations of compute thresholds as a governance strategy
Sara Hooker · 2024
Closest in time.
Governance of artificial intelligence (AI)
House of Commons Science, Innovation and Technology Committee · 2024
Closest in time.
Safe LoRA: The silver lining of reducing safety risks when fine-tuning large language models
Chia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen, Chia-Mu Yu, and Chun-Ying Huang · 2024
Closest in time.
Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Hp · 2024
Closest in time.
The underground network sneaking nvidia chips into china, July 2024
Raffaele Huang · 2024
Closest in time.
Sleeper agents: Training deceptive LLMs that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
Gated datasets, 2024
Hugging Face · 2024
Closest in time.
Beyond static AI evaluations: Advancing human interaction evaluations for LLM harms and risks
Lujain Ibrahim, Saffron Huang, Lama Ahmad, and Markus Anderljung · 2024
Closest in time.
Remote ATtestation ProcedureS (rats), 2024
IETF Datatracker · 2024
Closest in time.
Uncovering deceptive tendencies in language models: A simulated company AI assistant
Olli Järviniemi and Evan Hubinger · 2024
Closest in time.
Enabling research with publicly accessible platform data: Early DSA compliance issues and suggestions for improvement
Julian Jaursch, Jakob Ohme, and Ulrike Klinger · 2024
Closest in time.
Investigating data contamination for pre-training language models
Minhao Jiang, Ken Ziyu Liu, Ming Zhong, Rylan Schaeffer, Siru Ouyang, Jiawei Han, and Sanmi Koyejo · 2024
Closest in time.
Code competition FAQ, 2024
kaggle · 2024
Closest in time.
Governing AI agents, 2024
Noam Kolt · 2024
Closest in time.
Responsible reporting for frontier AI development
Noam Kolt, Markus Anderljung, Joslyn Barnhart, Asher Brass, Kevin Esvelt, Gillian K Hadfield, Lennart Heim, Mikel Rodriguez, Jonas B Sandbrink, and Thomas Woodside · 2024
Closest in time.
Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090
Gabriel Kulp, Daniel Gonzales, Everett Smith, Lennart Heim, Prateek Puri, Michael J D Vermeer, and Zev Winkelman · 2024
Closest in time.
The value chain of general-purpose AI, 2023
Sabrina Küspert, Nicolas Moës, and Connor Dunlop · 2024
Closest in time.
Frontier AI ethics: Anticipating and evaluating the societal impacts of generative agents
Seth Lazar · 2024
Closest in time.
Talkin’ ’bout AI generation: Copyright and the Generative-AI supply chain (the short version)
Katherine Lee, A Feder Cooper, and James Grimmelmann · 2024
Closest in time.
David Leslie, Cami Rincon, Morgan Briggs, Antonella Perini, Smera Jayadeva, Ann Borda, S J Bennett, Christopher Burr, Mhairi Aitken, Michael Katell, Claudia Fischer, Janis Wong, and Ismael Kherroubi Garcia · 2024
Closest in time.
The ai regulatory pyramid: A taxonomy & analysis of the emerging toolbox in the global race for the regulation and governance of artificial intelligence
Orly Lobel · 2024
Closest in time.
Energy star ratings for AI models, 2024
Sasha Luccioni · 2024
Closest in time.
Eight methods to evaluate robust unlearning in LLMs
Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell · 2024
Closest in time.
Quantifying variance in evaluation benchmarks
Lovish Madaan, Aaditya K Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes · 2024
Closest in time.
Generative AI has a visual plagiarism problem, January 2024
Gary Marcus and Reid Southen · 2024
Closest in time.
The AI index 2024 annual report
Nestor Maslej, Loredana Fattorini, Raymond Perrault, Vanessa Parli, Anka Reuel, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Russell Wald, and Jack Clark · 2024
Closest in time.
HarmBench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
Ethical ai governance: Methods for evaluating trustworthy ai
Louise McCormack and Malika Bendechache · 2024
Closest in time.
Inadequacies of large language model benchmarks in the era of generative artificial intelligence
Timothy R McIntosh, Teo Susnjak, Tong Liu, Paul Watters, and Malka N Halgamuge · 2024
Closest in time.
Machine learning with confidential computing: A systematization of knowledge
Fan Mo, Zahra Tarkhani, and Hamed Haddadi · 2024
Closest in time.
Researcher access to social media data: Lessons from clinical trial data sharing
Christopher Morten, Gabriel Nicholas, and Salome Viljoen · 2024
Closest in time.
Applications, 2024
MULTIBEAM · 2024
Closest in time.
AI safety is not a model property, 2024
Arvind Narayanan and Sayash Kapoor · 2024
Closest in time.
Model alignment protects against accidental harms, not intentional ones, 2023
Arvind Narayanan, Sayash Kapoor, and Seth Lazar · 2024
Closest in time.
Biden-Harris administration announces new NIST public working group on AI, June 2023
National Institute of Standards and Technology (NIST) · 2024
Closest in time.
Biden-Harris administration announces First-Ever consortium dedicated to AI safety, February 2024
National Institute of Standards and Technology (NIST) · 2024
Closest in time.
Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models
Sella Nevo, Dan Lahav, Ajay Karpur, Yogev Bar-On, Henry Alexander Bradley, and Jeff Alstott · 2024
Closest in time.
Secure data environment, 2024
NHS Research SDE Network · 2024
Closest in time.
Grounding ai policy: Towards researcher access to ai usage data
Gabriel Nicholas · 2024
Closest in time.
NIST AIRC - Crosswalk Documents, 2023a
NIST · 2024
Closest in time.
NIST AI RMF playbook, 2023b
NIST · 2024
Closest in time.
Artificial intelligence risk management framework: Generative artificial intelligence profile, 2024
NIST · 2024
Closest in time.
AI governance needs sociotechnical expertise: Why the humanities and social sciences are critical to government efforts
Serena Oduro and Tamara Kneese · 2024
Closest in time.
OECD AI incidents monitor, 2024
OECD.AI Policy Observatory · 2024
Closest in time.
Towards AI accountability infrastructure: Gaps and opportunities in AI audit tooling
Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji · 2024
Closest in time.
Interpretability dreams, 2023
Chris Olah · 2024
Closest in time.
Preparedness, 2024a
OpenAI · 2024
Closest in time.
Reimagining secure infrastructure for advanced AI, 2024b
OpenAI · 2024
Closest in time.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, C J Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph · 2024
Closest in time.
AI deception: A survey of examples, risks, and potential solutions
Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2024
Closest in time.
PAI’s guidance for safe foundation model deployment, October 2023
Partnership on AI · 2024
Closest in time.
China AI & semiconductors rise: US sanctions have failed, 2023
Dylan Patel · 2024
Closest in time.
Energy and emissions of machine learning on smartphones vs. the cloud
David Patterson, Jeffrey M Gilbert, Marco Gruteser, Efren Robles, Krishna Sekar, Yong Wei, and Tenghui Zhu · 2024
Closest in time.
The FineWeb datasets: Decanting the web for the finest text data at scale
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Closest in time.
Navigating the safety landscape: Measuring risks in finetuning large language models
Shengyun Peng, Pin-Yu Chen, Matthew Hull, and Duen Horng Chau · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe, and Toby Shevlane · 2024
Closest in time.
LLM self defense: By self examination, LLMs know they are being tricked
Mansi Phute, Alec Helbling, Matthew Daniel Hull, Shengyun Peng, Sebastian Szyller, Cory Cornelius, and Duen Horng Chau · 2024
Closest in time.
The EU’s AI act is barreling toward AI standards that do not exist, 2023
Hadrien Pouget · 2024
Closest in time.
Recite, reconstruct, recollect: Memorization in LMs as a multifaceted phenomenon
Usvsn Sai Prashanth, Alvin Deng, Kyle O’Brien, S V Jyothir, Mohammad Aflah Khan, Jaydeep Borkar, Christopher A Choquette-Choo, Jacob Ray Fuehne, Stella Biderman, Tracy Ke, Katherine Lee, and Naomi Saphra · 2024
Closest in time.
Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts - analysis of the final compromise text with a view to agreement
Presidency of the Council of the European Union · 2024
Closest in time.
Project oak, 2024
Project Oak · 2024
Closest in time.
Insight into the U.S. semiconductor export controls update, 2023
William Alan Reinsch, Matthew Schleich, and Thibault Denamiel · 2024
Closest in time.
Generative AI needs adaptive governance
Anka Reuel and Trond Arne Undheim · 2024
Closest in time.
ImageNot: A contrast with ImageNet preserves model rankings
Olawale Salaudeen and Moritz Hardt · 2024
Closest in time.
Computing power and the governance of artificial intelligence
Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Julian Hazell, Cullen O’Keefe, Gillian K Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F Trager, Shahar Avin, Adrian Weller, Yoshua Bengio, and Diane Coyle · 2024
Closest in time.
Why has predicting downstream capabilities of frontier AI models with scale remained elusive?
Rylan Schaeffer, Hailey Schoelkopf, Brando Miranda, Gabriel Mukobi, Varun Madan, Adam Ibrahim, Herbie Bradley, Stella Biderman, and Sanmi Koyejo · 2024
Closest in time.
Future-Proofing frontier AI regulation: Projecting future compute for frontier AI models
Paul Scharre · 2024
Closest in time.
Large language models can strategically deceive their users when put under pressure
Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn · 2024
Closest in time.
From principles to rules: A regulatory approach for frontier AI
Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler, and Ben Garfinkel · 2024
Closest in time.
AI incident reporting: Addressing a gap in the UK’s regulation of AI
Tommy Shaffer Shane · 2024
Closest in time.
ShareGPT, 2022
ShareGPT · 2024
Closest in time.
MUSE: Machine Unlearning Six-Way Evaluation for Language Models
Weijia Shi, Jaechan Lee, Yangsibo Huang, Sadhika Malladi, Jieyu Zhao, Ari Holtzman, Daogao Liu, Luke Zettlemoyer, Noah A Smith, and Chiyuan Zhang · 2024
Closest in time.
One model many scores: Using multiverse analysis to prevent fairness hhacking and evaluate the influence of model design decisions
Jan Simson, Florian Pfisterer, and Christoph Kern · 2024
Closest in time.
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2024
Closest in time.
A StrongREJECT for empty jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer · 2024
Closest in time.
Verifiable evaluations of machine learning models using zkSNARKs
Tobin South, Alexander Camuto, Shrey Jain, Shayla Nguyen, Robert Mahari, Christian Paquin, Jason Morton, and Alex ’sandy’ · 2024
Closest in time.
Detecting AI fingerprints: A guide to watermarking and beyond
Siddarth Srinivasan · 2024
Closest in time.
Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap
Saurabh Srivastava, Annarose M B, Anto P V, Shashank Menon, Ajay Sukumar, Adwaith Samod T, Alan Philipose, Stevin Prince, and Sooraj Thomas · 2024
Closest in time.
Mission Impossible: A statistical perspective on jailbreaking LLMs
Jingtong Su, Julia Kempe, and Karen Ullrich · 2024
Closest in time.
zkLLM: Zero knowledge proofs for large language models
Haochen Sun, Jason Li, and Hongyang Zhang · 2024
Closest in time.
Tamper-resistant safeguards for open-weight LLMs
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, and Mantas Mazeika · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C Daniel Freeman, Theodore R Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
WildChat, 2024
The Allen Institute for Artificial Intelligence · 2024
Closest in time.
FACT SHEET: President biden issues executive order on safe, secure, and trustworthy artificial intelligence, October 2023b
The White House · 2024
Closest in time.
Building safe AI: A tutorial for encrypted deep learning, 2017
Andrew Trask · 2024
Closest in time.
How to audit an AI model owned by someone else (part 1), 2023
Andrew Trask, Akshay Sukumar, Antti Kalliokoski, Bennett Farkas, Callis Ezenwaka, Carmen Popa, Curtis Mitchell, Dylan Hrebenach, George-Cristian Muraru, Ionesio Junior, Irina Bejan, Ishan Mishra, Ivoline Ngong, Jack Bandy, Jess Stahl, Julian Cardonnet, Kellye Trask, Kellye Trask, Khoa Nguyen, Kien Dang, Koen van der Veen, Kyoko Eng, Lacey Strahm, Laura Ayre, Madhava Jay, Oleksandr Lytvyn, Osam Kyemenu-Sarsah, Peter Chung, Peter Smith, S Rasswanth, Ronnie Falcon, Shubham Gupta, Stephen Gabriel, Teo Milea, Theresa Thoraldson, Thiago Porto, Tudor Cebere, Yash Gorana, and Zarreen Reza · 2024
Closest in time.
Vishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma, Philip H S Torr, Adel Bibi, Samuel Albanie, and Matthias Bethge · 2024
Closest in time.
Inspect, 2024
UK AI Safety Institute · 2024
Closest in time.
£300 million to launch first phase of new AI research resource, 2023
UK Research and Innovation · 2024
Closest in time.
Artificial intelligence: Accountability policy report
U.S. National Telecommunications and Information Administration · 2024
Closest in time.
Aya model: An instruction finetuned Open-Access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker · 2024
Closest in time.
A taxonomy of systemic risks from general-purpose ai
Risto Uuk, Carlos Ignacio Gutierrez, Daniel Guppy, Lode Lauwaert, Atoosa Kasirzadeh, Lucia Velasco, Peter Slattery, and Carina Prunkl · 2024
Closest in time.
SGX.Fail, 2022
Stephan van Schaik, Adam Batori, Alex Seto, Bader AlBassam, Christina Garman, Thomas Yurek, Andrew Miller, Daniel Genkin, Eyal Ronen, and Yuval Yarom · 2024
Closest in time.
Adversarial machine learning: A taxonomy and terminology of attacks and mitigations
Apostol Vassilev, Alina Oprea, Alie Fordyce, and Hyrum Anderson · 2024
Closest in time.
Trustless audits without revealing data or models
Suppakit Waiwitlikhit, Ion Stoica, Yi Sun, Tatsunori Hashimoto, and Daniel Kang · 2024
Closest in time.
Proving membership in LLM pretraining data via data watermarks
Johnny Tian-Zheng Wei, Ryan Yixiang Wang, and Robin Jia · 2024
Closest in time.
STAR: SocioTechnical approach to red teaming language models
Laura Weidinger, John Mellor, Bernat Guillen Pegueroles, Nahema Marchal, Ravin Kumar, Kristian Lum, Canfer Akbulut, Mark Diaz, Stevie Bergman, Mikel Rodriguez, Verena Rieser, and William Isaac · 2024
Closest in time.
joint and several liability, 2023
Wex Definitions Team · 2024
Closest in time.
Livebench: A challenging, contamination-free llm benchmark, June 2024
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblum · 2024
Closest in time.
Instructional fingerprinting of large language models
Jiashu Xu, Fei Wang, Mingyu Derek Ma, Pang Wei Koh, Chaowei Xiao, and Muhao Chen · 2024
Closest in time.
Fairproof : Confidential and certifiable fairness for neural networks
Chhavi Yadav, Amrita Roy Chowdhury, Dan Boneh, and Kamalika Chaudhuri · 2024
Closest in time.
AI risk categorization decoded (AIR 2024): From government regulations to corporate policies
Yi Zeng, Kevin Klyman, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, and Bo Li · 2024
Closest in time.
Removing RLHF protections in GPT-4 via Fine-Tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang · 2024
Closest in time.
A careful examination of large language model performance on grade school arithmetic
Hugh Zhang, Jeff Da, Dean Lee, Vaughn Robinson, Catherine Wu, Will Song, Tiffany Zhao, Pranav Raja, Dylan Slack, Qin Lyu, Sean Hendryx, Russell Kaplan, Michele Lunati, and Summer Yue · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks
Andy Zhou, Bo Li, and Haohan Wang · 2024
Closest in time.
Multi-agent risks from advanced AI
Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak, The Anh Han, Edward Hughes, Vojtěch Kovařík, Jan Kulveit, Joel Z. Leibo, Caspar Oesterheld, Christian Schroeder de Witt, Nisarg Shah, Michael Wellman, Paolo Bova, Theodor Cimpeanu, Carson Ezell, Quentin Feuillade-Montixi, Matija Franklin, Esben Kran, Igor Krawczuk, Max Lamparth, Niklas Lauffer, Alexander Meinke, Sumeet Motwani, Anka Reuel, Vincent Conitzer, Michael Dennis, Iason Gabriel, Adam Gleave, Gillian Hadfield, Nika Haghtalab, Atoosa Kasirzadeh, Sébastien Krier, Kate Larson, Joel Lehman, David C. Parkes, Georgios Piliouras, and Iyad Rahwan · 2025
Closest in time.
Data centre water consumption
David Mytton · 2059
Closest in time.
Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market For Digital Services and amending Directive 2000/31/EC (Digital Services Act) (Text with EEA relevance), October 2022
Digital Services Act · 2065
Closest in time.
Scheduling fair resource allocation policies for cloud computing through flow control
Stavros Souravlas and Stefanos Katsavounis · 2079
Closest in time.