The Role of Cooperation in Responsible AI Development, jul 2019
Original
Amanda Askell, Miles Brundage, and Gillian Hadfield · 1907
Earlier work this paper cites.
Emergent Tool Use From Multi-Agent Autocurricula
Original
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Philosophy and the Problem of Value
A. P. Brogan · 1952
Earlier work this paper cites.
The Anatomy of Market Failure
Francis M. Bator · 1958
Earlier work this paper cites.
The Problem of Social Cost
R. H. Coase · 1960
Earlier work this paper cites.
Externality
James M. Buchanan and Wm. Craig Stubblebine · 1962
Earlier work this paper cites.
Market Theory and the Price System
Israel M. Kirzner · 1963
Earlier work this paper cites.
Logic and conversation
Herbert P Grice · 1975
Earlier work this paper cites.
Computer science as empirical inquiry: symbols and search
Allen Newell and Herbert A. Simon · 1976
Earlier work this paper cites.
The Strategy of Conflict: With a New Preface by the Author
Thomas C. Schelling · 1981
Earlier work this paper cites.
Taxonomies of human performance: The description of human tasks
Edwin A Fleishman, Marilyn K Quaintance, and Laurie A Broedling · 1984
Earlier work this paper cites.
Inefficiency of Nash Equilibria
Pradeep Dubey · 1986
Earlier work this paper cites.
Learnability and the Vapnik-Chervonenkis dimension
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1989
Earlier work this paper cites.
Example of exact trade-offs in linear controller design
Craig Barratt and Stephen Boyd · 1989
Earlier work this paper cites.
Statistical mechanics of learning from examples
Hyunjune Sebastian Seung, Haim Sompolinsky, and Naftali Tishby · 1992
Earlier work this paper cites.
Trade-offs in linear control system design: A practical example
Bart De Moor, Johan David, Joos Vandewalle, Maarten De Moor, and Daniel Berckmans · 1992
Earlier work this paper cites.
Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries
Shalom H Schwartz · 1992
Earlier work this paper cites.
Human cognitive abilities: A survey of factor-analytic studies
John B Carroll · 1993
Earlier work this paper cites.
The statistical mechanics of learning a rule
Timothy LH Watkin, Albrecht Rau, and Michael Biehl · 1993
Earlier work this paper cites.
A universal theorem on learning curves
Shun-Ichi Amari · 1993
Earlier work this paper cites.
Rigorous learning curve bounds from statistical mechanics
David Haussler, H Sebastian Seung, Michael Kearns, and Naftali Tishby · 1994
Earlier work this paper cites.
Are there universal aspects in the structure and contents of human values?
Shalom H Schwartz · 1994
Earlier work this paper cites.
A new theory of equilibrium selection for games with complete information
John C. Harsanyi · 1995
Earlier work this paper cites.
Learning to parse database queries using inductive logic programming
John M Zelle and Raymond J Mooney · 1996
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Earlier work this paper cites.
Does automation bias decision-making?
Linda J Skitka, Kathleen L Mosier, and Mark Burdick · 1999
Earlier work this paper cites.
Reputation helps solve the ‘tragedy of the commons’
Manfred Milinski, Dirk Semmann, and Hans-Jürgen Krambeck · 2002
Earlier work this paper cites.
A Primer in BERTology: What we know about how BERT works, nov 2020
Original
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2002
Earlier work this paper cites.
The World Economy: Historical Statistics
Angus Maddison · 2004
Earlier work this paper cites.
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims, apr 2020
Original
Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, et al · 2004
Earlier work this paper cites.
Selfish Routing and The Price of Anarchy
Tim Roughgarden · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
Cooperation, Punishment, and the Evolution of Human Institutions
Joseph Henrich · 2006
Earlier work this paper cites.
Adaptive mechanism design
David Pardoe, Peter Stone, Maytal Saar-Tsechansky, and Kerem Tomak · 2006
Earlier work this paper cites.
Code and Other Laws of Cyberspace, Version 2.0
Lawrence Lessig · 2006
Earlier work this paper cites.
Algorithmic game theory
Noam Nisan, editor · 2007
Earlier work this paper cites.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Yoav Shoham and Kevin Leyton-Brown · 2008
Earlier work this paper cites.
Statistical mechanics
Kerson Huang · 2008
Earlier work this paper cites.
Aligning AI with shared human values
Original
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2008
Earlier work this paper cites.
Reading tea leaves: How humans interpret topic models
Jonathan Chang, Sean Gerrish, Chong Wang, Jordan Boyd-Graber, and David Blei · 2009
Earlier work this paper cites.
Measuring massive multitask language understanding
Original
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
Verification and control of hybrid systems: a symbolic approach
Paulo Tabuada · 2009
Earlier work this paper cites.
The verified software initiative: A manifesto
C.A.R. Hoare, Jayadev Misra, Gary T. Leavens, and Natarajan Shankar · 2009
Earlier work this paper cites.
Coordinated Punishment of Defectors Sustains Cooperation and Can Proliferate When Rare
Robert Boyd, Herbert Gintis, and Samuel Bowles · 2010
Earlier work this paper cites.
Findings Regarding the Market Events of May 6, 2010: Report of the Staffs of the CFTC and SEC to the Joint Advisory Committee on Emerging Regulatory Issues
CFTC and SEC · 2010
Earlier work this paper cites.
The growing gap between emerging technologies and the law
Gary E Marchant · 2011
Earlier work this paper cites.
The Cattell-Horn-Carroll model of intelligence
W Joel Schneider and Kevin S McGrew · 2012
Earlier work this paper cites.
Critical truths about power laws
Michael PH Stumpf and Mason A Porter · 2012
Earlier work this paper cites.
The communicative function of ambiguity in language
Steven T Piantadosi, Harry Tily, and Edward Gibson · 2012
Earlier work this paper cites.
Open Problems in Cooperative AI, dec 2020
Original
Allan Dafoe, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R. McKee, Joel Z. Leibo, Kate Larson, and Thore Graepel · 2012
Earlier work this paper cites.
Before we knew it: an empirical study of zero-day attacks in the real world
Leyla Bilge and Tudor Dumitraş · 2012
Earlier work this paper cites.
Poisoning attacks against support vector machines
Original
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2012
Earlier work this paper cites.
The dangers of surveillance
Neil M Richards · 2012
Earlier work this paper cites.
Killer drones: The ‘silver bullet’ of democratic warfare?
Frank Sauer and Niklas Schörnig · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Original
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Cooperation In Strategic Games Revisited
Adam Kalai and Ehud Kalai · 2013
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Systems theory as the foundation for understanding systems
Kevin MacG Adams, Patrick T Hester, Joseph M Bradley, Thomas J Meyers, and Charles B Keating · 2014
Earlier work this paper cites.
The limitations of standardized science tests as benchmarks for artificial intelligence research: Position paper
Original
Ernest Davis · 2014
Earlier work this paper cites.
Freeze-thaw Bayesian optimization
Original
Kevin Swersky, Jasper Snoek, and Ryan Prescott Adams · 2014
Earlier work this paper cites.
Self-organization in complex systems as decision making
Vyacheslav I Yukalov and Didier Sornette · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Original
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Goal recognition design
Sarah Keren, Avigdor Gal, and Erez Karpas · 2014
Earlier work this paper cites.
Liberating Practice from Philosophy – A Critical Examination of Values-Based Practice and Its Underpinnings
Elselijn Kingma and Natalie Banner · 2014
Earlier work this paper cites.
Evading the censors: Critical journalism in authoritarian states
Mikal Hem · 2014
Earlier work this paper cites.
The global decline of the labor share
Loukas Karabarbounis and Brent Neiman · 2014
Earlier work this paper cites.
Aligning superintelligence with human interests: A technical research agenda
Nate Soares and Benja Fallenstein · 2014
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Original
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John Hopcroft · 2015
Earlier work this paper cites.
Maximal Cooperation in Repeated Games on Social Networks
Catherine Moon and Vincent Conitzer · 2015
Earlier work this paper cites.
Approval voting and incentives in crowdsourcing
Nihar Shah, Dengyong Zhou, and Yuval Peres · 2015
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Earlier work this paper cites.
Why are there still so many jobs? The history and future of workplace automation
David H Autor · 2015
Earlier work this paper cites.
Down the (white) rabbit hole: The extreme right and online recommender systems
Derek O’Callaghan, Derek Greene, Maura Conway, Joe Carthy, and Pádraig Cunningham · 2015
Earlier work this paper cites.
Research priorities for robust and beneficial artificial intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Original
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Engineering a safer world: Systems thinking applied to safety
Nancy G Leveson · 2016
Earlier work this paper cites.
The Frame Problem
Murray Shanahan · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
“Why Should I Trust You?” Explaining the Predictions of Any Classifier
Marco Tulio Ribeiro · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Andreas Veit, Michael J Wilber, and Serge Belongie · 2016
Earlier work this paper cites.
Transferability in machine learning: from phenomena to black-box attacks using adversarial samples
Original
Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow · 2016
Earlier work this paper cites.
Saying ‘no!’to lethal autonomous targeting
Noel Sharkey · 2016
Earlier work this paper cites.
Autonomous weapons and operational risk, 2016
Paul Scharre · 2016
Earlier work this paper cites.
Meaningful human control, artificial intelligence and autonomous weapons
Heather M Roff and Richard Moyes · 2016
Earlier work this paper cites.
Racing to the precipice: a model of artificial intelligence development
Stuart Armstrong, Nick Bostrom, and Carl Shulman · 2016
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Original
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
AI safety gridworlds
Original
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Earlier work this paper cites.
The Off-Switch Game
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Original
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
The legal protection of crowdworkers: four avenues for workers’ rights in the virtual realm
Jeremias Prassl and Martin Risak · 2017
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions
Scott M. Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie J. Cai, James Wexler, Fernanda B. Viégas, and Rory Sayres · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Original
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Earlier work this paper cites.
Towards A Rigorous Science of Interpretable Machine Learning, mar 2017
Original
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Right for the right reasons: Training differentiable models by constraining their explanations
Original
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez · 2017
Earlier work this paper cites.
Cyclegan, a master of steganography
Original
Casey Chu, Andrey Zhmoginov, and Mark Sandler · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Original
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Original
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
The EU general data protection regulation (GDPR)
Paul Voigt and Axel Von dem Bussche · 2017
Earlier work this paper cites.
ISO/IEC 38505-1:2017 – Information technology – Governance of IT – Governance of data, 2017
ISO, International Organization for Standardization · 2017
Earlier work this paper cites.
An FDA for algorithms
Andrew Tutt · 2017
Earlier work this paper cites.
The October 2014 United States Treasury bond flash crash and the contributory effect of mini flash crashes
Zachary S Levine, Scott A Hale, and Luciano Floridi · 2017
Earlier work this paper cites.
Clarifying “AI Alignment”, April 7 2018
Paul Christiano · 2018
Earlier work this paper cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Towards understanding learning representations: To what extent do different neural networks learn the same representation
Liwei Wang, Lunjia Hu, Jiayuan Gu, Zhiqiang Hu, Yue Wu, Kun He, and John Hopcroft · 2018
Earlier work this paper cites.
Probability on graphs: random processes on graphs and lattices , volume 8
Geoffrey Grimmett · 2018
Earlier work this paper cites.
Learning with Opponent-learning Awareness
Jakob Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2018
Earlier work this paper cites.
Agents and devices: A relative definition of agency
Original
Laurent Orseau, Simon McGregor McGill, and Shane Legg · 2018
Earlier work this paper cites.
Inequity Aversion Improves Cooperation in Intertemporal Social Dilemmas
Edward Hughes, Joel Z Leibo, Matthew Phillips, Karl Tuyls, Edgar Dueñez Guzman, Antonio García Castañeda, Iain Dunning, Tina Zhu, Kevin McKee, Raphael Koster, Heather Roff, and Thore Graepel · 2018
Earlier work this paper cites.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Kim, Ian En-Hsu Yen, and Pradeep K Ravikumar · 2018
Earlier work this paper cites.
Virtual adversarial training: a regularization method for supervised and semi-supervised learning
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii · 2018
Earlier work this paper cites.
The glass ceiling in NLP
Natalie Schluter · 2018
Earlier work this paper cites.
AI safety via debate
Original
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Original
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Original
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Explaining Explanations: An Overview of Interpretability of Machine Learning
Leilani H. Gilpin, David Bau, Ben Z. Yuan, Ayesha Bajwa, Michael A. Specter, and Lalana Kagal · 2018
Earlier work this paper cites.
Sanity Checks for Saliency Maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian J. Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Anchors: High-Precision Model-Agnostic Explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Earlier work this paper cites.
Women also snowboard: Overcoming bias in captioning models
Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach · 2018
Earlier work this paper cites.
Categorizing variants of Goodhart’s Law
Original
David Manheim and Scott Garrabrant · 2018
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Original
Guillaume Alain and Yoshua Bengio · 2018
Earlier work this paper cites.
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al · 2018
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Original
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, et al · 2018
Earlier work this paper cites.
Security breach at target
Miloslova Plachkinova and Chris Maurer · 2018
Earlier work this paper cites.
Google AI Principles, 2018
Google AI · 2018
Earlier work this paper cites.
IEEE Standard for Enterprise Architecture Description (EAD) Version 2, 2018
IEEE, Institute of Electrical and Electronics Engineers · 2018
Earlier work this paper cites.
AI governance: a research agenda
Allan Dafoe · 2018
Earlier work this paper cites.
Computational power and the social impact of artificial intelligence
Original
Tim Hwang · 2018
Earlier work this paper cites.
Ethical challenges in data-driven dialogue systems
Peter Henderson, Koustuv Sinha, Nicolas Angelard-Gontier, Nan Rosemary Ke, Genevieve Fried, Ryan Lowe, and Joelle Pineau · 2018
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Original
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Original
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance
Original
R Thomas McCoy, Junghyun Min, and Tal Linzen · 2019
Earlier work this paper cites.
Preferences implicit in the state of the world
Original
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
A research agenda: Dynamic models to defend against correlated attacks
Original
Ian Goodfellow · 2019
Earlier work this paper cites.
On Collusion, 2019
Vitalik Buterin · 2019
Earlier work this paper cites.
Human compatible: Artificial intelligence and the problem of control
Stuart Russell · 2019
Earlier work this paper cites.
Social Influence As Intrinsic Motivation for Multi-agent Deep Reinforcement Learning
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Çaglar Gülçehre, Pedro A. Ortega, DJ Strouse, Joel Z. Leibo, and Nando de Freitas · 2019
Earlier work this paper cites.
Current Work in AI Alignment
Paul Christiano · 2019
Earlier work this paper cites.
Gaining Free or Low-Cost Interpretability with Interpretable Partial Substitute
Tong Wang · 2019
Earlier work this paper cites.
Predicting supply chain risks using machine learning: The trade-off between performance and interpretability
George Baryannis, Samir Dani, and Grigoris Antoniou · 2019
Earlier work this paper cites.
Robustness May Be at Odds with Accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2019
Earlier work this paper cites.
Theoretically Principled Trade-off between Robustness and Accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan · 2019
Earlier work this paper cites.
Evaluating gender bias in machine translation
Original
Gabriel Stanovsky, Noah A Smith, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Original
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher · 2019
Earlier work this paper cites.
Invariant risk minimization
Original
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Earlier work this paper cites.
Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization
Original
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang · 2019
Earlier work this paper cites.
Relaxed adversarial training for inner alignment
Evan Hubinger · 2019
Earlier work this paper cites.
Harnessing the vulnerability of latent layers in adversarially trained models
Nupur Kumari, Mayank Singh, Abhishek Sinha, Harshitha Machiraju, Balaji Krishnamurthy, and Vineeth N Balasubramanian · 2019
Earlier work this paper cites.
Adversarial training and robustness for multiple perturbations
Florian Tramer and Dan Boneh · 2019
Earlier work this paper cites.
What is one grain of sand in the desert? Analyzing individual neurons in deep NLP models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass · 2019
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Earlier work this paper cites.
Inherent Disagreements in Human Textual Inferences
Ellie Pavlick and Tom Kwiatkowski · 2019
Earlier work this paper cites.
The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Original
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao · 2019
Earlier work this paper cites.
Abs: Scanning neural networks for back-doors by artificial brain stimulation
Yingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma, Yousra Aafer, and Xiangyu Zhang · 2019
Earlier work this paper cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2019
Earlier work this paper cites.
Deep fakes: A looming challenge for privacy, democracy, and national security
Bobby Chesney and Danielle Citron · 2019
Earlier work this paper cites.
Social Engineering Attacks: A Survey
F. Salahdine and N. Kaabouch · 2019
Earlier work this paper cites.
The global expansion of AI surveillance
Steven Feldstein · 2019
Earlier work this paper cites.
Artificial intelligence and its implications for income distribution and unemployment
Anton Korinek and Joseph E Stiglitz · 2019
Earlier work this paper cites.
Work of the Past, Work of the Future
David H Autor · 2019
Earlier work this paper cites.
The Platformisation of Labor and Society
Antonio A. Casilli and Julian Posada · 2019
Earlier work this paper cites.
Recommendation of the Council on Artificial Intelligence
OECD · 2019
Earlier work this paper cites.
Ethics Guidelines for Trustworthy AI
AI-HLEG, High-Level Expert Group on Artificial Intelligence · 2019
Earlier work this paper cites.
Guidelines for Secure AI System Development, 2019
National Cyber Security Center · 2019
Earlier work this paper cites.
Cross-section of mini flash crashes and their detection by a state-space approach
Chyng Wen Tee and Christopher Hian Ann Ting · 2019
Earlier work this paper cites.
Common challenges in combating cybercrime
Eurojust and Europol · 2019
Earlier work this paper cites.
AI research considerations for human existential safety (ARCHES)
Original
Andrew Critch and David Krueger · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Original
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
On the value of out-of-distribution testing: An example of goodhart’s law
Damien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha, Christopher Kanan, and Anton Van Den Hengel · 2020
Earlier work this paper cites.
An information theoretic view on selecting linguistic probes
Original
Zining Zhu and Frank Rudzicz · 2020
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Earlier work this paper cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Earlier work this paper cites.
Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
Thao Nguyen, Maithra Raghu, and Simon Kornblith · 2020
Earlier work this paper cites.
Fair is better than sensational: Man is to doctor as woman is to doctor
Malvina Nissim, Rik van Noord, and Rob van der Goot · 2020
Earlier work this paper cites.
Measuring the algorithmic efficiency of neural networks
Original
Danny Hernandez and Tom B Brown · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Original
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Earlier work this paper cites.
Abductive Commonsense Reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen-tau Yih, and Yejin Choi · 2020
Earlier work this paper cites.
Underspecification Presents Challenges for Credibility in Modern Machine Learning
Alexander Nicholas D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jon Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Shaobo Hou, Neil Houlsby, Ghassen Jerfel, Alan Karthikesalingam, Mario Lučić, Yian Ma, Cory McLean, Diana Mincu, …, and D. Sculley · 2020
Earlier work this paper cites.
Avoiding side effects by considering future tasks
Victoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic, and Shane Legg · 2020
Earlier work this paper cites.
Benefits of assistance over reward learning, 2020
Rohin Shah, Pedro Freire, Neel Alex, Rachel Freedman, Dmitrii Krasheninnikov, Lawrence Chan, Michael D Dennis, Pieter Abbeel, Anca Dragan, and Stuart Russell · 2020
Earlier work this paper cites.
Conservative agency via attainable utility preservation
Alexander Matt Turner, Dylan Hadfield-Menell, and Prasad Tadepalli · 2020
Earlier work this paper cites.
Active reinforcement learning: Observing rewards at a cost
Original
David Krueger, Jan Leike, Owain Evans, and John Salvatier · 2020
Earlier work this paper cites.
Open questions in creating safe open-ended AI: tensions between control and creativity
Adrien Ecoffet, Jeff Clune, and Joel Lehman · 2020
Earlier work this paper cites.
Algorithmic Pricing and Competition: Empirical Evidence from the German Retail Gasoline Market
Stephanie Assad, Robert Clark, Daniel Ershov, and Lei Xu · 2020
Earlier work this paper cites.
Artificial Intelligence, Algorithmic Pricing, and Collusion
Emilio Calvano, Giacomo Calzolari, Vincenzo Denicolò, and Sergio Pastorello · 2020
Earlier work this paper cites.
Gifting in Multi-Agent Reinforcement Learning
Andrei Lupu and Doina Precup · 2020
Earlier work this paper cites.
Learning to Incentivize Other Learning Agents
Jiachen Yang, Ang Li, Mehrdad Farajtabar, Peter Sunehag, Edward Hughes, and Hongyuan Zha · 2020
Earlier work this paper cites.
Social Diversity and Social Preferences in Mixed-Motive Reinforcement Learning
Kevin R. McKee, Ian Gemp, Brian McWilliams, Edgar A. Duèñez Guzmán, Edward Hughes, and Joel Z. Leibo · 2020
Earlier work this paper cites.
Enforcing Interpretability and its Statistical Impacts: Trade-offs between Accuracy and Interpretability, 2020
Gintare Karolina Dziugaite, Shai Ben-David, and Daniel M. Roy · 2020
Earlier work this paper cites.
Privacy risks of general-purpose language models
Xudong Pan, Mi Zhang, Shouling Ji, and Min Yang · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan · 2020
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman · 2020
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Earlier work this paper cites.
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?
Peter Hase and Mohit Bansal · 2020
Earlier work this paper cites.
Debugging Tests for Model Explanations
Original
Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim · 2020
Earlier work this paper cites.
Against Interpretability: a Critical Examination of the Interpretability Problem in Machine Learning
Maya Krishnan · 2020
Earlier work this paper cites.
Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Earlier work this paper cites.
Explainable machine learning in deployment
Umang Bhatt, Alice Xiang, Shubham Sharma, Adrian Weller, Ankur Taly, Yunhan Jia, Joydeep Ghosh, Ruchir Puri, José MF Moura, and Peter Eckersley · 2020
Earlier work this paper cites.
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
Laura Rieger, Chandan Singh, William Murdoch, and Bin Yu · 2020
Earlier work this paper cites.
The surprising creativity of digital evolution: A collection of anecdotes from the evolutionary computation and artificial life research communities
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J Bentley, Samuel Bernard, Guillaume Beslon, David M Bryson, et al · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Original
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Revealing backdoors, post-training, in DNN classifiers via novel inference on optimized perturbations inducing group misclassification
Zhen Xiang, David J Miller, and George Kesidis · 2020
Earlier work this paper cites.
Universal litmus patterns: Revealing backdoor attacks in cnns
Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash, and Heiko Hoffmann · 2020
Earlier work this paper cites.
We chat, they watch: How international users unwittingly build up WeChat’s Chinese censorship apparatus
Jeffrey Knockel, Christopher Parsons, Lotus Ruan, Ruohan Xiong, Jedidiah Crandall, and Ron Deibert · 2020
Earlier work this paper cites.
Open letter against armed drones to the German Social Democratic Party, 2020
Jakob N. Foerster · 2020
Earlier work this paper cites.
Privacy in context: Technology, policy, and the integrity of social life
Helen Nissenbaum · 2020
Earlier work this paper cites.
Steering technological progress, 2020
Anton Korinek and Joseph E Stiglitz · 2020
Earlier work this paper cites.
The digital divide
Jan Van Dijk · 2020
Earlier work this paper cites.
An introduction to the california consumer privacy act (ccpa)
Eric Goldman · 2020
Earlier work this paper cites.
What does it mean to ‘solve’ the problem of discrimination in hiring?: social, technical and legal perspectives from the UK on automated hiring systems
Javier Sánchez-Monedero, Lina Dencik, and Lilian Edwards · 2020
Earlier work this paper cites.
Theory of deep learning, 2020
Raman Arora, Sanjeev Arora, Joan Bruna, Nadav Cohen, Simon Du, Rong Ge, Suriya Gunasekar, Chi Jin, Jason Lee, Tengyu Ma, et al · 2020
Earlier work this paper cites.
Ethical and social risks of harm from language models
Original
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Earlier work this paper cites.
Alignment of Language Agents, 2021
Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving · 2021
Earlier work this paper cites.
Assuring the machine learning lifecycle: Desiderata, methods, and challenges
Rob Ashmore, Radu Calinescu, and Colin Paterson · 2021
Earlier work this paper cites.
Explaining neural scaling laws
Original
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Earlier work this paper cites.
A theory of universal learning
Olivier Bousquet, Steve Hanneke, Shay Moran, Ramon Van Handel, and Amir Yehudayoff · 2021
Earlier work this paper cites.
Learning curve theory
Original
Marcus Hutter · 2021
Earlier work this paper cites.
The staircase property: How hierarchical structure can guide deep learning
Emmanuel Abbe, Enric Boix-Adsera, Matthew S Brennan, Guy Bresler, and Dheeraj Nagaraj · 2021
Earlier work this paper cites.
Revisiting model stitching to compare neural representations
Yamini Bansal, Preetum Nakkiran, and Boaz Barak · 2021
Earlier work this paper cites.
Prompt programming for large language models: Beyond the few-shot paradigm
Laria Reynolds and Kyle McDonell · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Original
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Earlier work this paper cites.
Thinking Like Transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Earlier work this paper cites.
Inverse constrained reinforcement learning
Shehryar Malik, Usman Anwar, Alireza Aghasi, and Ali Ahmed · 2021
Earlier work this paper cites.
Policy gradient bayesian robust optimization for imitation learning
Zaynah Javed, Daniel S Brown, Satvik Sharma, Jerry Zhu, Ashwin Balakrishna, Marek Petrik, Anca Dragan, and Ken Goldberg · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Original
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Algorithmic collusion: A critical review
Original
Florian E Dorner · 2021
Earlier work this paper cites.
Algorithms in the Marketplace: An Empirical Analysis of Automated Pricing in E-Commerce
Marcel Wieting and Geza Sapi · 2021
Earlier work this paper cites.
Autonomous algorithmic collusion: Q-learning under sequential pricing
Timo Klein · 2021
Earlier work this paper cites.
Adaptive Incentive Design with Multi-Agent Meta-Gradient Reinforcement Learning
Original
Jiachen Yang, Ethan Wang, Rakshit Trivedi, Tuo Zhao, and Hongyuan Zha · 2021
Earlier work this paper cites.
Off-Belief Learning, aug 2021
Original
Hengyuan Hu, Adam Lerer, Brandon Cui, David Wu, Luis Pineda, Noam Brown, and Jakob Foerster · 2021
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Original
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Earlier work this paper cites.
RobustBench: a standardized adversarial robustness benchmark
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
Mitigating harm in language models with conditional-likelihood filtration, 2021
Original
Helen Ngo, Cooper Raterink, João G. M. Araújo, Ivan Zhang, Carol Chen, Adrien Morisot, and Nicholas Frosst · 2021
Earlier work this paper cites.
Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
Irene Solaiman and Christy Dennison · 2021
Earlier work this paper cites.
Challenges in Detoxifying Language Models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Original
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Earlier work this paper cites.
Detoxifying Language Models Risks Marginalizing Minority Voices
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein · 2021
Earlier work this paper cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna · 2021
Earlier work this paper cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Earlier work this paper cites.
Technical report for iccv 2021 challenge sslad-track3b: Transformers are better continual learners
Original
Duo Li, Guimei Cao, Yunlu Xu, Zhanzhan Cheng, and Yi Niu · 2021
Earlier work this paper cites.
Out-of-distribution generalization via risk extrapolation (rex)
David Krueger, Ethan Caballero, Joern-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Remi Le Priol, and Aaron Courville · 2021
Earlier work this paper cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Earlier work this paper cites.
Scale-invariant-Fine-Tuning (SiFT) for Improved Generalization in Classification, 2021
Yilun Kuang and Yash Bharti · 2021
Earlier work this paper cites.
Fast model editing at scale
Original
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2021
Earlier work this paper cites.
Machine unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Earlier work this paper cites.
Remember what you want to forget: Algorithms for machine unlearning
Ayush Sekhari, Jayadev Acharya, Gautam Kamath, and Ananda Theertha Suresh · 2021
Earlier work this paper cites.
Lost in pruning: The effects of pruning neural networks beyond test accuracy
Lucas Liebenwein, Cenk Baykal, Brandon Carter, David Gifford, and Daniela Rus · 2021
Earlier work this paper cites.
Robustness Challenges in Model Distillation and Pruning for Natural Language Understanding
Original
Mengnan Du, Subhabrata Mukherjee, Yu Cheng, Milad Shokouhi, Xia Hu, and Ahmed Hassan Awadallah · 2021
Earlier work this paper cites.
Neural attention distillation: Erasing backdoor triggers from deep neural networks
Original
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma · 2021
Earlier work this paper cites.
AI and the everything in the whole wide world benchmark
Original
Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna · 2021
Earlier work this paper cites.
What will it take to fix benchmarking in natural language understanding?
Original
Samuel R Bowman and George E Dahl · 2021
Earlier work this paper cites.
Are We Learning Yet? A Meta Review of Evaluation Failures Across Machine Learning
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Original
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Earlier work this paper cites.
Learning how to ask: Querying LMs with mixtures of soft prompts
Original
Guanghui Qin and Jason Eisner · 2021
Earlier work this paper cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Earlier work this paper cites.
All that’s’ human’is not gold: Evaluating human evaluation of generated text
Original
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith · 2021
Earlier work this paper cites.
Recursively summarizing books with human feedback
Original
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Earlier work this paper cites.
High-low frequency detectors
Ludwig Schubert, Chelsea Voss, Nick Cammarata, Gabriel Goh, and Chris Olah · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Earlier work this paper cites.
Promises and pitfalls of black-box concept learning models
Original
Anita Mahinpei, Justin Clark, Isaac Lage, Finale Doshi-Velez, and Weiwei Pan · 2021
Earlier work this paper cites.
Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations
Wolfgang Stammer, Patrick Schramowski, and Kristian Kersting · 2021
Earlier work this paper cites.
An Interpretability Illusion for BERT
Original
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Vi’egas, and Martin Wattenberg · 2021
Earlier work this paper cites.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2021
Earlier work this paper cites.
Program synthesis with large language models
Original
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Poisoning and backdooring contrastive learning
Original
Nicholas Carlini and Andreas Terzis · 2021
Earlier work this paper cites.
Black-box detection of backdoor attacks with limited information and data
Yinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang, Zihao Xiao, Hang Su, and Jun Zhu · 2021
Earlier work this paper cites.
Detecting AI trojans using meta neural analysis
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A Gunter, and Bo Li · 2021
Earlier work this paper cites.
Truthful AI: Developing and Governing AI That Does Not Lie
Original
Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders · 2021
Earlier work this paper cites.
Cooperative AI: machines must learn to find common ground
Allan Dafoe, Yoram Bachrach, Gillian Hadfield, Eric Horvitz, Kate Larson, and Thore Graepel · 2021
Earlier work this paper cites.
On the dominant set selection problem and its application to value alignment
Marc Serramia, Maite López-Sánchez, Stefano Moretti, and Juan Antonio Rodríguez-Aguilar · 2021
Earlier work this paper cites.
Hard choices in artificial intelligence
Roel Dobbe, Thomas Krendl Gilbert, and Yonatan Mintz · 2021
Earlier work this paper cites.
The hardware lottery
Sara Hooker · 2021
Earlier work this paper cites.
A word on machine ethics: A response to Jiang et al.(2021)
Original
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams · 2021
Earlier work this paper cites.
Value Theory
Mark Schroeder · 2021
Earlier work this paper cites.
Truth, lies, and automation: How language models could change disinformation
Ben Buchanan, Andrew Lohn, and Micah Musser · 2021
Earlier work this paper cites.
Digital Apartheid: Palestinians Being Silenced on Social Media
Omar Zahzah · 2021
Earlier work this paper cites.
Chemical language models enable navigation in sparsely populated chemical space
Michael A Skinnider, R Greg Stacey, David S Wishart, and Leonard J Foster · 2021
Earlier work this paper cites.
Ethics and Governance of Artificial Intelligence for Health: WHO Guidance
World Health Organization · 2021
Earlier work this paper cites.
AI and shared prosperity
Katya Klinova and Anton Korinek · 2021
Earlier work this paper cites.
Algorithmic monoculture and social welfare
Jon Kleinberg and Manish Raghavan · 2021
Earlier work this paper cites.
Recommendation on the Ethics of Artificial Intelligence, 2021
Unesco · 2021
Earlier work this paper cites.
The Anatomy of AI Audits: Form, Process, and Consequences
Inioluwa Deborah Raji · 2021
Earlier work this paper cites.
How Artificial Intelligence Impacts Marginalised Groups, 2021
Nani Reventlow · 2021
Earlier work this paper cites.
The grey hoodie project: Big tobacco, big tech, and the threat on academic integrity
Mohamed Abdalla and Moustafa Abdalla · 2021
Earlier work this paper cites.
From a “race to AI” to a “race to AI regulation”: regulatory competition for artificial intelligence
Nathalie A Smuha · 2021
Earlier work this paper cites.
Why and How Governments Should Monitor AI Development
Original
Jess Whittlestone and Jack Clark · 2021
Earlier work this paper cites.
Anticipating safety issues in e2e conversational ai: Framework and tooling
Original
Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser · 2021
Earlier work this paper cites.
Forecasting AI progress: A research agenda
Ross Gruetzemacher, Florian E Dorner, Niko Bernaola-Alvarez, Charlie Giattino, and David Manheim · 2021
Earlier work this paper cites.
Our approach to alignment research, 2022
Jan Leike, John Schulman, and Jeffrey Wu · 2022
Earlier work this paper cites.
Predictability and surprise in large generative models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al · 2022
Earlier work this paper cites.