Fetching the paper…
Reading the bibliography…
AI alignment aims to make AI systems behave in line with human intentions and values.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020b · 1901
Earlier work this paper cites.
An evaluation of the human-interpretability of explanation
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. 2019 · 1902
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. 2019c · 1906
Earlier work this paper cites.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov. 2019 · 1906
Earlier work this paper cites.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019 · 1907
Earlier work this paper cites.
On the weaknesses of reinforcement learning for neural machine translation
Leshem Choshen, Lior Fox, Zohar Aizenbud, and Omri Abend. 2019 · 1907
Earlier work this paper cites.
Distributionally robust optimization: A review
Hamed Rahimian and Sanjay Mehrotra. 2019 · 1908
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019 · 1908
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
Label ranking by learning pairwise preferences
Eyke Hüllermeier, Johannes Fürnkranz, Weiwei Cheng, and Klaus Brinker. 2008 · 1916
Earlier work this paper cites.
Asimov’s laws
Asimov. 1942 · 1942
Earlier work this paper cites.
Deontic logic
Georg Henrik Von Wright. 1951 · 1951
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Becoming: Basic considerations for a psychology of personality , volume 20
Gordon Willard Allport. 1955 · 1955
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel Bowman. 2020 · 1967
Earlier work this paper cites.
Motivational bases of choice in experimental games
David M Messick and Charles G McClintock. 1968 · 1968
Earlier work this paper cites.
The nature of human values
Milton Rokeach. 1973 · 1973
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett. 1975 · 1975
Earlier work this paper cites.
Social values and rules of fairness: A theoretical perspective
Charles G McClintock and Eddy Van Avermaet. 1982 · 1982
Earlier work this paper cites.
Replicator dynamics
Peter Schuster and Karl Sigmund. 1983 · 1983
Earlier work this paper cites.
Problems of monetary management: the UK experience
Charles AE Goodhart and CAE Goodhart. 1984 · 1984
Earlier work this paper cites.
The effect of social motives, communication and group size on behaviour in an n-person multi-stage mixed-motive game
Wim BG Liebrand. 1984 · 1984
Earlier work this paper cites.
Social choice theory
Amartya Sen. 1986 · 1986
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau. 1991 · 1991
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vladimir Vapnik. 1991 · 1991
Earlier work this paper cites.
Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries
Shalom H Schwartz. 1992 · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan. 1993 · 1993
Earlier work this paper cites.
Are there universal aspects in the structure and contents of human values?
Shalom H Schwartz. 1994 · 1994
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut. 1995 · 1995
Earlier work this paper cites.
Robot see, robot do: An overview of robot imitation
Paul Bakker, Yasuo Kuniyoshi, et al. 1996 · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore. 1996 · 1996
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal. 1996 · 1996
Earlier work this paper cites.
Development of prosocial, individualistic, and competitive orientations: theory and preliminary evidence
Paul AM Van Lange, Ellen De Bruin, Wilma Otten, and Jeffrey A Joireman. 1997 · 1997
Earlier work this paper cites.
Evolutionary game theory
Jörgen W Weibull. 1997 · 1997
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999 · 1999
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal. 1999 · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart Russell, et al. 2000 · 2000
Earlier work this paper cites.
Multiagent planning with factored mdps
Carlos Guestrin, Daphne Koller, and Ronald Parr. 2001 · 2001
Earlier work this paper cites.
A social reinforcement learning agent
Charles Isbell, Christian R Shelton, Michael Kearns, Satinder Singh, and Peter Stone. 2001 · 2001
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Simulation-based optimization of markov reward processes
Peter Marbach and John N Tsitsiklis. 2001 · 2001
Earlier work this paper cites.
Neuroscience, 2nd edition
Dale Purves, George J Augustine, David Fitzpatrick, Lawrence C Katz, Anthony-Samuel LaMantia, James O McNamara, and S Mark. Williams. 2001 · 2001
Earlier work this paper cites.
Evolution of digital organisms at high mutation rates leads to survival of the flattest
Claus O Wilke, Jia Lan Wang, Charles Ofria, Richard E Lenski, and Christoph Adami. 2001 · 2001
Earlier work this paper cites.
Agent-based modeling: Methods and techniques for simulating human systems
Eric Bonabeau. 2002 · 2002
Earlier work this paper cites.
Early stopping-but when?
Lutz Prechelt. 2002 · 2002
Earlier work this paper cites.
Interactive machine learning
Jerry Alan Fails and Dan R Olsen Jr. 2003 · 2003
Earlier work this paper cites.
Pairwise preference learning and ranking
Johannes Fürnkranz and Eyke Hüllermeier. 2003 · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng. 2004 · 2004
Earlier work this paper cites.
Toward trustworthy ai development: mechanisms for supporting verifiable claims
Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, et al. 2020 · 2004
Earlier work this paper cites.
About the role of the environment in multi-agent simulations
Franziska Klügl, Manuel Fehler, and Rainer Herrler. 2005 · 2004
Earlier work this paper cites.
Adversarial training for large neural language models
Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao. 2020 · 2004
Earlier work this paper cites.
" grabcut" interactive foreground extraction using iterated graph cuts
Carsten Rother, Vladimir Kolmogorov, and Andrew Blake. 2004 · 2004
Earlier work this paper cites.
The evolution of cooperation
Joel L Sachs, Ulrich G Mueller, Thomas P Wilcox, and James J Bull. 2004 · 2004
Earlier work this paper cites.
Towards machine ethics: Implementing two action-based ethical theories
Michael Anderson, Susan Anderson, and Chris Armen. 2005 · 2005
Earlier work this paper cites.
Toward ethical robots via mechanized deontic logic
Konstantine Arkoudas, Selmer Bringsjord, and Paul Bello. 2005 · 2005
Earlier work this paper cites.
The logic of scientific discovery
Karl Popper. 2005 · 2005
Earlier work this paper cites.
Invariant visual representation by single neurons in the human brain
R Quian Quiroga, Leila Reddy, Gabriel Kreiman, Christof Koch, and Itzhak Fried. 2005 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Ai research considerations for human existential safety (arches)
Andrew Critch and David Krueger. 2020 · 2006
Earlier work this paper cites.
The status of machine ethics: a report from the aaai symposium
Michael Anderson and Susan Leigh Anderson. 2007 · 2007
Earlier work this paper cites.
Multi-principal assistance games
Arnaud Fickinger, Simon Zhuang, Dylan Hadfield-Menell, and Stuart Russell. 2020 · 2007
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir. 2007 · 2007
Earlier work this paper cites.
Toward harnessing user feedback for machine learning
Simone Stumpf, Vidya Rajaram, Lida Li, Margaret Burnett, Thomas Dietterich, Erin Sullivan, Russell Drummond, and Jonathan Herlocker. 2007 · 2007
Earlier work this paper cites.
Multi-label classification: An overview
Grigorios Tsoumakas and Ioannis Katakis. 2007 · 2007
Earlier work this paper cites.
Adaptive control
Karl Johan Åström and Björn Wittenmark. 2008 · 2008
Earlier work this paper cites.
Tamer: Training an agent manually via evaluative reinforcement
W Bradley Knox and Peter Stone. 2008 · 2008
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty
Joel Lehman, Kenneth O Stanley, et al. 2008 · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
The black swan: the impact of the highly improbable
Nassim Nicholas. 2008 · 2008
Earlier work this paper cites.
The basic ai drives
Stephen M Omohundro. 2008 · 2008
Earlier work this paper cites.
Apprenticeship learning using linear programming
Umar Syed, Michael Bowling, and Robert E Schapire. 2008 · 2008
Earlier work this paper cites.
Teachable robots: Understanding human teaching behavior to build more effective robot learners
Andrea L Thomaz and Cynthia Breazeal. 2008 · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al. 2008 · 2008
Earlier work this paper cites.
Robust optimization , volume 28
Aharon Ben-Tal, Laurent El Ghaoui, and Arkadi Nemirovski. 2009 · 2009
Earlier work this paper cites.
L2 regularization for learning kernels
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh. 2009 · 2009
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2009 · 2009
Earlier work this paper cites.
Overview of supervised learning
Trevor Hastie, Robert Tibshirani, Jerome Friedman, Trevor Hastie, Robert Tibshirani, and Jerome Friedman. 2009 · 2009
Earlier work this paper cites.
Hidden incentives for auto-induced distributional shift
David Krueger, Tegan Maharaj, and Jan Leike. 2020 · 2009
Earlier work this paper cites.
Interacting meaningfully with machine learning systems: Three experiments
Simone Stumpf, Vidya Rajaram, Lida Li, Weng-Keen Wong, Margaret Burnett, Thomas Dietterich, Erin Sullivan, and Jonathan Herlocker. 2009 · 2009
Earlier work this paper cites.
Increasing robotic wheelchair safety with collaborative control: Evidence from secondary task experiments
Tom Carlson and Yiannis Demiris. 2010 · 2010
Earlier work this paper cites.
Predicting partial orders: ranking with abstention
Weiwei Cheng, Michaël Rademaker, Bernard De Baets, and Eyke Hüllermeier. 2010c · 2010
Earlier work this paper cites.
On the consistency of ranking algorithms
John C Duchi, Lester W Mackey, and Michael I Jordan. 2010 · 2010
Earlier work this paper cites.
Preference Learning
Johannes Fürnkranz and Eyke Hüllermeier. 2010 · 2010
Earlier work this paper cites.
Robust solutions to stackelberg games: Addressing bounded rationality and limited observations in human cognition
James Pita, Manish Jain, Milind Tambe, Fernando Ordónez, and Sarit Kraus. 2010 · 2010
Earlier work this paper cites.
Is this paper dangerous? balancing secrecy and openness in counterterrorism
Jacob N Shapiro and David A Siegel. 2010 · 2010
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Peter Stone, Gal Kaminka, Sarit Kraus, and Jeffrey Rosenschein. 2010 · 2010
Earlier work this paper cites.
Reliability and safety engineering , volume 43
Ajit Kumar Verma, Srividya Ajit, Durga Rao Karanki, et al. 2010 · 2010
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2010
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag. 2011 · 2011
Earlier work this paper cites.
Machine ethics
Michael Anderson and Susan Leigh Anderson. 2011 · 2011
Earlier work this paper cites.
Global catastrophic risks
Nick Bostrom and Milan M Cirkovic. 2011 · 2011
Earlier work this paper cites.
Measuring social value orientation
Ryan O Murphy, Kurt A Ackermann, and Michel JJ Handgraaf. 2011 · 2011
Earlier work this paper cites.
Delusion, survival, and intelligent agents
Mark Ring and Laurent Orseau. 2011 · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
A Short Introduction to Preferences: Between AI and Social Choice
Francesca Rossi, Kristen Brent Venable, and Toby Walsh. 2011 · 2011
Earlier work this paper cites.
A survey of crowdsourcing systems
Man-Ching Yuen, Irwin King, and Kwong-Sak Leung. 2011 · 2011
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
Riad Akrour, Marc Schoenauer, and Michèle Sebag. 2012 · 2012
Earlier work this paper cites.
Social choice and individual values , volume 12
Kenneth J Arrow. 2012 · 2012
Earlier work this paper cites.
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Nick Bostrom. 2012 · 2012
Earlier work this paper cites.
Open problems in cooperative ai
Allan Dafoe, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R McKee, Joel Z Leibo, Kate Larson, and Thore Graepel. 2020 · 2012
Earlier work this paper cites.
Aum shinrikyo: insights into how terrorists develop biological and chemical weapons
Richard Danzig. 2012 · 2012
Earlier work this paper cites.
An overview of 11 proposals for building safe advanced ai
Evan Hubinger. 2020 · 2012
Earlier work this paper cites.
Reinforcement learning from simultaneous human and mdp reward
W Bradley Knox and Peter Stone. 2012 · 2012
Earlier work this paper cites.
Learning from human-generated reward
William Bradley Knox. 2012 · 2012
Earlier work this paper cites.
Shadow attacks: automatically evading system-call-behavior based malware detection
Weiqin Ma, Pu Duan, Sanmin Liu, Guofei Gu, and Jyh-Charn Liu. 2012 · 2012
Earlier work this paper cites.
A game-theoretic model and best-response learning method for ad hoc coordination in multiagent systems
Stefano V Albrecht and Subramanian Ramamoorthy. 2013 · 2013
Earlier work this paper cites.
Existential risk prevention as global priority
Nick Bostrom. 2013 · 2013
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, and Andrea L Thomaz. 2013 · 2013
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
Ashesh Jain, Brian Wojcik, Thorsten Joachims, and Ashutosh Saxena. 2013 · 2013
Earlier work this paper cites.
Learning non-myopically from human-generated reward
W Bradley Knox and Peter Stone. 2013 · 2013
Earlier work this paper cites.
Training a robot via human feedback: A case study
W Bradley Knox, Peter Stone, and Cynthia Breazeal. 2013 · 2013
Earlier work this paper cites.
Data-efficient generalization of robot skills with contextual policy search
Andras Kupcsik, Marc Deisenroth, Jan Peters, and Gerhard Neumann. 2013 · 2013
Earlier work this paper cites.
After virtue
Alasdair MacIntyre. 2013 · 2013
Earlier work this paper cites.
Decentralized stochastic control with partial history sharing: A common information approach
Ashutosh Nayyar, Aditya Mahajan, and Demosthenis Teneketzis. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013 · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Preference-based reinforcement learning: A preliminary survey
Christian Wirth and Johannes Fürnkranz. 2013 · 2013
Earlier work this paper cites.
Power to the people: The role of humans in interactive machine learning
Saleema Amershi, Maya Cakmak, William Bradley Knox, and Todd Kulesza. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Agent-based models
Scott De Marchi and Scott E Page. 2014 · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Social value orientation: Theoretical and measurement issues in the study of social preferences
Ryan O Murphy and Kurt A Ackermann. 2014 · 2014
Earlier work this paper cites.
Visualizing mnist: An exploration of dimensionality reduction
Chris Olah. 2014 · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman. 2014 · 2014
Earlier work this paper cites.
Norms as a basis for governing sociotechnical systems
Munindar P Singh. 2014 · 2014
Earlier work this paper cites.
Aligning superintelligence with human interests: A technical research agenda
Nate Soares and Benja Fallenstein. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission
Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad. 2015 · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández. 2015 · 2015
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015 · 2015
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks
Alexander Mordvintsev, Chris Olah, and Mike Tyka. 2015 · 2015
Earlier work this paper cites.
Deep neural networks are easily fooled: High confidence predictions for unrecognizable images
Anh Nguyen, Jason Yosinski, and Jeff Clune. 2015 · 2015
Earlier work this paper cites.
Visualizing representations: Deep learning and human beings
Chris Olah. 2015 · 2015
Earlier work this paper cites.
Causal inference using invariant prediction: identification and confidence intervals. arxiv
J Peters, Peter Buhlmann, and N Meinshausen. 2015 · 2015
Earlier work this paper cites.
Research priorities for robust and beneficial artificial intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark. 2015 · 2015
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky. 2015 · 2015
Earlier work this paper cites.
https://jsteinhardt.wordpress.com/2015/06/24/long-term-and-short-term-challenges-to-ensuring-the-safety-of-ai-systems , title = Long-Term and Short-Term Challenges to Ensuring the Safety of AI Systems
Jacob Steinhardt. 2015 · 2015
Earlier work this paper cites.
Falling rule lists
Fulton Wang and Cynthia Rudin. 2015 · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merriënboer, Armand Joulin, and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
Reinforcement learning as a framework for ethical decision making
David Abel, James MacGlashan, and Michael L Littman. 2016 · 2016
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Earlier work this paper cites.
Racing to the precipice: a model of artificial intelligence development
Stuart Armstrong, Nick Bostrom, and Carl Shulman. 2016 · 2016
Earlier work this paper cites.
Formalizing convergent instrumental goals
Tsvi Benson-Tilsen and Nate Soares. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Handbook of computational social choice
Felix Brandt, Vincent Conitzer, Ulle Endriss, Jérôme Lang, and Ariel D Procaccia. 2016 · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Formal verification of ethical choices in autonomous systems
Louise Dennis, Michael Fisher, Marija Slavkovik, and Matt Webster. 2016 · 2016
Earlier work this paper cites.
Avoiding wireheading with value reinforcement learning
Tom Everitt and Marcus Hutter. 2016 · 2016
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando De Freitas, and Shimon Whiteson. 2016 · 2016
Earlier work this paper cites.
Scott Garrabrant, Tsvi Benson-Tilsen, Andrew Critch, Nate Soares, and Jessica Taylor. 2016 · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon. 2016 · 2016
Earlier work this paper cites.
Interactive machine learning for health informatics: when do we need the human-in-the-loop?
Andreas Holzinger. 2016 · 2016
Earlier work this paper cites.
Modeling human ad hoc coordination
Peter Krafft, Chris Baker, Alex Pentland, and Joshua Tenenbaum. 2016 · 2016
Earlier work this paper cites.
Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning
Robert Loftin, Bei Peng, James MacGlashan, Michael L Littman, Matthew E Taylor, Jeff Huang, and David L Roberts. 2016 · 2016
Earlier work this paper cites.
Formal verication of ethical properties in multiagent systems
Bruno Mermet and Gaële Simon. 2016 · 2016
Earlier work this paper cites.
Anh Nguyen, Jason Yosinski, and Jeff Clune. 2016 · 2016
Earlier work this paper cites.
Towards the science of security and privacy in machine learning
Nicolas Papernot, Patrick McDaniel, Arunesh Sinha, and Michael Wellman. 2016 · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. 2016 · 2016
Earlier work this paper cites.
Deep interactive object selection
Ning Xu, Brian Price, Scott Cohen, Jimei Yang, and Thomas S Huang. 2016 · 2016
Earlier work this paper cites.
Improving the robustness of deep neural networks via stability training
Stephan Zheng, Yang Song, Thomas Leung, and Ian Goodfellow. 2016 · 2016
Earlier work this paper cites.
From reinforcement learning to deep reinforcement learning: An overview
Forest Agostinelli, Guillaume Hocquet, Sameer Singh, and Pierre Baldi. 2018 · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Learning from human preferences
Dario Amodei, Paul Christiano, and Alex Ray. 2017 · 2017
Earlier work this paper cites.
A convex framework for fair regression
Richard Berk, Hoda Heidari, Shahin Jabbari, Matthew Joseph, Michael Kearns, Jamie Morgenstern, Seth Neel, and Aaron Roth. 2017 · 2017
Earlier work this paper cites.
A declarative modular framework for representing and applying ethical principles
Fiona Berreby, Gauvain Bourgne, and Jean-Gabriel Ganascia. 2017 · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Moral decision making frameworks for artificial intelligence
Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg, Yuan Deng, and Max Kramer. 2017 · 2017
Earlier work this paper cites.
Conscientious classification: A data scientist’s guide to discrimination-aware classification
Brian d’Alessandro, Cathy O’Neil, and Tom LaGatta. 2017 · 2017
Earlier work this paper cites.
Steps toward robust artificial intelligence
Thomas G Dietterich. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Feeling the force: Integrating force and pose for fluent discovery through imitation learning to open medicine bottles
Mark Edmonds, Feng Gao, Xu Xie, Hangxin Liu, Siyuan Qi, Yixin Zhu, Brandon Rothrock, and Song-Chun Zhu. 2017 · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
Tom Everitt, Victoria Krakovna, Laurent Orseau, and Shane Legg. 2017 · 2017
Earlier work this paper cites.
Input switched affine networks: An rnn architecture designed for interpretability
Jakob N Foerster, Justin Gilmer, Jascha Sohl-Dickstein, Jan Chorowski, and David Sussillo. 2017 · 2017
Earlier work this paper cites.
What do we need to build explainable ai systems for the medical domain?
Andreas Holzinger, Chris Biemann, Constantinos S Pattichis, and Douglas B Kell. 2017 · 2017
Earlier work this paper cites.
Adversarial attacks on neural network policies
Sandy H. Huang, Nicolas Papernot, Ian J. Goodfellow, Yan Duan, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne. 2017 · 2017
Earlier work this paper cites.
The flash crash: High-frequency trading in an electronic market
Andrei Kirilenko, Albert S Kyle, Mehrdad Samadi, and Tugkan Tuzun. 2017 · 2017
Earlier work this paper cites.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. 2017 · 2017
Earlier work this paper cites.
Interactive visualization and manipulation of attention-based neural machine translation
Jaesong Lee, Joong-Hwi Shin, and Jun-Seok Kim. 2017 · 2017
Earlier work this paper cites.
A review of dynamic stackelberg game models
Tao Li and Suresh P Sethi. 2017 · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman. 2017 · 2017
Earlier work this paper cites.
A future that works: Ai, automation, employment, and productivity
James Manyika, Michael Chui, Mehdi Miremadi, Jacques Bughin, Katy George, Paul Willmott, and Martin Dewhurst. 2017 · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah et al. 2017 · 2017
Earlier work this paper cites.
A multi-agent reinforcement learning model of common-pool resource appropriation
Julien Perolat, Joel Z Leibo, Vinicius Zambaldi, Charles Beattie, Karl Tuyls, and Thore Graepel. 2017 · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. 2017 · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta. 2017 · 2017
Earlier work this paper cites.
The neural lasso: Local linear sparsity for interpretable explanations
Andrew Ross, Isaac Lage, and Finale Doshi-Velez. 2017 · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al. 2017 · 2017
Earlier work this paper cites.
Agent foundations for aligning machine intelligence with human interests: a technical research agenda
Nate Soares and Benya Fallenstein. 2017 · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, Johannes Fürnkranz, et al. 2017 · 2017
Earlier work this paper cites.
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017 · 2017
Earlier work this paper cites.
Artificial intelligence, automation, and work
Daron Acemoglu and Pascual Restrepo. 2018 · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
David Alvarez Melis and Tommi Jaakkola. 2018 · 2018
Earlier work this paper cites.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross. 2018 · 2018
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2018 · 2018
Earlier work this paper cites.
Global overview of imitation learning
Alexandre Attia and Sharone Dayan. 2018 · 2018
Earlier work this paper cites.
The moral machine experiment
Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-François Bonnefon, and Iyad Rahwan. 2018 · 2018
Earlier work this paper cites.
Learning from physical human corrections, one feature at a time
Andrea Bajcsy, Dylan P Losey, Marcia K O’Malley, and Anca D Dragan. 2018 · 2018
Earlier work this paper cites.
Recognition in terra incognita
Sara Beery, Grant Van Horn, and Pietro Perona. 2018 · 2018
Earlier work this paper cites.
Rachel KE Bellamy, Kuntal Dey, Michael Hind, Samuel C Hoffman, Stephanie Houde, Kalapriya Kannan, Pranay Lohia, Jacquelyn Martino, Sameep Mehta, Aleksandra Mojsilovic, et al. 2018 · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning for control: Performance, stability, and deep approximators
Lucian Buşoniu, Tim De Bruin, Domagoj Tolić, Jens Kober, and Ivana Palunko. 2018 · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Paul Christiano, Buck Shlegeris, and Dario Amodei. 2018 · 2018
Earlier work this paper cites.
Iterated distillation and amplification
Ajeya Cotra. 2018 · 2018
Earlier work this paper cites.
Building safe artificial intelligence: specification, robustness, and assurance
DeepMind. 2018 · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht. 2018 · 2018
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Earlier work this paper cites.
Agi safety literature review
Tom Everitt, Gary Lea, and Marcus Hutter. 2018 · 2018
Earlier work this paper cites.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
Ruth Fong and Andrea Vedaldi. 2018 · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson. 2018 · 2018
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich. 2018 · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Drew Arad Hudson and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei. 2018 · 2018
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei. 2018 · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018 · 2018
Earlier work this paper cites.
Learning how to explain neural networks: Patternnet and patternattribution
Pieter-Jan Kindermans, Kristof T Schütt, Maximilian Alber, Klaus-Robert Müller, Dumitru Erhan, Been Kim, and Sven Dähne. 2018 · 2018
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
Svetlana Kiritchenko and Saif M Mohammad. 2018 · 2018
Earlier work this paper cites.
Human-in-the-loop interpretability prior
Isaac Lage, Andrew Ross, Samuel J Gershman, Been Kim, and Finale Doshi-Velez. 2018 · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018 · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton. 2018 · 2018
Earlier work this paper cites.
Visual interrogation of attention-based models for natural language inference and machine comprehension
Shusen Liu, Tao Li, Zhimin Li, Vivek Srikumar, Valerio Pascucci, and Peer-Timo Bremer. 2018 · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018 · 2018
Earlier work this paper cites.
Deep learning: A critical appraisal
Gary Marcus. 2018 · 2018
Earlier work this paper cites.
Transparency by design: Closing the gap between performance and interpretability in visual reasoning
David Mascharka, Philip Tran, Ryan Soklaski, and Arjun Majumdar. 2018 · 2018
Earlier work this paper cites.
A comment on the ida-alphagozero metaphor; capabilities versus alignment
Alex Mennen. 2018 · 2018
Earlier work this paper cites.
On the importance of single directions for generalization
Ari S Morcos, David GT Barrett, Neil C Rabinowitz, and Matthew Botvinick. 2018 · 2018
Earlier work this paper cites.
Algorithms of oppression
Safiya Umoja Noble. 2018 · 2018
Earlier work this paper cites.
A voting-based system for ethical decision making
Ritesh Noothigattu, Snehalkumar Gaikwad, Edmond Awad, Sohan Dsouza, Iyad Rahwan, Pradeep Ravikumar, and Ariel Procaccia. 2018 · 2018
Earlier work this paper cites.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev. 2018 · 2018
Earlier work this paper cites.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J Andrew Bagnell, Pieter Abbeel, Jan Peters, et al. 2018 · 2018
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Aequitas: A bias and fairness audit toolkit
Pedro Saleiro, Benedict Kuester, Loren Hinkson, Jesse London, Abby Stevens, Ari Anisfeld, Kit T Rodolfa, and Rayid Ghani. 2018 · 2018
Earlier work this paper cites.
Learning when to communicate at scale in multiagent cooperative and competitive tasks
Amanpreet Singh, Tushar Jain, and Sainbayar Sukhbaatar. 2018 · 2018
Earlier work this paper cites.
The value learning problem
Nate Soares. 2018 · 2018
Earlier work this paper cites.
Seq2seq-vis: A visual debugging tool for sequence-to-sequence models
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch, Adam Perer, Hanspeter Pfister, and Alexander M Rush. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Earlier work this paper cites.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone. 2018 · 2018
Earlier work this paper cites.
Fairness definitions explained
Sahil Verma and Julia Rubin. 2018 · 2018
Earlier work this paper cites.
Mind the gap: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018 · 2018
Earlier work this paper cites.
The future of work: Robots, AI, and automation
Darrell M West. 2018 · 2018
Earlier work this paper cites.
A low-cost ethics shaping approach for designing reinforcement learning agents
Yueh-Hua Wu and Shou-De Lin. 2018 · 2018
Earlier work this paper cites.
Fairgan: Fairness-aware generative adversarial networks
Depeng Xu, Shuhan Yuan, Lu Zhang, and Xintao Wu. 2018a · 2018
Cited alongside, same era.
Building ethics into artificial intelligence
Han Yu, Zhiqi Shen, Chunyan Miao, Cyril Leung, Victor R Lesser, and Qiang Yang. 2018 · 2018
Cited alongside, same era.
Towards sample efficient reinforcement learning
Yang Yu. 2018 · 2018
Cited alongside, same era.
Challenges to christiano’s capability amplification proposal
E Yudkowsky. 2018 · 2018
Cited alongside, same era.
Mitigating unwanted biases with adversarial learning
Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. 2018a · 2018
Cited alongside, same era.
Visual interpretability for deep learning: a survey
Quan-shi Zhang and Song-Chun Zhu. 2018 · 2018
Cited alongside, same era.
On the sensitivity of reward inference to misspecified human models
Joey Hong, Kush Bhatia, and Anca Dragan. 2022 · 2022
Later among the works it cites.
Building a virtual machine inside chatgpt
Jonas DeGrave. 2022 · 2022
Later among the works it cites.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra. 2022 · 2022
Later among the works it cites.
Threat model literature review
Zachary Kenton, Rohin Shah, David Lindner, Vikrant Varma, Victoria Krakovna, Mary Phuong, Ramana Kumar, and Elliot Catt. 2022 · 2022
Later among the works it cites.
Planning to avoid side effects
Toryn Q Klassen, Sheila A McIlraith, Christian Muise, and Jarvis Xu. 2022 · 2022
Later among the works it cites.
Paradigms of ai alignment: components and enablers
Victoria Krakovna. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
Tianhao Zhang, Zoe McCarthy, Owen Jow, Dennis Lee, Xi Chen, Ken Goldberg, and Pieter Abbeel. 2018d · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
Revisiting the importance of individual units in cnns via ablation
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba. 2018 · 2018
Cited alongside, same era.
Social integration of artificial intelligence: functions, automation allocation logic and human-autonomy trust
Hussein A Abbass. 2019 · 2019
Cited alongside, same era.
problems with ai debate
Stuart Armstrong. 2019 · 2019
Cited alongside, same era.
Ilastik: interactive machine learning for (bio) image analysis
Stuart Berg, Dominik Kutra, Thorben Kroeger, Christoph N Straehle, Bernhard X Kausler, Carsten Haubold, Martin Schiegg, Janez Ales, Thorsten Beier, Markus Rudy, et al. 2019 · 2019
Cited alongside, same era.
The disagreement problem in explainable machine learning: A practitioner’s perspective
Satyapriya Krishna, Tessa Han, Alex Gu, Javin Pombra, Shahin Jabbari, Steven Wu, and Himabindu Lakkaraju. 2022 · 2022
Later among the works it cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Yann LeCun. 2022 · 2022
Later among the works it cites.
Query-efficient and scalable black-box adversarial attacks on discrete sequential data via bayesian optimization
Deokjae Lee, Seungyong Moon, Junhyeok Lee, and Hyun Oh Song. 2022 · 2022
Later among the works it cites.
A proposal for improving societys values
Jan Leike. 2022 · 2022
Later among the works it cites.
Ai and biological weapons
Filippa Lentzos. 2022 · 2022
Later among the works it cites.
Pseudoclick: Interactive Image Segmentation with Click Imitation
Qin Liu, Meng Zheng, Benjamin Planche, Srikrishna Karanam, Terrence Chen, Marc Niethammer, and Ziyan Wu. 2022 · 2022
Later among the works it cites.
Physical interaction as communication: Learning robot objectives online from human corrections
Dylan P Losey, Andrea Bajcsy, Marcia K O’Malley, and Anca D Dragan. 2022 · 2022
Later among the works it cites.
A rigorous study of integrated gradients method and extensions to internal neuron attributions
Daniel D Lundstrom, Tianjian Huang, and Meisam Razaviyayn. 2022 · 2022
Later among the works it cites.
Elign: Expectation alignment as a multi-agent intrinsic reward
Zixian Ma, Rose Wang, Fei-Fei Li, Michael Bernstein, and Ranjay Krishna. 2022 · 2022
Later among the works it cites.
Aligning ai regulation to sociotechnical change
Matthijs M Maas. 2021 · 2022
Later among the works it cites.
Post-hoc interpretability for neural nlp: A survey
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2022 · 2022
Later among the works it cites.
Enhance the visual representation via discrete adversarial training
Xiaofeng Mao, Yuefeng Chen, Ranjie Duan, Yao Zhu, Gege Qi, Xiaodan Li, Rong Zhang, Hui Xue, et al. 2022 · 2022
Later among the works it cites.
On the fragility of learned reward functions
Lev E McKinney, Yawen Duan, David Krueger, and Adam Gleave. 2022 · 2022
Later among the works it cites.
What do nlp researchers believe? results of the nlp community metasurvey
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, et al. 2022 · 2022
Later among the works it cites.
Democratizing ai, stable diffusion & generative models
Emad Mostaque. 2022 · 2022
Later among the works it cites.
Equivariant networks for zero-shot coordination
Darius Muglich, Christian Schroeder de Witt, Elise van der Pol, Shimon Whiteson, and Jakob Foerster. 2022 · 2022
Later among the works it cites.
Iterated distillation-amplification, gato, and proto-agi
Gabriel Mukobi. 2022 · 2022
Later among the works it cites.
Situational awareness: techniques, challenges, and prospects
Arslan Munir, Alexander Aved, and Erik Blasch. 2022 · 2022
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2022 · 2022
Later among the works it cites.
Safe pareto improvements for delegated game playing
Caspar Oesterheld and Vincent Conitzer. 2022 · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
Modeling and mitigating human annotation errors to design efficient stream processing systems with human-in-the-loop machine learning
Rahul Pandey, Hemant Purohit, Carlos Castillo, and Valerie L Shalin. 2022 · 2022
Later among the works it cites.
The unsurprising effectiveness of pre-trained vision models for control
Simone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, and Abhinav Gupta. 2022 · 2022
Later among the works it cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2022 · 2022
Later among the works it cites.
Investigations of performance and bias in human-ai teamwork in hiring
Andi Peng, Besmira Nushi, Emre Kiciman, Kori Inkpen, and Ece Kamar. 2022 · 2022
Later among the works it cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
Deep networks on toroids: removing symmetries reveals the structure of flat regions in the landscape geometry
Fabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer, Carlo Baldassi, and Riccardo Zecchina. 2022 · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. 2022 · 2022
Later among the works it cites.
Linear adversarial concept erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell. 2022 · 2022
Later among the works it cites.
A generalist agent
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexander Novikov, Gabriel Barth-maron, Mai Giménez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. 2022 · 2022
Later among the works it cites.
Gradient hacking
Richard Ngo. 2022 · 2022
Later among the works it cites.
Attention-based interpretability with concept transformers
Mattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind, and Paolo Scotton. 2022 · 2022
Later among the works it cites.
A survey of evaluation metrics used for nlg systems
Ananya B Sai, Akash Kumar Mohankumar, and Mitesh M Khapra. 2022 · 2022
Later among the works it cites.
Self-critiquing models for assisting human evaluators
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022 · 2022
Later among the works it cites.
More examples of gmg
Rohin Shah and Vikrant Varma. 2022 · 2022
Later among the works it cites.
Correcting robot plans with natural language feedback
Pratyusha Sharma, Balakumar Sundaralingam, Valts Blukis, Chris Paxton, Tucker Hermans, Antonio Torralba, Jacob Andreas, and Dieter Fox. 2022 · 2022
Later among the works it cites.
Defining and detecting toxicity on social media: context and knowledge are key
Amit Sheth, Valerie L Shalin, and Ugur Kursuncu. 2022 · 2022
Later among the works it cites.
How to diversify conceptual alignment: the model behind refine
Adam Shimi. 2022 · 2022
Later among the works it cites.
Why so toxic? measuring and triggering toxic behavior in open-domain chatbots
Wai Man Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro, Gianluca Stringhini, Savvas Zannettou, and Yang Zhang. 2022 · 2022
Later among the works it cites.
Defining and characterizing reward gaming
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. 2022 · 2022
Later among the works it cites.
What overarching ethical principle should a superintelligent ai follow?
Atle Ottesen Søvik. 2022 · 2022
Later among the works it cites.
expert survey on progress in ai
Zach Stein-Perlman, Benjamin Weinstein-Raun, and Katja Grace. 2022 · 2022
Later among the works it cites.
Causal confusion and reward misidentification in preference-based reward learning
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca Dragan, and Daniel S Brown. 2022 · 2022
Later among the works it cites.
Inner and outer alignment decompose one hard problem into two extremely hard problems
Alex Turner. 2022 · 2022
Later among the works it cites.
Dual use of artificial-intelligence-powered drug discovery
Fabio Urbina, Filippa Lentzos, Cédric Invernizzi, and Sean Ekins. 2022 · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2022 · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Later among the works it cites.
In situ bidirectional human-robot value alignment
Luyao Yuan, Xiaofeng Gao, Zilong Zheng, Mark Edmonds, Ying Nian Wu, Federico Rossano, Hongjing Lu, Yixin Zhu, and Song-Chun Zhu. 2022 · 2022
Later among the works it cites.
Constructing highly inductive contexts for dialogue safety through controllable reverse generation
Zhexin Zhang, Jiale Cheng, Hao Sun, Jiawen Deng, Fei Mi, Yasheng Wang, Lifeng Shang, and Minlie Huang. 2022 · 2022
Later among the works it cites.
Domain generalization: A survey
Kaiyang Zhou, Ziwei Liu, Yu Qiao, Tao Xiang, and Chen Change Loy. 2022 · 2022
Later among the works it cites.
Adversarial training for high-stakes reliability
Daniel Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Benjamin Weinstein-Raun, Daniel de Haas, et al. 2022 · 2022
Later among the works it cites.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022 · 2022
Later among the works it cites.
Moral foundations of large language models
Marwa Abdulhai, Clément Crepy, Daria Valter, John Canny, and Natasha Jaques. 2022 · 2023
Closest in time.
Ai safety summit 2023: Roundtable chairs’ summaries, 1 november
AI Safety Summit. 2023 · 2023
Closest in time.
Coordinated pausing: An evaluation-based coordination scheme for frontier ai developers
Jide Alaga and Jonas Schuett. 2023 · 2023
Closest in time.
Frontier ai regulation: Managing emerging risks to public safety
Markus Anderljung, Joslyn Barnhart, Jade Leung, Anton Korinek, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, et al. 2023 · 2023
Closest in time.
Circuits updates - july 2023
Anthropic. 2023b · 2023
Closest in time.
Update on ARC’s recent eval efforts
ARC Evals. 2023 · 2023
Closest in time.
Ai4people
Atomium-EISMD. 2023 · 2023
Closest in time.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos. 2023 · 2023
Closest in time.
A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023 · 2023
Closest in time.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023 · 2023
Closest in time.
How rogue ais may arise
Yoshua Bengio. 2023 · 2023
Closest in time.
Pause giant ai experiments: An open letter
Yoshua Bengio, Stuart Russell, Elon Musk, and Future of Life Institute. 2023 · 2023
Closest in time.
Taken out of context: On measuring situational awareness in llms
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans. 2023 · 2023
Closest in time.
Accurate medium-range global weather forecasting with 3d neural networks
Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. 2023 · 2023
Closest in time.
Bipartisan framework for u.s. ai act
Richard Blumenthal and Josh Hawley. 2023 · 2023
Closest in time.
Beyond shared autonomy: Joint perception and action for human-in-the-loop mobile robot navigation systems
Hamed Bozorgi and Trung Dung Ngo. 2023 · 2023
Closest in time.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, et al. 2023 · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023 · 2023
Closest in time.
Deep reinforcement learning from hierarchical weak preference feedback
Alexander Bukharin, Yixiao Li, Pengcheng He, Weizhu Chen, and Tuo Zhao. 2023 · 2023
Closest in time.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al. 2023 · 2023
Closest in time.
Zeno: An interactive framework for behavioral evaluation of machine learning
Ángel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein, Ameet Talwalkar, Jason I Hong, and Adam Perer. 2023 · 2023
Closest in time.
Center for ai safety: Statement on ai risk
CAIS. 2023 · 2023
Closest in time.
’deepfake’ scam in china fans worries over ai-driven fraud
Ella Cao and Eduardo Baptista. 2023 · 2023
Closest in time.
Teaching large language models to zip their lips
Andrew Carr. 2023 · 2023
Closest in time.
Deceptive alignment monitoring
Andres Carranza, Dhruv Pai, Rylan Schaeffer, Arnuv Tandon, and Sanmi Koyejo. 2023 · 2023
Closest in time.
Characterizing manipulation from ai systems
Micah Carroll, Alan Chan, Henry Ashton, and David Krueger. 2023 · 2023
Closest in time.
Moving Forward: 11th post of The Engineer’s Interpretability Sequence
Stephen Casper. 2023 · 2023
Closest in time.
Harms from increasingly agentic algorithmic systems
Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, et al. 2023 · 2023
Closest in time.
An ai challenge: Balancing open and closed systems
Pablo Chavez. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023 · 2023
Closest in time.
Thoughts on the impact of rlhf research
Paul Christiano. 2023 · 2023
Closest in time.
Get it in writing: Formal contracts mitigate social dilemmas in multi-agent rl
Phillip JK Christoffersen, Andreas A Haupt, and Dylan Hadfield-Menell. 2023 · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda. 2023 · 2023
Closest in time.
New lw feature debates
Ruby RobertM GPT-4 Claude. 2023 · 2023
Closest in time.
Introducing the collective intelligence project
Collective Intelligence Project. 2023 · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine N Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023 · 2023
Closest in time.
Tasra: A taxonomy and analysis of societal-scale risks from ai
Andrew Critch and Stuart Russell. 2023 · 2023
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2023 · 2023
Closest in time.
Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. 2023 · 2023
Closest in time.
Analyzing transformers in embedding space
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2023 · 2023
Closest in time.
Learning Dexterous Manipulation from Exemplar Object Trajectories and Pre-Grasps
Sudeep Dasari, Abhinav Gupta, and Vikash Kumar. 2023 · 2023
Closest in time.
RAFT: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, KaShun SHUM, and Tong Zhang. 2023 · 2023
Closest in time.
Cooperative multi-agent learning in a complex world: challenges and solutions
Yali Du. 2023 · 2023
Closest in time.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2023 · 2023
Closest in time.
Who’s harry potter? approximate unlearning in llms
Ronen Eldan and Mark Russinovich. 2023 · 2023
Closest in time.
Eu ai act: first regulation on artificial intelligence
European Parliament. 2023 · 2023
Closest in time.
Bing chat is blatantly, aggressively misaligned
Evan Hubinger. 2023 · 2023
Closest in time.
Google’s ai red team: the ethical hackers making ai safer
Daniel Fabian. 2023 · 2023
Closest in time.
Bridging the gap: A survey on integrating (human) feedback for natural language generation
Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José GC de Souza, Shuyan Zhou, Tongshuang Wu, Graham Neubig, et al. 2023 · 2023
Closest in time.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton. 2023 · 2023
Closest in time.
Josh A Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova. 2023 · 2023
Closest in time.
Frontier ai: capabilities and risks – discussion paper
Government of the United Kingdom. 2023 · 2023
Closest in time.
In ai race, microsoft and google choose speed over caution
Nico Grant and Karen Weise. 2023 · 2023
Closest in time.
Ai control: Improving safety despite intentional subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. 2023 · 2023
Closest in time.
Studying large language model generalization with influence functions
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al. 2023 · 2023
Closest in time.
Ground (less) truth: A causal framework for proxy labels in human-algorithm decision-making
Luke Guerdan, Amanda Coston, Zhiwei Steven Wu, and Kenneth Holstein. 2023 · 2023
Closest in time.
Reinforced self-training (rest) for language modeling
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al. 2023 · 2023
Closest in time.
Language models represent space and time
Wes Gurnee and Max Tegmark. 2023 · 2023
Closest in time.
Secretary-general’s remarks to the security council on artificial intelligence
António Guterres. 2023 · 2023
Closest in time.
A circuit for Python docstrings in a 4-layer attention-only transformer
Stefan Heimersheim and Janiak Jett. 2023 · 2023
Closest in time.
Natural selection favors ais over humans
Dan Hendrycks. 2023 · 2023
Closest in time.
An overview of catastrophic ai risks
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside. 2023 · 2023
Closest in time.
International institutions for advanced ai
Lewis Ho, Joslyn Barnhart, Robert Trager, Yoshua Bengio, Miles Brundage, Allison Carnegie, Rumman Chowdhury, Allan Dafoe, Gillian Hadfield, Margaret Levi, et al. 2023 · 2023
Closest in time.
Ai safety and the age of dislightenment
Jeremy Howard. 2023 · 2023
Closest in time.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al. 2023 · 2023
Closest in time.
Emergent deception and emergent optimization
Jacob Steinhardt. 2023 · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Closest in time.
Automatically auditing large language models via discrete optimization
Erik Jones, Anca D. Dragan, Aditi Raghunathan, and Jacob Steinhardt. 2023 · 2023
Closest in time.
Preference Transformer: Modeling Human Preferences using Transformers for RL
Changyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee. 2023 · 2023
Closest in time.
Evaluating language-model agents on realistic autonomous tasks
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, Elizabeth Barnes, and Paul Christiano. 2023 · 2023
Closest in time.
Reward (mis) design for autonomous driving
W Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone. 2023 · 2023
Closest in time.
Leonie Koessler and Jonas Schuett. 2023 · 2023
Closest in time.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. 2023 · 2023
Closest in time.
Auditing the ai auditors: A framework for evaluating fairness and bias in high stakes ai predictive models
Richard N Landers and Tara S Behrend. 2023 · 2023
Closest in time.
Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research]
Chan Lawrence, Adria Garriga-alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Thomas Nate. 2023 · 2023
Closest in time.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen. 2023b · 2023
Closest in time.
Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik. 2023 · 2023
Closest in time.
LLM-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Closest in time.
Jailbreaking chatgpt via prompt engineering: An empirical study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023 · 2023
Closest in time.
Mechanistic mode connectivity
Ekdeep Singh Lubana, Eric J Bigelow, Robert P Dick, David Krueger, and Hidenori Tanaka. 2023 · 2023
Closest in time.
Interactive language: Talking to robots in real time
Corey Lynch, Ayzaan Wahid, Jonathan Tompson, Tianli Ding, James Betker, Robert Baruch, Travis Armstrong, and Pete Florence. 2023 · 2023
Closest in time.
Introducing democratic fine-tuning
MAI. 2023 · 2023
Closest in time.
Faster sorting algorithms discovered using deep reinforcement learning
Daniel J Mankowitz, Andrea Michi, Anton Zhernov, Marco Gelmi, Marco Selvi, Cosmin Paduraru, Edouard Leurent, Shariq Iqbal, Jean-Baptiste Lespiau, Alex Ahern, et al. 2023 · 2023
Closest in time.
The hydra effect: Emergent self-repair in language model computations
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. 2023 · 2023
Closest in time.
The risks associated with artificial general intelligence: A systematic review
Scott McLean, Gemma JM Read, Jason Thompson, Chris Baber, Neville A Stanton, and Paul M Salmon. 2023 · 2023
Closest in time.
Fairness, accountability, transparency, and ethics (fate) in artificial intelligence (ai), and higher education: A systematic review
Bahar Memarian and Tenzin Doleck. 2023 · 2023
Closest in time.
Meta and microsoft introduce the next generation of llama
Meta. 2023 · 2023
Closest in time.
Recent advances in natural language processing via large pre-trained language models: A survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth. 2023 · 2023
Closest in time.
The threat of offensive ai to organizations
Yisroel Mirsky, Ambra Demontis, Jaidip Kotak, Ram Shankar, Deng Gelei, Liu Yang, Xiangyu Zhang, Maura Pintor, Wenke Lee, Yuval Elovici, et al. 2023 · 2023
Closest in time.
Model-based reinforcement learning: A survey
Thomas M Moerland, Joost Broekens, Aske Plaat, Catholijn M Jonker, et al. 2023 · 2023
Closest in time.
Auditing large language models: a three-layered approach
Jakob Mökander, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi. 2023 · 2023
Closest in time.
Probabilistic Machine Learning: Advanced Topics
Kevin P. Murphy. 2023 · 2023
Closest in time.
Interpretability dreams
Chris Olah. 2023 · 2023
Closest in time.
Introducing superalignment
OpenAI. 2023c · 2023
Closest in time.
Committing to bridging the digital divide in least developed countries
Robert Opp. 2023 · 2023
Closest in time.
A review of cooperative multi-agent deep reinforcement learning
Afshin Oroojlooy and Davood Hajinezhad. 2023 · 2023
Closest in time.
Nvidia ai red team: An introduction
Will Pearce and Joseph Lucas. 2023 · 2023
Closest in time.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. 2023 · 2023
Closest in time.
Investigating emergent goal-like behaviour in large language models using experimental economics
Steve Phelps and Yvan I. Russell. 2023 · 2023
Closest in time.
Question decomposition improves the faithfulness of model-generated reasoning
Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Carson Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, et al. 2023 · 2023
Closest in time.
Microsoft ai red team building future of safer ai
Ram Shankar Siva Kumar. 2023 · 2023
Closest in time.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. 2023 · 2023
Closest in time.
Data distillation: A survey
Noveen Sachdeva and Julian McAuley. 2023 · 2023
Closest in time.
Jonas B Sandbrink. 2023 · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Towards best practices in agi safety and governance: A survey of expert opinion
Jonas Schuett, Noemi Dreksler, Markus Anderljung, David McCaffary, Lennart Heim, Emma Bluemke, and Ben Garfinkel. 2023 · 2023
Closest in time.
Elizabeth Seger, Noemi Dreksler, Richard Moulange, Emily Dardaman, Jonas Schuett, K Wei, Christoph Winter, Mackenzie Arnold, Seán Ó hÉigeartaigh, Anton Korinek, et al. 2023 · 2023
Closest in time.
Against almost every theory of impact of interpretability
Charbel-Raphael Segerie. 2023 · 2023
Closest in time.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Rusheb Shah, Quentin Feuillade Montixi, Soroush Pour, Arush Tagade, and Javier Rando. 2023 · 2023
Closest in time.
Videodex: Learning dexterity from internet videos
Kenneth Shaw, Shikhar Bahl, and Deepak Pathak. 2023 · 2023
Closest in time.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023 · 2023
Closest in time.
Model evaluation for extreme risks
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al. 2023 · 2023
Closest in time.
Benchmarks and algorithms for offline preference-based reward learning
Daniel Shin, Anca D. Dragan, and Daniel S. Brown. 2023 · 2023
Closest in time.
Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy
Buck Shlegeris and Ryan Greenblatt. 2023 · 2023
Closest in time.
Dual rl: Unification and new methods for reinforcement and imitation learning
Harshit Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum. 2023 · 2023
Closest in time.
Invariance in policy optimisation and partial identifiability in reward learning
Joar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave. 2023 · 2023
Closest in time.
Can large language models democratize access to dual-use biotechnology?
Emily H. Soice, Rafael Rocha, Kimberlee Cordova, Michael Specter, and Kevin M. Esvelt. 2023 · 2023
Closest in time.
Ai model gpt-3 (dis)informs us better than humans
Giovanni Spitale, Nikola Biller-Andorno, and Federico Germani. 2023 · 2023
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2023 · 2023
Closest in time.
Assessing the ethical and social concerns of artificial intelligence in neuroinformatics research: An empirical test of the european union assessment list for trustworthy ai (altai)
Bernd Carsten Stahl and Tonii Leach. 2023 · 2023
Closest in time.
The bletchley declaration by countries attending the ai safety summit, 1-2 november 2023
AI Safety Summit. 2023 · 2023
Closest in time.
The global governance of artificial intelligence: Next steps for empirical and normative research
Jonas Tallberg, Eva Erman, Markus Furendal, Johannes Geith, Mark Klamberg, and Magnus Lundgren. 2023 · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Closest in time.
Provably safe systems: the only path to controllable agi
Max Tegmark and Steve Omohundro. 2023 · 2023
Closest in time.
Fact sheet: Biden-harris administration secures voluntary commitments from leading artificial intelligence companies to manage the risks posed by ai
The White House. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
International governance of civilian ai: A jurisdictional certification approach
Robert Trager, Ben Harack, Anka Reuel, Allison Carnegie, Lennart Heim, Lewis Ho, Sarah Kreps, Ranjit Lall, Owen Larter, Seán Ó hÉigeartaigh, et al. 2023 · 2023
Closest in time.
What is ai capability control & why does it matter?
UniteAI. 2023 · 2023
Closest in time.
Population of global offline continues steady decline to 2.6 billion people in 2023
United Nations, ITU. 2023 · 2023
Closest in time.
Using the veil of ignorance to align ai systems with principles of justice
Laura Weidinger, Kevin R McKee, Richard Everett, Saffron Huang, Tina O Zhu, Martin J Chadwick, Christopher Summerfield, and Iason Gabriel. 2023 · 2023
Closest in time.
Adversarial attacks on llms
Lilian Weng. 2023a · 2023
Closest in time.
Llm powered autonomous agents
Lilian Weng. 2023b · 2023
Closest in time.
Ensuring safe, secure, and trustworthy ai
White House. 2023 · 2023
Closest in time.
Emergence of maps in the memories of blind navigation agents
Erik Wijmans, Manolis Savva, Irfan Essa, Stefan Lee, Ari S. Morcos, and Dhruv Batra. 2023 · 2023
Closest in time.
Existential risk from artificial general intelligence
wikipedia. 2023 · 2023
Closest in time.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2023 · 2023
Closest in time.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023 · 2023
Closest in time.
Rt-2: New model translates vision and language into action
Tianhe Yu Yevgen Chebotar. 2023 · 2023
Closest in time.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2023 · 2023
Closest in time.
Language to rewards for robotic skill synthesis
Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montserrat Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, brian ichter, Ted Xiao, Peng Xu, Andy Zeng, Tingnan Zhang, Nicolas Heess, Dorsa Sadigh, Jie Tan, Yuval Tassa, and Fei Xia. 2023 · 2023
Closest in time.
Democratic inputs to ai
Wojciech Zaremb, Arka Dhar, Lama Ahmad, Tyna Eloundou, Shibani Santurkar, Sandhini Agarwal, and Jade Leung. 2023 · 2023
Closest in time.
The wisdom of hindsight makes language models better instruction followers
Tianjun Zhang, Fangchen Liu, Justin Wong, Pieter Abbeel, and Joseph E. Gonzalez. 2023a · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023 · 2023
Closest in time.
Dyval: Graph-informed dynamic evaluation of large language models
Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong, Diyi Yang, and Xing Xie. 2023 · 2023
Closest in time.
Sycophancy to subterfuge: Investigating reward tampering in language models
anthropic. 2024 · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al. 2024 · 2024
Closest in time.
High-dimension human value representation in large language models
Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji, Etsuko Ishii, and Pascale Fung. 2024 · 2024
Closest in time.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei W Koh, Daphne Ippolito, Florian Tramer, and Ludwig Schmidt. 2024 · 2024
Closest in time.
PARL: A unified framework for policy alignment in reinforcement learning
Souradip Chakraborty, Amrit Bedi, Alec Koppel, Huazheng Wang, Dinesh Manocha, Mengdi Wang, and Furong Huang. 2024 · 2024
Closest in time.
Can LLM-generated misinformation be detected?
Canyu Chen and Kai Shu. 2024 · 2024
Closest in time.
Safeguarded ai: Constructing guaranteed safety
David Dalrymple. 2024 · 2024
Closest in time.
Towards guaranteed safe ai: A framework for ensuring robust and reliable ai systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, et al. 2024 · 2024
Closest in time.
Inverse constitutional ai: Compressing preferences into principles
Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier, Samuel Albanie, and Robert Mullins. 2024 · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024 · 2024
Closest in time.
Gemma scope: Helping the safety community shed light on the inner workings of language models
Google DeepMind. 2024 · 2024
Closest in time.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien. 2024 · 2024
Closest in time.
What’s in your" safe" data?: Identifying benign data that breaks safety
Luxi He, Mengzhou Xia, and Peter Henderson. 2024 · 2024
Closest in time.
Contrastive preference learning: Learning from human feedback without reinforcement learning
Joey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn, Scott Niekum, W. Bradley Knox, and Dorsa Sadigh. 2024 · 2024
Closest in time.
Do large language models know about facts?
Xuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo, Lijie Wen, Philip S. Yu, and Zhijiang Guo. 2024 · 2024
Closest in time.
Catastrophic jailbreak of open-source LLMs via exploiting generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. 2024 · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al. 2024 · 2024
Closest in time.
On scalable oversight with weak llms judging strong llms
Zachary Kenton, Noah Y Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D Goodman, et al. 2024 · 2024
Closest in time.
Debating with more persuasive llms leads to more truthful answers
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R Bowman, Tim Rocktäschel, and Ethan Perez. 2024 · 2024
Closest in time.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2024 · 2024
Closest in time.
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, et al. 2024 · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al. 2024 · 2024
Closest in time.
Leon Lang, Davis Foote, Stuart Russell, Anca Dragan, Erik Jenner, and Scott Emmons. 2024 · 2024
Closest in time.
Can LLMs keep a secret? testing privacy implications of language models via contextual integrity theory
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2024 · 2024
Closest in time.
The alignment problem from a deep learning perspective: A position paper
Richard Ngo, Lawrence Chan, and Sören Mindermann. 2024 · 2024
Closest in time.
How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions
Lorenzo Pacchiardi, Alex James Chan, Sören Mindermann, Ilan Moscovitz, Alexa Yue Pan, Yarin Gal, Owain Evans, and Jan M. Brauner. 2024 · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2024 · 2024
Closest in time.
Rethinking information structures in rlhf: Reward generalization from a graph theory perspective
Tianyi Qiu, Fanzhi Zeng, Jiaming Ji, Dong Yan, Kaile Wang, Jiayi Zhou, Han Yang, Josef Dai, Xuehai Pan, and Yaodong Yang. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
Safetywashing: Do ai safety benchmarks actually measure safety progress?
Richard Ren, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan H Kim, et al. 2024 · 2024
Closest in time.
Evaluating the moral beliefs encoded in llms
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. 2024 · 2024
Closest in time.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2024 · 2024
Closest in time.
Beyond memorization: Violating privacy via inference with large language models
Robin Staab, Mark Vero, Mislav Balunovic, and Martin Vechev. 2024 · 2024
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2024 · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. 2024 · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Closest in time.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024 · 2024
Closest in time.
Beyond reverse KL: Generalizing direct preference optimization with diverse divergence constraints
Chaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu, and Yuxin Chen. 2024 · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2024 · 2024
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2024 · 2024
Closest in time.
Improving out-of-domain generalization with domain relations
Huaxiu Yao, Xinyu Yang, Xinyi Pan, Shengchao Liu, Pang Wei Koh, and Chelsea Finn. 2024 · 2024
Closest in time.
Rrhf: Rank responses to align language models with human feedback
Hongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2024 · 2024
Closest in time.
On evaluating adversarial robustness of large vision-language models
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Man Cheung, and Min Lin. 2024 · 2024
Closest in time.
Improving generalization of alignment with human preferences through group invariant learning
Rui Zheng, Wei Shen, Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Haoran Huang, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024 · 2024
Closest in time.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2024 · 2024
Closest in time.
Few-Shot Preference Learning for Human-in-the-Loop RL
Donald Joseph Hejna III and Dorsa Sadigh. 2022 · 2025
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.
Bbq: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al. 2018 · 2087
Closest in time.