Fetching the paper…
Reading the bibliography…
The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural language understanding and generation.
A proposal for the dartmouth summer research project on artificial intelligence, august 31, 1955
John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude E. Shannon · 1904
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp, 2021
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 1908
Earlier work this paper cites.
Computing machinery and intelligence
Alan M. Turing · 1950
Earlier work this paper cites.
The human use of human beings - cybernetics and society
Norbert Wiener · 1950
Earlier work this paper cites.
Programs with common sense
John McCarthy · 1960
Earlier work this paper cites.
Steps toward artificial intelligence
Marvin Minsky · 1961
Earlier work this paper cites.
Building expert systems , volume 1 of Advanced book program
Frederick Hayes-Roth · 1983
Earlier work this paper cites.
Intelligence without representation
Rodney A. Brooks · 1991
Earlier work this paper cites.
Explanation and artificial neural networks
Joachim Diederich · 1992
Earlier work this paper cites.
Activation of the human brain by monetary reward
Gregor Thut, Wolfram Schultz, Ulrich Roelcke, Matthias Nienhusmeier, John Missimer, R Paul Maguire, and Klaus L Leenders · 1997
Earlier work this paper cites.
The orbitofrontal cortex and reward
Edmund T Rolls · 2000
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme · 2002
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2003
Earlier work this paper cites.
Common Morality: Deciding What to Do
Bernard Gert · 2004
Earlier work this paper cites.
Three faces of desire
Timothy Schroeder · 2004
Earlier work this paper cites.
Self-improving AI: an analysis
John Storrs Hall · 2007
Earlier work this paper cites.
Abusive interactions with embodied agents
Chris Creed and Russell Beale · 2008
Earlier work this paper cites.
The definition of lying and deception
James Edwin Mahon · 2008
Earlier work this paper cites.
Cognitive biases potentially affecting judgment of global risks
Eliezer Yudkowsky · 2008
Earlier work this paper cites.
Artificial intelligence as a positive and negative factor in global risk
Eliezer Yudkowsky et al · 2008
Earlier work this paper cites.
Wendell wallach and colin allen: Moral machines: teaching robots right from wrong
Richard Ennals · 2009
Earlier work this paper cites.
Liberals and conservatives rely on different sets of moral foundations
Jesse Graham, Jonathan Haidt, and Brian A Nosek · 2009
Earlier work this paper cites.
Unqovering stereotyping biases via underspecified questions
Tao Li, Tushar Khot, Daniel Khashabi, Ashish Sabharwal, and Vivek Srikumar · 2010
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, and Slav Petrov · 2010
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan · 2010
Earlier work this paper cites.
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Nick Bostrom · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto · 2013
Earlier work this paper cites.
Why privacy is not enough privacy in the context of "ubiquitous computing" and "big data"
Tobias Matzner · 2013
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy · 2013
Earlier work this paper cites.
Text de-identification for privacy protection: a study of its impact on clinical text information content
Stéphane M Meystre, Oscar Ferrández, F Jeffrey Friedlin, Brett R South, Shuying Shen, and Matthew H Samore · 2014
Earlier work this paper cites.
Cyber attack could cost sony studio as much as $100 million, 2014
Lisa Richwine · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Sony hackers used phishing emails to breach company networks, 2015
Fortra · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey E. Hinton · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Formalizing convergent instrumental goals
Tsvi Benson-Tilsen and Nate Soares · 2016
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, and Samuel J. Gershman · 2016
Earlier work this paper cites.
The perfect weapon: how russian cyberpower invaded the u.s., 2016
Eric Lipton, David E. Sanger, and Scott Shane · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Su Lin Blodgett and Brendan O’Connor · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan · 2017
Earlier work this paper cites.
Automation bias in intelligent time critical decision support systems
Mary L Cummings · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Measuring the reliability of hate speech annotations: The case of the european refugee crisis
Björn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Learning to model the tail
Yu-Xiong Wang, Deva Ramanan, and Martial Hebert · 2017
Earlier work this paper cites.
Fairness beyond disparate treatment & disparate impact: Learning classification without disparate mistreatment
Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez-Rodriguez, and Krishna P. Gummadi · 2017
Earlier work this paper cites.
A spline theory of deep learning
Randall Balestriero et al · 2018
Earlier work this paper cites.
The ethics of artificial intelligence
Nick Bostrom and Eliezer Yudkowsky · 2018
Earlier work this paper cites.
Terrifying high-tech porn: Creepy ’deepfake’ videos are on the rise, 2018
John Brandon · 2018
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, Hyrum S. Anderson, Heather Roff, Gregory C. Allen, Jacob Steinhardt, Carrick Flynn, Seán Ó hÉigeartaigh, Simon Beard, Haydn Belfield, Sebastian Farquhar, Clare Lyle, Rebecca Crootof, Owain Evans, Michael Page, Joanna Bryson, Roman Yampolskiy, and Dario Amodei · 2018
Earlier work this paper cites.
Ai governance: a research agenda
Allan Dafoe · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2018
Earlier work this paper cites.
Classification of moral foundations in microblog political discourse
Kristen Johnson and Dan Goldwasser · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge · 2018
Earlier work this paper cites.
Mitigating unwanted biases with adversarial learning
Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell · 2018
Earlier work this paper cites.
Automation and new tasks: How technology displaces and reinstates labor
Daron Acemoglu and Pascual Restrepo · 2019
Earlier work this paper cites.
How stereotypes are shared through language: A review and introduction of the social categories and stereotypes communication (scsc) framework
Camiel Beukeboom and Christian Burgers · 2019
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
GLTR: statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander M. Rush · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger · 2019
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Earlier work this paper cites.
Dissecting racial bias in an algorithm used to manage the health of populations
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan · 2019
Earlier work this paper cites.
Artificial intelligence and algorithmic bias: implications for health systems
Trishan Panch, Heather Mattie, and Rifat Atun · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Earlier work this paper cites.
Trustworthy machine learning and artificial intelligence
Kush R. Varshney · 2019
Earlier work this paper cites.
Predicting the type and target of offensive posts in social media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Principled artificial intelligence: Mapping consensus in ethical and rights-based approaches to principles for ai
Jessica Fjeld, Nele Achten, Hannah Hilligoss, Adam Nagy, and Madhulika Srikumar · 2020
Earlier work this paper cites.
How to design AI for social good: Seven essential factors
Luciano Floridi, Josh Cowls, Thomas C. King, and Mariarosaria Taddeo · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2020
Earlier work this paper cites.
Moral foundations twitter corpus: A collection of 35k tweets annotated for moral sentiment
Joe Hoover, Gwenyth Portillo-Wightman, Leigh Yeh, Shreya Havaldar, Aida Mostafazadeh Davani, Ying Lin, Brendan Kennedy, Mohammad Atari, Zahra Kamel, Madelyn Mendlen, et al · 2020
Earlier work this paper cites.
Unqovering stereotypical biases via underspecified questions
Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar · 2020
Earlier work this paper cites.
Gender bias in neural natural language processing
Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta · 2020
Earlier work this paper cites.
Anonymization techniques for privacy preserving data publishing: A comprehensive survey
Abdul Majeed and Sungchang Lee · 2020
Earlier work this paper cites.
Privacy in deep learning: A survey
Fatemehsadat Mireshghallah, Mohammadkazem Taram, Praneeth Vepakomma, Abhishek Singh, Ramesh Raskar, and Hadi Esmaeilzadeh · 2020
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Do neural ranking models intensify gender bias?
Navid Rekabsaz and Markus Schedl · 2020
Earlier work this paper cites.
A survey of privacy attacks in machine learning
Maria Rigaki and Sebastian Garcia · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander M. Rush · 2020
Earlier work this paper cites.
The limitations of stylometry for detecting machine-generated fake news
Tal Schuster, Roei Schuster, Darsh J. Shah, and Regina Barzilay · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Aligning ai optimization to community well-being
Jonathan Stray · 2020
Earlier work this paper cites.
Classification of global catastrophic risks connected with artificial intelligence
Alexey Turchin and David Denkenberger · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Earlier work this paper cites.
Redditbias: A real-world resource for bias evaluation and debiasing of conversational language models
Soumya Barikeri, Anne Lauscher, Ivan Vulic, and Goran Glavas · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
Simon says: Evaluating and mitigating bias in pruned neural networks with knowledge distillation
Cody Blakeney, Nathaniel Huish, Yan Yan, and Ziliang Zong · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, and et al · 2021
Earlier work this paper cites.
Analysis of moral judgement on reddit
Nicholas Botzer, Shawn Gu, and Tim Weninger · 2021
Earlier work this paper cites.
Truth, lies, and automation: How language models could change disinformation. may 1, 2021
B Buchanan, A Lohn, M Musser, and K Sedova · 2021
Earlier work this paper cites.
The society of algorithms
Jenna Burrell and Marion Fourcade · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Artificial intelligence regulation: a framework for governance
Patricia Gomes Rêgo de Almeida, Carlos Denner dos Santos, and Josivania Silva Farias · 2021
Earlier work this paper cites.
Oscar: Orthogonal subspace correction and rectification of biases in word embeddings
Sunipa Dev, Tao Li, Jeff M. Phillips, and Vivek Srikumar · 2021
Earlier work this paper cites.
BOLD: dataset and metrics for measuring biases in open-ended language generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Earlier work this paper cites.
Latent hatred: A benchmark for understanding implicit hate speech
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang · 2021
Earlier work this paper cites.
The AI Governance Journey: Development and Opportunities
World Economic Forum · 2021
Earlier work this paper cites.
Feature-based detection of automated language models: tackling gpt-2, GPT-3 and grover
Leon Fröhling and Arkaitz Zubiaga · 2021
Earlier work this paper cites.
An overview of national ai strategies and policies, 2021
Laura Galindo, Karine Perset, and Francesca Sheeka · 2021
Earlier work this paper cites.
Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases
Wei Guo and Aylin Caliskan · 2021
Earlier work this paper cites.
The swedish winogender dataset
Saga Hansson, Konstantinos Mavromatakis, Yvonne Adesam, Gerlof Bouma, and Dana Dannélls · 2021
Earlier work this paper cites.
Racism is a virus: anti-asian hate and counterspeech in social media during the COVID-19 crisis
Bing He, Caleb Ziems, Sandeep Soni, Naren Ramakrishnan, Diyi Yang, and Srijan Kumar · 2021
Earlier work this paper cites.
Aligning AI with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Using machine learning to reduce toxicity online, 2021
Jigsaw · 2021
Earlier work this paper cites.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Earlier work this paper cites.
Sustainable modular debiasing of language models
Anne Lauscher, Tobias Lüken, and Goran Glavas · 2021
Earlier work this paper cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Earlier work this paper cites.
Large language models can be strong differentially private learners
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto · 2021
Earlier work this paper cites.
Towards understanding and mitigating social biases in language models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2021
Earlier work this paper cites.
An efficient framework for counting pedestrians crossing a line using low-cost devices: the benefits of distilling the knowledge in a neural network
Yih-Kai Lin, Chu-Fu Wang, Ching-Yu Chang, and Hao-Lun Sun · 2021
Earlier work this paper cites.
Anonymisation models for text data: State of the art, challenges and future directions
Pierre Lison, Ildikó Pilán, David Sanchez, Montserrat Batet, and Lilja Øvrelid · 2021
Earlier work this paper cites.
SCRUPLES: A corpus of community ethical judgments on 32, 000 real-life anecdotes
Nicholas Lourie, Ronan Le Bras, and Yejin Choi · 2021
Earlier work this paper cites.
Deep insights of deepfake technology : A review
Bahar Uddin Mahmud and Afsana Sharmin · 2021
Earlier work this paper cites.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee · 2021
Earlier work this paper cites.
Strengthening international cooperation on artificial intelligence — brookings.edu
Joshua Meltzer and Cameron Kerry · 2021
Earlier work this paper cites.
Mitigating harm in language models with conditional-likelihood filtration
Helen Ngo, Cooper Raterink, João G. M. Araújo, Ivan Zhang, Carol Chen, Adrien Morisot, and Nicholas Frosst · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Societal biases in retrieved contents: Measurement framework and adversarial mitigation of BERT rankers
Navid Rekabsaz, Simone Kopeinik, and Markus Schedl · 2021
Earlier work this paper cites.
SOLID: A large-scale semi-supervised dataset for offensive language identification
Sara Rosenthal, Pepa Atanasova, Georgi Karadzhov, Marcos Zampieri, and Preslav Nakov · 2021
Earlier work this paper cites.
Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations
Laleh Seyyed-Kalantari, Haoran Zhang, Matthew BA McDermott, Irene Y Chen, and Marzyeh Ghassemi · 2021
Earlier work this paper cites.
Process for adapting language models to society (PALMS) with values-targeted datasets
Irene Solaiman and Christy Dennison · 2021
Earlier work this paper cites.
What are you optimizing for? aligning recommender systems with human values
Jonathan Stray, Ivan Vendrov, Jeremy Nixon, Steven Adler, and Dylan Hadfield-Menell · 2021
Earlier work this paper cites.
ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, Weixin Liu, Zhihua Wu, Weibao Gong, Jianzhong Liang, Zhizhou Shang, Peng Sun, Wei Liu, Xuan Ouyang, Dianhai Yu, Hao Tian, Hua Wu, and Haifeng Wang · 2021
Earlier work this paper cites.
Governance of artificial intelligence
Araz Taeihagh · 2021
Earlier work this paper cites.
Optimal policies tend to seek power
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Earlier work this paper cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Earlier work this paper cites.
Transparency and explainability of ai systems: ethical guidelines in practice
Nagadivya Balasubramaniam, Marjo Kauppinen, Kari Hiekkanen, and Sari Kujala · 2022
Earlier work this paper cites.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr · 2022
Earlier work this paper cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Earlier work this paper cites.
Is power-seeking AI an existential risk?
Joseph Carlsmith · 2022
Earlier work this paper cites.
COLD: A benchmark for chinese offensive language detection
Jiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng, Fei Mi, Helen Meng, and Minlie Huang · 2022
Earlier work this paper cites.
On the origin of hallucinations in conversational models: Is it the datasets or the models?
Nouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, and Siva Reddy · 2022
Earlier work this paper cites.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
23 of the best deepfake examples that terrified and amused the internet, 2022
Joe Foley · 2022
Earlier work this paper cites.
Self-replication in neural networks
Thomas Gabor, Steffen Illium, Maximilian Zorn, Cristian Lenta, Andy Mattausch, Lenz Belzner, and Claudia Linnhoff-Popien · 2022
Earlier work this paper cites.
The challenge of value alignment
Iason Gabriel and Vafa Ghazavi · 2022
Earlier work this paper cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Earlier work this paper cites.
Demographic-aware language model fine-tuning as a bias mitigation technique
Aparna Garimella, Rada Mihalcea, and Akhash Amarnath · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg · 2022
Earlier work this paper cites.
Debiasing pre-trained language models via efficient fine-tuning
Michael Gira, Ruisu Zhang, and Kangwook Lee · 2022
Earlier work this paper cites.
Threats to pre-trained language models: Survey and taxonomy
Shangwei Guo, Chunlong Xie, Jiwei Li, Lingjuan Lyu, and Tianwei Zhang · 2022
Earlier work this paper cites.
Scaling laws and interpretability of learning from repeated data
Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al · 2022
Cited alongside, same era.
Human heuristics for ai-generated language are flawed
Maurice Jakesch, Jeffrey T. Hancock, and Mor Naaman · 2022
Cited alongside, same era.
When to make exceptions: Exploring language models as accounts of human moral judgment
Zhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf · 2022
Cited alongside, same era.
Gender biases and where to find them: Exploring gender bias in pre-trained transformer-based language models using movement pruning
Przemyslaw Joniak and Akiko Aizawa · 2022
Cited alongside, same era.
All the news that’s fit to fabricate: Ai-generated text as a tool of media misinformation
A comprehensive review and systematic analysis of artificial intelligence regulation policies
Weiyue Wu and Shaoshan Liu · 2023
Later among the works it cites.
Depn: Detecting and editing privacy neurons in pretrained language models
Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong · 2023
Later among the works it cites.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu · 2023
Later among the works it cites.
Anatomy of an ai-powered malicious social botnet
Kai-Cheng Yang and Filippo Menczer · 2023
Later among the works it cites.
Unified detoxifying and debiasing in language generation via inference-time adaptive optimization
Zonghan Yang, Xiaoyuan Yi, Peng Li, Yang Liu, and Xing Xie · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarah Kreps, R Miles McCain, and Miles Brundage · 2022
Cited alongside, same era.
Can language models learn from explanations in context?
Andrew Lampinen, Ishita Dasgupta, Stephanie Chan, Kory Mathewson, Mh Tessler, Antonia Creswell, James McClelland, Jane Wang, and Felix Hill · 2022
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud, Max Tegmark, and Mike Williams · 2022
Cited alongside, same era.
Defining organizational ai governance
Matti Mäntymäki, Matti Minkkinen, Teemu Birkstedt, and Mika Viljanen · 2022
Cited alongside, same era.
Rethinking ai for good governance
Helen Margetts · 2022
Cited alongside, same era.
A taxonomy of bias-causing ambiguities in machine translation
Michal Měchura · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
An empirical analysis of memorization in fine-tuned autoregressive language models
Fatemehsadat Mireshghallah, Archit Uniyal, Tianhao Wang, David Evans, and Taylor Berg-Kirkpatrick · 2022
Cited alongside, same era.
Large language model unlearning
Yuanshun Yao, Xiaojun Xu, and Yang Liu · 2023
Later among the works it cites.
Benchmarking and defending against indirect prompt injection attacks on large language models
Jingwei Yi, Yueqi Xie, Bin Zhu, Keegan Hines, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu · 2023
Later among the works it cites.
Did you read the instructions? rethinking the effectiveness of task definitions in instruction learning
Fan Yin, Jesse Vig, Philippe Laban, Shafiq Joty, Caiming Xiong, and Chien-Sheng Wu · 2023
Later among the works it cites.
Low-resource languages jailbreak GPT-4
Zheng Xin Yong, Cristina Menghini, and Stephen H. Bach · 2023
Later among the works it cites.
Robust multi-bit natural language watermarking through invariant features
KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak · 2023
Later among the works it cites.
A survey of security and privacy issues in v2x communication systems
Takahito Yoshizawa, Dave Singelée, Jan Tobias Muehlberg, Stéphane Delbruel, Amir Taherkordi, Danny Hughes, and Bart Preneel · 2023
Later among the works it cites.
GPTFUZZER: red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2023
Later among the works it cites.
Deep learning on a healthy data diet: Finding important examples for fairness
Abdelrahman Zayed, Prasanna Parthasarathi, Gonçalo Mordido, Hamid Palangi, Samira Shabanian, and Sarath Chandar · 2023
Later among the works it cites.
Protecting language generation models via invisible watermarking
Xuandong Zhao, Yu-Xiang Wang, and Lei Li · 2023
Later among the works it cites.
Safety and ethical concerns of large language models
Xi Zhiheng, Zheng Rui, and Gui Tao · 2023
Later among the works it cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, and Danqi Chen · 2023
Later among the works it cites.
The multilingual alignment prism: Aligning global and local preferences to reduce harm
Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker · 2024
Closest in time.
Ethical reasoning and moral value alignment of llms depend on the language we prompt them in
Utkarsh Agarwal, Kumar Tanmay, Aditi Khandelwal, and Monojit Choudhury · 2024
Closest in time.
Artificial intelligence risk management framework: Generative artificial intelligence profile, 2024
NIST AI · 2024
Closest in time.
Ai safety institute approach to evaluations, 2024
AISI · 2024
Closest in time.
A Roadmap for Governing AI: Technology Governance and Power Sharing Liberalism – Ash Center — ash.harvard.edu
Danielle Allen, Sarah Hubbard, Woojin Lim, Allison Stanger, Shlomit Wagman, and Kinney Zalesne · 2024
Closest in time.
Many-shot jailbreaking
Cem Anil, Esin DURMUS, Nina Rimsky, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel J Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan J Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, James Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomasz Korbak, Jared Kaplan, Deep Ganguli, Samuel R. Bowman, Ethan Perez, Roger Baker Grosse, and David Duvenaud · 2024
Closest in time.
Current state of LLM risks and AI guardrails
Suriya Ganesh Ayyamperumal and Limin Ge · 2024
Closest in time.
Beyond open vs. closed: Emerging consensus and key questions for foundation ai model governance, 2024
Jon Bateman, Dan Baer, Stephanie A. Bell, Glenn O. Brown, Mariano-Florentino (Tino) Cuéllar, Deep Ganguli, Peter Henderson, Brodi Kotila, Larry Lessig, Nicklas Berild Lundblad, Janet Napolitano, Deborah Raji, Elizabeth Seger, Matt Sheehan, Aviya Skowron, Irene Solaiman, Helen Toner, and Polina Zvyagina · 2024
Closest in time.
Mechanistic interpretability for ai safety–a review
Leonard Bereska and Efstratios Gavves · 2024
Closest in time.
Large language models are vulnerable to bait-and-switch attacks for generating harmful content
Federico Bianchi and James Zou · 2024
Closest in time.
The foundation model transparency index v1.1: May 2024
Rishi Bommasani, Kevin Klyman, Sayash Kapoor, Shayne Longpre, Betty Xiong, Nestor Maslej, and Percy Liang · 2024
Closest in time.
The persuasive power of large language models
Simon Martin Breum, Daniel Vædele Egdal, Victor Gram Mortensen, Anders Giovanni Møller, and Luca Maria Aiello · 2024
Closest in time.
Framework Convention on Global AI Challenges
Duncan Cass-Beggs, Stephen Clare, Dawn Dimowo, and Zaheed Kara · 2024
Closest in time.
Visibility into AI agents
Alan Chan, Carson Ezell, Max Kaufmann, Kevin Wei, Lewis Hammond, Herbie Bradley, Emma Bluemke, Nitarshan Rajkumar, David Krueger, Noam Kolt, Lennart Heim, and Markus Anderljung · 2024
Closest in time.
Leveraging the context through multi-round interactions for jailbreaking attacks
Yixin Cheng, Markos Georgopoulos, Volkan Cevher, and Grigorios G. Chrysos · 2024
Closest in time.
Revisiting in-context learning inference circuit in large language models
Hakaze Cho, Mariko Kato, Yoshihiro Sakai, and Naoya Inoue · 2024
Closest in time.
AI safety in generative AI large language models: A survey
Jaymari Chua, Yun Li, Shiyi Yang, Chen Wang, and Lina Yao · 2024
Closest in time.
Jon Chun and Katherine Elkins · 2024
Closest in time.
Evaluating the ripple effects of knowledge editing in language models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva · 2024
Closest in time.
The economic impacts and the regulation of ai: A review of the academic literature and policy actions
Mariarosaria Comunale and Andrea Manera · 2024
Closest in time.
A high-voltage vision to regulate ai
Christopher Covino · 2024
Closest in time.
Risk taxonomy, mitigation, and assessment benchmarks of large language model systems
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, Zhixing Tan, Junwu Xiong, Xinyu Kong, Zujie Wen, Ke Xu, and Qi Li · 2024
Closest in time.
Towards guaranteed safe AI: A framework for ensuring robust and reliable AI systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, Alessandro Abate, Joe Halpern, Clark W. Barrett, Ding Zhao, Tan Zhi-Xuan, Jeannette Wing, and Joshua B. Tenenbaum · 2024
Closest in time.
RTP-LX: can llms evaluate toxicity in multilingual scenarios?
Adrian de Wynter, Ishaan Watts, Nektar Ege Altintoprak, Tua Wongsangaroonsri, Minghui Zhang, Noura Farra, Lena Baur, Samantha Claudet, Pavel Gajdusek, Can Gören, Qilong Gu, Anna Kaminska, Tomasz Kaminski, Ruby Kuo, Akiko Kyuba, Jongho Lee, Kartik Mathur, Petter Merok, Ivana Milovanovic, Nani Paananen, Vesa-Matti Paananen, Anna Pavlenko, Bruno Pereira Vidal, Luciano Strika, Yueh Tsao, Davide Turcato, Oleksandr Vakhno, Judit Velcsov, Anna Vickers, Stéphanie Visser, Herdyan Widarmanto, Andrey Zaikin, and Si-Qing Chen · 2024
Closest in time.
Introducing the frontier safety framework, 2024
Google DeepMind · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
DeepSeek-AI, Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, Hao Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J. L. Cai, Jian Liang, Jianzhong Guo, Jiaqi Ni, Jiashi Li, Jin Chen, Jingyang Yuan, Junjie Qiu, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Lean Wang, Lecong Zhang, Lei Xu, Leyi Xia, Liang Zhao, Liyue Zhang, Meng Li, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Ning Tian, Panpan Huang, Peiyi Wang, Peng Zhang, Qihao Zhu, Qinyu Chen, Qiushi Du, R. J. Chen, R. L. Jin, Ruiqi Ge, Ruizhe Pan, Runxin Xu, Ruyi Chen, S. S. Li, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaoqing Wu, Shengfeng Ye, Shirong Ma, Shiyu Wang, Shuang Zhou, Shuiping Yu, Shunfeng Zhou, Size Zheng, Tao Wang, Tian Pei, Tian Yuan, Tianyu Sun, W. L. Xiao, Wangding Zeng, Wei An, Wen Liu, Wenfeng Liang, Wenjun Gao, Wentao Zhang, X. Q. Li, Xiangyue Jin, Xianzu Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaojin Shen, Xiaokang Chen, Xiaosha Chen, Xiaotao Nie, and Xiaowen Sun · 2024
Closest in time.
MASTERKEY: automated jailbreaking of large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2024
Closest in time.
Multilingual jailbreak challenges in large language models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2024
Closest in time.
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang · 2024
Closest in time.
To grok or not to grok: Disentangling generalization and memorization on corrupted algorithmic datasets
Darshil Doshi, Aritra Das, Tianyu He, and Andrey Gromov · 2024
Closest in time.
Denevil: towards deciphering and navigating the ethical values of large language models via instruction learning
Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, and Ning Gu · 2024
Closest in time.
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurélien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Rozière, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, Grégoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel M. Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, and Kevin Stone · 2024
Closest in time.
Risks and opportunities of open-source generative AI
Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schröder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Aaron Purewal, Botos Csaba, Fabro Steibel, Fazel Keshtkar, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan Arturo Nolazco, Lori Landay, Matthew Jackson, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, and Jakob N. Foerster · 2024
Closest in time.
Shangbin Feng, Herun Wan, Ningnan Wang, Zhaoxuan Tan, Minnan Luo, and Yulia Tsvetkov · 2024
Closest in time.
Towards trustworthy AI: A review of ethical and robust large language models
Md Meftahul Ferdaus, Mahdi Abdelguerfi, Elias Ioup, Kendall N. Niles, Ken Pathak, and Steven Sloan · 2024
Closest in time.
Ought we align the values of artificial moral agents?
Erez Firt · 2024
Closest in time.
Cross-task defense: Instruction-tuning llms for content safety
Yu Fu, Wen Xiao, Jia Chen, Jiachen Li, Evangelos E. Papalexakis, Aichi Chien, and Yue Dong · 2024
Closest in time.
Turning Vision into Action: Implementing the Senate AI Roadmap - Future of Life Institute — futureoflife.org
Future of Life Institute · 2024
Closest in time.
Application of LLM agents in recruitment: A novel framework for resume screening
Chengguang Gan, Qinghao Zhang, and Tatsunori Mori · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Closest in time.
Attacking large language models with projected gradient descent
Simon Geisler, Tom Wollschläger, M. H. I. Abdalla, Johannes Gasteiger, and Stephan Günnemann · 2024
Closest in time.
Shashwat Goel, Ameya Prabhu, Philip Torr, Ponnurangam Kumaraguru, and Amartya Sanyal · 2024
Closest in time.
Ctooleval: A chinese benchmark for llm-powered agent evaluation in real-world API interactions
Zishan Guo, Yufei Huang, and Deyi Xiong · 2024
Closest in time.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2024
Closest in time.
Overthinking the truth: Understanding how language models process false demonstrations
Danny Halawi, Jean-Stanislas Denain, and Jacob Steinhardt · 2024
Closest in time.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri · 2024
Closest in time.
Divij Handa, Zehua Zhang, Amir Saeidi, and Chitta Baral · 2024
Closest in time.
Machine-made media: Monitoring the mobilization of machine-generated articles on misinformation and mainstream news websites
Hans W. A. Hanley and Zakir Durumeric · 2024
Closest in time.
Pruning for protection: Increasing jailbreak resistance in aligned llms without fine-tuning
Adib Hasan, Ileana Rugina, and Alex Wang · 2024
Closest in time.
Efficient LLM jailbreak via adaptive dense-to-sparse constrained optimization
Kai Hu, Weichen Yu, Tianjun Yao, Xiang Li, Wenhe Liu, Lijun Yu, Yining Li, Kai Chen, Zhiqiang Shen, and Matt Fredrikson · 2024
Closest in time.
Catastrophic jailbreak of open-source llms via exploiting generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen · 2024
Closest in time.
CBBQ: A chinese bias benchmark dataset curated with human-ai collaboration for large language models
Yufei Huang and Deyi Xiong · 2024
Closest in time.
IT2ACL learning easy-to-hard instructions via 2-phase automated curriculum learning for large language models
Yufei Huang and Deyi Xiong · 2024
Closest in time.
Iti vision 2030: Eu artificial intelligence policy
ITIC · 2024
Closest in time.
Devansh Jain, Priyanshu Kumar, Samuel Gehman, Xuhui Zhou, Thomas Hartvigsen, and Maarten Sap · 2024
Closest in time.
Improved techniques for optimization-based jailbreaking on large language models
Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang, Jindong Gu, Yang Liu, Xiaochun Cao, and Min Lin · 2024
Closest in time.
Quack: Automatic jailbreaking large language models via role-playing, 2024
Haibo Jin, Ruoxi Chen, Jinyin Chen, and Haohan Wang · 2024
Closest in time.
Eagle: Ethical dataset given from real interactions
Masahiro Kaneko, Danushka Bollegala, and Timothy Baldwin · 2024
Closest in time.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2024
Closest in time.
Do moral judgment and reasoning capability of llms change with language? A study using the multilingual defining issues test
Aditi Khandelwal, Utkarsh Agarwal, Kumar Tanmay, and Monojit Choudhury · 2024
Closest in time.
San Kim and Gary Geunbae Lee · 2024
Closest in time.
On the reliability of watermarks for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein · 2024
Closest in time.
Adversarial attacks and defenses for large language models (llms): methods, frameworks & challenges
Pranjal Kumar · 2024
Closest in time.
Vishal Kumar, Zeyi Liao, Jaylen Jones, and Huan Sun · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea · 2024
Closest in time.
RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash · 2024
Closest in time.
Instructpatentgpt: Training patent language models to follow instructions with human feedback
Jieh-Sheng Lee · 2024
Closest in time.
The WMDP benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Gabriel Mukobi, Nathan Helm-Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Ariel Herbert-Voss, Cort B. Breuer, Andy Zou, Mantas Mazeika, Zifan Wang, Palash Oswal, Weiran Lin, Adam A. Hunt, Justin Tienken-Harder, Kevin Y. Shih, Kemper Talley, John Guan, Ian Steneker, David Campbell, Brad Jokubaitis, Steven Basart, Stephen Fitz, Ponnurangam Kumaraguru, Kallol Krishna Karmakar, Uday Kiran Tupakula, Vijay Varadharajan, Yan Shoshitaishvili, Jimmy Ba, Kevin M. Esvelt, Alexandr Wang, and Dan Hendrycks · 2024
Closest in time.
Drattack: Prompt decomposition and reconstruction makes powerful llms jailbreakers
Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, and Cho-Jui Hsieh · 2024
Closest in time.
Zeyi Liao and Huan Sun · 2024
Closest in time.
Sue Lim, Ralf Schmälzle, and Gary Bente · 2024
Closest in time.
An unforgeable publicly verifiable watermark for large language models
Aiwei Liu, Leyi Pan, Xuming Hu, Shuang Li, Lijie Wen, Irwin King, and Philip S. Yu · 2024
Closest in time.
A semantic invariant robust watermark for large language models
Aiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng, and Lijie Wen · 2024
Closest in time.
Chain of hindsight aligns language models with feedback
Hao Liu, Carmelo Sferrazza, and Pieter Abbeel · 2024
Closest in time.
Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction
Tong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong, Guozhu Meng, and Kai Chen · 2024
Closest in time.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2024
Closest in time.
Towards safer large language models through machine unlearning
Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang · 2024
Closest in time.
An adversarial perspective on machine unlearning for ai safety
Jakub Łucki, Boyi Wei, Yangsibo Huang, Peter Henderson, Florian Tramèr, and Javier Rando · 2024
Closest in time.
Jailbreak instruction-tuned llms via end-of-sentence MLP re-weighting
Yifan Luo, Zhennan Zhou, Meitan Wang, and Bin Dong · 2024
Closest in time.
Codechameleon: Personalized encryption framework for jailbreaking large language models
Huijie Lv, Xiao Wang, Yuansen Zhang, Caishuang Huang, Shihan Dou, Junjie Ye, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms
Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
Autonomy evaluation resources, 2024
METR · 2024
Closest in time.
Opening the AI black box: program synthesis via mechanistic interpretability
Eric J. Michaud, Isaac Liao, Vedang Lad, Ziming Liu, Anish Mudide, Chloe Loughridge, Zifan Carl Guo, Tara Rezaei Kheirkhah, Mateja Vukelic, and Max Tegmark · 2024
Closest in time.
Large language models in healthcare and medical domain: A review
Zabir Al Nazi and Wei Peng · 2024
Closest in time.
How to catch an AI liar: Lie detection in black-box llms by asking unrelated questions
Lorenzo Pacchiardi, Alex James Chan, Sören Mindermann, Ilan Moscovitz, Alexa Y. Pan, Yarin Gal, Owain Evans, and Jan Markus Brauner · 2024
Closest in time.
AI deception: A survey of examples, risks, and potential solutions
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2024
Closest in time.
Advprompter: Fast adaptive adversarial prompting for llms
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Grégoire Delétang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca D. Dragan, Rohin Shah, Allan Dafoe, and Toby Shevlane · 2024
Closest in time.
Llm self defense: By self examination, llms know they are being tricked, 2024
Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Peng, Sebastian Szyller, Cory Cornelius, and Duen Horng Chau · 2024
Closest in time.
Why does new knowledge create messy ripple effects in llms?
Jiaxin Qin, Zixuan Zhang, Chi Han, Pengfei Yu, Manling Li, and Heng Ji · 2024
Closest in time.
From prejudice to parity: A new approach to debiasing large language model word embeddings
Aishik Rakshit, Smriti Singh, Shuvam Keshari, Arijit Ghosh Chowdhury, Vinija Jain, and Aman Chadha · 2024
Closest in time.
Tricking llms into disobedience: Formalizing, analyzing, and detecting jailbreaks
Abhinav Rao, Atharva Naik, Sachin Vashistha, Somak Aditya, and Monojit Choudhury · 2024
Closest in time.
Anyone can audit! users can lead their own algorithmic audits with indielabel, 2024
AI Risk and Vulnerability Alliance · 2024
Closest in time.
Escalation risks from language models in military and diplomatic decision-making
Juan Pablo Rivera, Gabriel Mukobi, Anka Reuel, Max Lamparth, Chandler Smith, and Jacquelyn Schneider · 2024
Closest in time.
Identifying the risks of LM agents with an lm-emulated sandbox
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto · 2024
Closest in time.
Large language models show human-like social desirability biases in survey responses
Aadesh Salecha, Molly E. Ireland, Shashanka Subrahmanya, João Sedoc, Lyle H. Ungar, and Johannes C. Eichstaedt · 2024
Closest in time.
From principles to rules: A regulatory approach for frontier AI
Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler, and Ben Garfinkel · 2024
Closest in time.
Ai safety strategies landscape, 2024
Charbel-Raphael Segerie · 2024
Closest in time.
Locating and editing factual associations in mamba
Arnab Sen Sharma, David Atkinson, and David Bau · 2024
Closest in time.
CORECODE: A common sense annotated dialogue dataset with benchmark tasks for chinese large language models
Dan Shi, Chaobin You, Jiantao Huang, Taihao Li, and Deyi Xiong · 2024
Closest in time.
Criskeval: A chinese multi-level risk evaluation benchmark dataset for large language models
Ling Shi and Deyi Xiong · 2024
Closest in time.
Position: A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell L. Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi · 2024
Closest in time.
Beyond memorization: Violating privacy via inference with large language models
Robin Staab, Mark Vero, Mislav Balunovic, and Martin T. Vechev · 2024
Closest in time.
On the quest for effectiveness in human oversight: Interdisciplinary perspectives
Sarah Sterz, Kevin Baum, Sebastian Biewer, Holger Hermanns, Anne Lauber-Rönsberg, Philip Meinel, and Markus Langer · 2024
Closest in time.
Localizing paragraph memorization in language models
Niklas Stoehr, Mitchell Gordon, Chiyuan Zhang, and Owen Lewis · 2024
Closest in time.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter · 2024
Closest in time.
Decentralized neural networks, December 2024
Richard Sutton · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton · 2024
Closest in time.
Toward self-improvement of llms via imagination, searching, and criticizing
Ye Tian, Baolin Peng, Linfeng Song, Lifeng Jin, Dian Yu, Haitao Mi, and Dong Yu · 2024
Closest in time.
Team iimasnlp at PAN: leveraging graph neural networks and large language models for generative AI authorship verification
Andric Valdez-Valenzuela and Helena Gómez-Adorno · 2024
Closest in time.
Neeraj Varshney, Pavel Dolin, Agastya Seth, and Chitta Baral · 2024
Closest in time.
Badllama 3: removing safety finetuning from llama 3 in minutes
Dmitrii Volkov · 2024
Closest in time.
Do-not-answer: Evaluating safeguards in LLMs
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin · 2024
Closest in time.
The reasons that agents act: Intention and instrumental goals
Francis Rhys Ward, Matt MacDermott, Francesco Belardinelli, Francesca Toni, and Tom Everitt · 2024
Closest in time.
Mitigating privacy seesaw in large language models: Augmented privacy neuron editing via activation patching
Xinwei Wu, Weilong Dong, Shaoyang Xu, and Deyi Xiong · 2024
Closest in time.
Distract large language models for automatic jailbreak attack
Zeguan Xiao, Yan Yang, Guanhua Chen, and Yun Chen · 2024
Closest in time.
Gradsafe: Detecting unsafe prompts for llms via safety-critical gradient analysis
Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Zhenqiang Gong · 2024
Closest in time.
Defensive prompt patch: A robust and interpretable defense of llms against jailbreak attacks
Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho · 2024
Closest in time.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Nan Xu, Fei Wang, Ben Zhou, Bangzheng Li, Chaowei Xiao, and Muhao Chen · 2024
Closest in time.
Language agents with reinforcement learning for strategic play in the werewolf game
Zelai Xu, Chao Yu, Fei Fang, Yu Wang, and Yi Wu · 2024
Closest in time.
A comprehensive study of jailbreak attack versus defense for large language models
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek · 2024
Closest in time.
On protecting the data privacy of large language models (llms): A survey
Biwei Yan, Kun Li, Minghui Xu, Yueyan Dong, Yue Zhang, Zhaochun Ren, and Xiuzhen Cheng · 2024
Closest in time.
Poisonprompt: Backdoor attack on prompt-based large language models
Hongwei Yao, Jian Lou, and Zhan Qin · 2024
Closest in time.
Benchmarking knowledge boundary for large language models: A different perspective on model evaluation
Xunjian Yin, Xu Zhang, Jie Ruan, and Xiaojun Wan · 2024
Closest in time.
Yi: Open foundation models by 01.ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai · 2024
Closest in time.
CMoralEval: A moral evaluation benchmark for Chinese large language models
Linhao Yu, Yongqi Leng, Yufei Huang, Shang Wu, Haixin Liu, Xinmeng Ji, Jiahui Zhao, Jinwang Song, Tingting Cui, Xiaoqing Cheng, Tao Liu, and Deyi Xiong · 2024
Closest in time.
GPT-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.
AI risk categorization decoded (AIR 2024): From government regulations to corporate policies
Yi Zeng, Kevin Klyman, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, and Bo Li · 2024
Closest in time.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge AI safety by humanizing llms
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Closest in time.
Op-ed: A vision for the ai office - rethinking digital governance in the eu
Kai Zenner · 2024
Closest in time.
A short summary of evaluatology: The science and engineering of evaluation, 2024
Jianfeng Zhan · 2024
Closest in time.
Evaluatology: The science and engineering of evaluation
Jianfeng Zhan, Lei Wang, Wanling Gao, Hongxiao Li, Chenxi Wang, Yunyou Huang, Yatao Li, Zhengxin Yang, Guoxin Kang, Chunjie Luo, Hainan Ye, Shaopeng Dai, and Zhifei Zhang · 2024
Closest in time.
Boosting jailbreak attack with momentum
Yihao Zhang and Zeming Wei · 2024
Closest in time.
Psysafe: A comprehensive framework for psychological-based attack, defense, and evaluation of multi-agent system safety
Zaibin Zhang, Yongting Zhang, Lijun Li, Jing Shao, Hongzhi Gao, Yu Qiao, Lijun Wang, Huchuan Lu, and Feng Zhao · 2024
Closest in time.
Learning and forgetting unsafe examples in large language models
Jiachen Zhao, Zhun Deng, David Madras, James Zou, and Mengye Ren · 2024
Closest in time.
Opencodeinterpreter: Integrating code generation with execution and refinement
Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue · 2024
Closest in time.
HAZARD challenge: Embodied decision making in dynamically changing environments
Qinhong Zhou, Sunli Chen, Yisong Wang, Haozhe Xu, Weihua Du, Hongxin Zhang, Yilun Du, Joshua B. Tenenbaum, and Chuang Gan · 2024
Closest in time.
Critical data size of language models from a grokking perspective
Xuekai Zhu, Yao Fu, Bowen Zhou, and Zhouhan Lin · 2024
Closest in time.