Fetching the paper…
Reading the bibliography…
Large language models (LLMs), exemplified by ChatGPT, have gained considerable attention for their excellent natural language processing capabilities.
A theory of objective self awareness
Shelley Duval and Robert A Wicklund · 1972
Earlier work this paper cites.
How to generate and exchange secrets
Andrew Chi-Chih Yao · 1986
Earlier work this paper cites.
Accountability and control: Government officials and the exercise of power
Ian Thynne and John Goldring · 1987
Earlier work this paper cites.
An introduction to categorical data analysis
Alan Agresti · 1990
Earlier work this paper cites.
The levels of emotional awareness scale: A cognitive-developmental measure of emotion
Richard D Lane, Donald M Quinlan, Gary E Schwartz, Pamela A Walker, and Sharon B Zeitlin · 1990
Earlier work this paper cites.
An investigation of the therac-25 accidents
Nancy G Leveson and Clark S Turner · 1993
Earlier work this paper cites.
The study of pidgin and creole languages
Pieter Muysken, Norval Smith, et al · 1995
Earlier work this paper cites.
Accountability in a computerized society
Helen Nissenbaum · 1996
Earlier work this paper cites.
Learning in the presence of concept drift and hidden contexts
Gerhard Widmer and Miroslav Kubat · 1996
Earlier work this paper cites.
47 U.S.C. § 230, 1996
Protection for private blocking and screening of offensive material · 1996
Earlier work this paper cites.
‘accountability’: an ever-expanding concept?
Richard Mulgan · 2000
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Hidetoshi Shimodaira · 2000
Earlier work this paper cites.
Machines and mindlessness: Social responses to computers
Clifford Nass and Youngme Moon · 2000
Earlier work this paper cites.
Ehud reiter and robert dale. building natural language generation systems. cambridge university press, 2000
ADVAITH SIDDHARTHAN · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
What is transparency?
Richard W. Oliver · 2004
Earlier work this paper cites.
Guest editors’ introduction: machine ethics
Michael Anderson and Susan Leigh Anderson · 2006
Earlier work this paper cites.
The nature, importance, and difficulty of machine ethics
James H Moor · 2006
Earlier work this paper cites.
Promoting cultural diversity and cultural competency: self-assessment checklist for personnel providing behavioral health services and supports to children, youth and their families
T Goode · 2006
Earlier work this paper cites.
Machine ethics: Creating an ethical intelligent agent
Michael Anderson and Susan Leigh Anderson · 2007
Earlier work this paper cites.
Machine morality: bottom-up and top-down approaches for modelling human moral faculties
Wendell Wallach, Colin Allen, and Iva Smit · 2008
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence · 2008
Earlier work this paper cites.
Internet filtering and censorship
Samir N Hamade · 2008
Earlier work this paper cites.
Political Research Quarterly
Religious stereotyping and voter support for evangelical candidates · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Four kinds of ethical robots
James Moor et al · 2009
Earlier work this paper cites.
Two concepts of accountability: Accountability as a virtue and as a mechanism
Mark Bovens · 2010
Earlier work this paper cites.
Transfer learning
Lisa Torrey and Jude Shavlik · 2010
Earlier work this paper cites.
Self-awareness part 1: Definition, measures, effects, functions, and antecedents
Alain Morin · 2011
Earlier work this paper cites.
A unifying view on dataset shift in classification
Jose G Moreno-Torres, Troy Raeder, Rocío Alaiz-Rodríguez, Nitesh V Chawla, and Francisco Herrera · 2012
Earlier work this paper cites.
An overview of the schwartz theory of basic values
Shalom H Schwartz · 2012
Earlier work this paper cites.
Accountability: A synthesis
Essien E Akpanuko and Ikenna E Asogwa · 2013
Earlier work this paper cites.
Mapping accountability: core concept and subtypes
Staffan I Lindberg · 2013
Earlier work this paper cites.
Understanding the complex dynamics of transparency
Albert Meijer · 2013
Earlier work this paper cites.
Transparent predictions
Tal Z Zarsky · 2013
Earlier work this paper cites.
Enron email dataset, 2015
CMU · 2015
Earlier work this paper cites.
Accountable algorithms
Joshua Alexander Kroll · 2015
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
A survey of transfer learning
Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Doc: Deep open classification of text documents
Lei Shu, Hu Xu, and Bing Liu · 2017
Earlier work this paper cites.
Training confidence-calibrated classifiers for detecting out-of-distribution samples
Kimin Lee, Honglak Lee, Kibok Lee, and Jinwoo Shin · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Reluplex: An efficient smt solver for verifying deep neural networks
Guy Katz, Clark Barrett, David L Dill, Kyle Julian, and Mykel J Kochenderfer · 2017
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Algorithmic decision-making based on machine learning from big data: can transparency restore accountability?
Paul B De Laat · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning · 2018
Earlier work this paper cites.
Gender stereotypes
Naomi Ellemers · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods, 2018
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Scott Sorensen, Nithum Thain, and Lucy Vasserman · 2018
Earlier work this paper cites.
Adversarial over-sensitivity and over-stability strategies for dialogue models, 2018
Tong Niu and Mohit Bansal · 2018
Earlier work this paper cites.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin · 2018
Earlier work this paper cites.
Deep visual domain adaptation: A survey
Mei Wang and Weihong Deng · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R Bowman, and Noah A Smith · 2018
Earlier work this paper cites.
Multinomial adversarial networks for multi-domain text classification
Xilun Chen and Claire Cardie · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman · 2018
Earlier work this paper cites.
Output transparency vs. input transparency
Cass R Sunstein · 2018
Earlier work this paper cites.
Efficient neural network robustness certification with general activation functions
Huan Zhang, Tsui-Wei Weng, Pin-Yu Chen, Cho-Jui Hsieh, and Luca Daniel · 2018
Earlier work this paper cites.
Provable defenses against adversarial examples via the convex outer adversarial polytope
Eric Wong and Zico Kolter · 2018
Earlier work this paper cites.
On evaluating adversarial robustness, 2019
Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers, 2019
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Codah: An adversarially authored question-answer dataset for common sense
Michael Chen, Mike D’Arcy, Alisa Liu, Jared Fernandez, and Doug Downey · 2019
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2019
Earlier work this paper cites.
Avoiding reasoning shortcuts: Adversarial evaluation, training, and model development for multi-hop qa, 2019
Yichen Jiang and Mohit Bansal · 2019
Earlier work this paper cites.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Earlier work this paper cites.
Towards empathetic open-domain conversation models: a new benchmark and dataset, 2019
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
A right to reasonable inferences: re-thinking data protection law in the age of big data and ai
Sandra Wachter and Brent Mittelstadt · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
An abstract domain for certifying neural networks
Gagandeep Singh, Timon Gehr, Markus Püschel, and Martin Vechev · 2019
Earlier work this paper cites.
Provably robust boosted decision stumps and trees against adversarial attacks
Maksym Andriushchenko and Matthias Hein · 2019
Earlier work this paper cites.
Robustness verification of tree-based models
Hongge Chen, Huan Zhang, Si Si, Yang Li, Duane Boning, and Cho-Jui Hsieh · 2019
Earlier work this paper cites.
Scalable verified training for provably robust image classification
Sven Gowal, Krishnamurthy Dj Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy Mann, and Pushmeet Kohli · 2019
Earlier work this paper cites.
Towards stable and efficient training of verifiably robust neural networks
Huan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal, Robert Stanforth, Bo Li, Duane Boning, and Cho-Jui Hsieh · 2019
Earlier work this paper cites.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang · 2019
Earlier work this paper cites.
Achieving verified robustness to symbol substitutions via interval bound propagation
Po-Sen Huang, Robert Stanforth, Johannes Welbl, Chris Dyer, Dani Yogatama, Sven Gowal, Krishnamurthy Dvijotham, and Pushmeet Kohli · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models, 2020
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems, 2020
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2020
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Principled artificial intelligence: Mapping consensus in ethical and rights-based approaches to principles for ai
Jessica Fjeld, Nele Achten, Hannah Hilligoss, Ádám Nagy, and Madhulika Srikumar · 2020
Earlier work this paper cites.
One explanation does not fit all: The promise of interactive explanations for machine learning transparency
Kacper Sokol and Peter Flach · 2020
Earlier work this paper cites.
Explainable ai: A review of machine learning interpretability methods
Pantelis Linardatos, Vasilis Papastefanopoulos, and Sotiris Kotsiantis · 2020
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Beat the ai: Investigating adversarial human annotation for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp · 2020
Earlier work this paper cites.
Climate-fever: A dataset for verification of real-world climate claims
Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold · 2020
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi · 2020
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman · 2020
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2020
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang · 2020
Earlier work this paper cites.
On measuring and mitigating biased inferences of word embeddings
Sunipa Dev, Tao Li, Jeff M. Phillips, and Vivek Srikumar · 2020
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding, 2020
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Earlier work this paper cites.
Adv-bert: Bert is not robust on misspellings! generating nature adversarial samples on bert, 2020
Lichao Sun, Kazuma Hashimoto, Wenpeng Yin, Akari Asai, Jia Li, Philip Yu, and Caiming Xiong · 2020
Earlier work this paper cites.
Anomalous example detection in deep learning: A survey
Saikiran Bulusu, Bhavya Kailkhura, Bo Li, Pramod K Varshney, and Dawn Song · 2020
Earlier work this paper cites.
Stable learning via sample reweighting
Zheyan Shen, Peng Cui, Tong Zhang, and Kun Kunag · 2020
Earlier work this paper cites.
A comprehensive survey on transfer learning
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He · 2020
Earlier work this paper cites.
Towards transparency by design for artificial intelligence
Heike Felzmann, Eduard Fosch-Villaronga, Christoph Lutz, and Aurelia Tamò-Larrieux · 2020
Earlier work this paper cites.
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al · 2020
Earlier work this paper cites.
Human factors in model interpretability: Industry practices, challenges, and needs
Sungsoo Ray Hong, Jessica Hullman, and Enrico Bertini · 2020
Earlier work this paper cites.
Designing robots for care: Care centered value-sensitive design
Aimee Van Wynsberghe · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
The court recognized the ai-generated content as a work and entitled to copyright, 2020
2020
Earlier work this paper cites.
Automatic perturbation analysis for scalable certified robustness and beyond
Kaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang, Kai-Wei Chang, Minlie Huang, Bhavya Kailkhura, Xue Lin, and Cho-Jui Hsieh · 2020
Earlier work this paper cites.
Branch and bound for piecewise linear neural network verification
Rudy Bunel, Jingyue Lu, Ilker Turkaslan, Philip HS Torr, Pushmeet Kohli, and M Pawan Kumar · 2020
Earlier work this paper cites.
SAFER: A structure-free approach for certified robustness to adversarial word substitutions
Mao Ye, Chengyue Gong, and Qiang Liu · 2020
Earlier work this paper cites.
State-of-the-art artificial intelligence techniques for distributed smart grids: A review
Syed Saqib Ali and Bong Jun Choi · 2020
Earlier work this paper cites.
Applications of artificial intelligence for disaster management
Wenjuan Sun, Paolo Bocchini, and Brian D Davison · 2020
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Evaluation of text generation: A survey, 2021
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao · 2021
Earlier work this paper cites.
The gem benchmark: Natural language generation, its evaluation and metrics, 2021
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D. Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Rubungo Andre Niyongabo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou · 2021
Earlier work this paper cites.
Federated reconstruction: Partially local federated learning, 2021
Karan Singhal, Hakim Sidahmed, Zachary Garrett, Shanshan Wu, Keith Rush, and Sushant Prakash · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models (2021)
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Earlier work this paper cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Earlier work this paper cites.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan · 2021
Earlier work this paper cites.
A framework for understanding sources of harm throughout the machine learning life cycle
Harini Suresh and John Guttag · 2021
Earlier work this paper cites.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li · 2021
Earlier work this paper cites.
Privacy regularization: Joint privacy-utility optimization in language models
Fatemehsadat Mireshghallah, Huseyin A Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, and Robert Sim · 2021
Earlier work this paper cites.
A word on machine ethics: A response to jiang et al.(2021)
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams · 2021
Earlier work this paper cites.
Nine potential pitfalls when designing human-ai co-creative systems
Daniel Buschek, Lukas Mecke, Florian Lehmann, and Hai Dang · 2021
Earlier work this paper cites.
What do we want from explainable artificial intelligence (xai)?–a stakeholder perspective on xai and a conceptual model guiding interdisciplinary xai research
Markus Langer, Daniel Oster, Timo Speith, Holger Hermanns, Lena Kästner, Eva Schmidt, Andreas Sesing, and Kevin Baum · 2021
Earlier work this paper cites.
Beyond expertise and roles: A framework to characterize the stakeholders of interpretable machine learning and their needs
Harini Suresh, Steven R Gomez, Kevin K Nam, and Arvind Satyanarayan · 2021
Earlier work this paper cites.
Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking
Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, and Douwe Kiela · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A. Smith, and Mike Lewis · 2021
Earlier work this paper cites.
COVID-fact: Fact extraction and verification of real-world claims on COVID-19 pandemic
Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan · 2021
Earlier work this paper cites.
Evidence-based fact-checking of health-related claims
Mourad Sarrouti, Asma Ben Abacha, Yassine M’rabet, and Dina Demner-Fushman · 2021
Earlier work this paper cites.
What makes good in-context examples for gpt- 3 3 ?
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen · 2021
Earlier work this paper cites.
Learning to retrieve prompts for in-context learning
Ohad Rubin, Jonathan Herzig, and Jonathan Berant · 2021
Earlier work this paper cites.
Infosurgeon: Cross-media fine-grained information consistency checking for fake news detection
Yi Fung, Christopher Thomas, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji, Shih-Fu Chang, Kathleen McKeown, Mohit Bansal, and Avi Sil · 2021
Earlier work this paper cites.
Probing toxic content in large pre-trained language models
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit-Yan Yeung · 2021
Earlier work this paper cites.
Can machines learn morality? the delphi experiment
Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, et al · 2021
Earlier work this paper cites.
Understanding the capabilities, limitations, and societal impact of large language models
Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli · 2021
Earlier work this paper cites.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models, 2021
Hannah Kirk, Yennie Jun, Haider Iqbal, Elias Benussi, Filippo Volpin, Frederic A. Dreyer, Aleksandar Shtedritski, and Yuki M. Asano · 2021
Earlier work this paper cites.
Persistent anti-muslim bias in large language models, 2021
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Earlier work this paper cites.
On measures of biases and harms in nlp
Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, et al · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Earlier work this paper cites.
Robustness gym: Unifying the NLP evaluation landscape
Karan Goel, Nazneen Fatema Rajani, Jesse Vig, Zachary Taschdjian, Mohit Bansal, and Christopher Ré · 2021
Earlier work this paper cites.
OoD-Bench: Benchmarking and understanding out-of-distribution generalization datasets and algorithms
Nanyang Ye, Kaican Li, Lanqing Hong, Haoyue Bai, Yiting Chen, Fengwei Zhou, and Zhenguo Li · 2021
Earlier work this paper cites.
Maxime Peyrard, Sarvjeet Singh Ghotra, Martin Josifoski, Vidhan Agarwal, Barun Patra, Dean Carignan, Emre Kiciman, and Robert West · 2021
Earlier work this paper cites.
Generalized out-of-distribution detection: A survey
Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu · 2021
Earlier work this paper cites.
Towards out-of-distribution generalization: A survey
Zheyan Shen, Jiashuo Liu, Yue He, Xingxuan Zhang, Renzhe Xu, Han Yu, and Peng Cui · 2021
Earlier work this paper cites.
Learning models with uniform performance via distributionally robust optimization
John C Duchi and Hongseok Namkoong · 2021
Earlier work this paper cites.
Heterogeneous risk minimization
Jiashuo Liu, Zheyuan Hu, Peng Cui, Bo Li, and Zheyan Shen · 2021
Earlier work this paper cites.
Exploring the efficacy of automatically generated counterfactuals for sentiment analysis
Linyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang, Barry Smyth, and Ruihai Dong · 2021
Earlier work this paper cites.
Deep learning models are not robust against noise in clinical text
Milad Moradi, Kathrin Blagec, and Matthias Samwald · 2021
Earlier work this paper cites.
Robustness to spurious correlations in text classification via automatically generated counterfactuals
Zhao Wang and Aron Culotta · 2021
Earlier work this paper cites.
Combining feature and instance attribution to detect artifacts
Pouya Pezeshkpour, Sarthak Jain, Sameer Singh, and Byron C Wallace · 2021
Earlier work this paper cites.
Cross-lingual cross-domain nested named entity evaluation on english web texts
Barbara Plank · 2021
Earlier work this paper cites.
Keyword search based on unsupervised pre-trained acoustic models
Xiner Li, Jing Zhao, Wei-Qiang Zhang, Zhiqiang Lv, and Shen Huang · 2021
Earlier work this paper cites.
Measure and improve robustness in nlp models: A survey
Xuezhi Wang, Haohan Wang, and Diyi Yang · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
What would jiminy cricket do? towards agents that behave morally
Dan Hendrycks, Mantas Mazeika, Andy Zou, Sahil Patel, Christine Zhu, Jesus Navarro, Dawn Song, Bo Li, and Jacob Steinhardt · 2021
Earlier work this paper cites.
A survey on the explainability of supervised machine learning
Nadia Burkart and Marco F Huber · 2021
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2021
Earlier work this paper cites.
Differentiable prompt makes pre-trained language models better few-shot learners
Ningyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang, and Huajun Chen · 2021
Earlier work this paper cites.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg · 2021
Earlier work this paper cites.
Beta-crown: Efficient bound propagation with per-neuron split constraints for neural network robustness verification
Shiqi Wang, Huan Zhang, Kaidi Xu, Xue Lin, Suman Jana, Cho-Jui Hsieh, and J Zico Kolter · 2021
Earlier work this paper cites.
Fast certified robust training with short warmup
Zhouxing Shi, Yihan Wang, Huan Zhang, Jinfeng Yi, and Cho-Jui Hsieh · 2021
Earlier work this paper cites.
A toolkit for text extraction and analysis for natural language processing tasks
Tshephisho Joseph Sefara, Mahlatse Mbooi, Katlego Mashile, Thompho Rambuda, and Mapitsi Rangata · 2022
Earlier work this paper cites.
Wordcraft: story writing with large language models
Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito · 2022
Earlier work this paper cites.
Pathways: Asynchronous distributed dataflow for ml
Paul Barham, Aakanksha Chowdhery, Jeff Dean, Sanjay Ghemawat, Steven Hand, Daniel Hurt, Michael Isard, Hyeontaek Lim, Ruoming Pang, Sudip Roy, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, et al · 2022
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al · 2022
Earlier work this paper cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Large language models may leak personal data, 2022
Slator · 2022
Earlier work this paper cites.
What does it mean to align ai with human values?, 2022
Quanta Magazine · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
Generalizing to unseen domains: A survey on domain generalization, 2022
Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Tao Qin, Wang Lu, Yiqiang Chen, Wenjun Zeng, and Philip S. Yu · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark · 2022
Earlier work this paper cites.
Linyi Yang, Shuibai Zhang, Libo Qin, Yafu Li, Yidong Wang, Hanmeng Liu, Jindong Wang, Xing Xie, and Yue Zhang · 2022
Earlier work this paper cites.
Self-instruct: Aligning language model with self generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Earlier work this paper cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Earlier work this paper cites.
Moral mimicry: Large language models produce moral rationalizations tailored to political identity
Gabriel Simmons · 2022
Earlier work this paper cites.
Adversarial Robustness for Machine Learning
Pin-Yu Chen and Cho-Jui Hsieh · 2022
Earlier work this paper cites.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr · 2022
Earlier work this paper cites.
Are large pre-trained language models leaking your personal information?, 2022
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang · 2022
Earlier work this paper cites.
Ew-tune: A framework for privately fine-tuning large language models with differential privacy
Rouzbeh Behnia, Mohammadreza Reza Ebrahimi, Jason Pacheco, and Balaji Padmanabhan · 2022
Earlier work this paper cites.
Philip Feldman, Aaron Dant, and David Rosenbluth · 2022
Earlier work this paper cites.
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts
Tongshuang Wu, Michael Terry, and Carrie Jun Cai · 2022
Earlier work this paper cites.
Accountability in an algorithmic society: relationality, responsibility, and robustness in machine learning
A Feder Cooper, Emanuel Moss, Benjamin Laufer, and Helen Nissenbaum · 2022
Earlier work this paper cites.
Ddxplus: A new dataset for automatic medical diagnosis
Arsene Fansi Tchango, Rishab Goel, Zhi Wen, Julien Martel, and Joumana Ghosn · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Automatic chain of thought prompting in large language models, 2022
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola · 2022
Earlier work this paper cites.
Canyu Chen, Haoran Wang, Matthew Shapiro, Yunyu Xiao, Fei Wang, and Kai Shu · 2022
Earlier work this paper cites.
Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal · 2022
Earlier work this paper cites.
Attributed question answering: Evaluation and modeling for attributed large language models
Bernd Bohnet, Vinh Q Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, Kai Hui, et al · 2022
Earlier work this paper cites.
Improving cross-lingual fact checking with cross-lingual retrieval
Kung-Hsiang Huang, ChengXiang Zhai, and Heng Ji · 2022
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau · 2022
Earlier work this paper cites.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Shiyue Zhang, David Wan, and Mohit Bansal · 2022
Earlier work this paper cites.
Evaluating and improving factuality in multimodal abstractive summarization
David Wan and Mohit Bansal · 2022
Earlier work this paper cites.
Factpegasus: Factuality-aware pre-training and fine-tuning for abstractive summarization
David Wan and Mohit Bansal · 2022
Earlier work this paper cites.
Factgraph: Evaluating factuality in summarization with semantic graph representations
Leonardo FR Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, and Mohit Bansal · 2022
Earlier work this paper cites.
Evaluating the factual consistency of large language models through summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel · 2022
Earlier work this paper cites.
Exploring the limits of domain-adaptive training for detoxifying large-scale language models
Boxin Wang, Wei Ping, Chaowei Xiao, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Bo Li, Anima Anandkumar, and Bryan Catanzaro · 2022
Earlier work this paper cites.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang · 2022
Earlier work this paper cites.
https://old.reddit.com/r/ChatGPT/comments/zlcyr9/dan_is_my_new_friend/
Dan is my new friend, 2022 · 2022
Earlier work this paper cites.
Toxicity detection with generative prompt-based inference
Yau-Shian Wang and Yingshan Chang · 2022
Earlier work this paper cites.
On measures of biases and harms in nlp, 2022
Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, and Kai-Wei Chang · 2022
Earlier work this paper cites.
Out-of-distribution detection and selective generation for conditional language models
Jie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna, Mohammad Saleh, Balaji Lakshminarayanan, and Peter J Liu · 2022
Earlier work this paper cites.
Towards textual out-of-domain detection without in-domain labels
Di Jin, Shuyang Gao, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tür · 2022
Earlier work this paper cites.
Language models (mostly) know what they know, 2022
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan · 2022
Earlier work this paper cites.
GOOD: A graph out-of-distribution benchmark
Shurui Gui, Xiner Li, Limei Wang, and Shuiwang Ji · 2022
Earlier work this paper cites.
Extending the scope of out-of-domain: Examining qa models in multiple subdomains
Chenyang Lyu, Jennifer Foster, and Yvette Graham · 2022
Earlier work this paper cites.
Id and ood performance are sometimes inversely correlated on real-world datasets
Damien Teney, Yong Lin, Seong Joon Oh, and Ehsan Abbasnejad · 2022
Earlier work this paper cites.
Enabling classifiers to make judgements explicitly aligned with human values
Yejin Bang, Tiezheng Yu, Andrea Madotto, Zhaojiang Lin, Mona Diab, and Pascale Fung · 2022
Earlier work this paper cites.
Do multilingual language models capture differing moral norms?, 2022
Katharina Hämmerl, Björn Deiseroth, Patrick Schramowski, Jindřich Libovický, Alexander Fraser, and Kristian Kersting · 2022
Earlier work this paper cites.
A benchmark model for language models towards increased transparency
AyseKok Arslan · 2022
Earlier work this paper cites.
Interactive model cards: A human-centered approach to model documentation
Anamaria Crisan, Margaret Drouhard, Jesse Vig, and Nazneen Rajani · 2022
Earlier work this paper cites.
Kasia S Chmielinski, Sarah Newman, Matt Taylor, Josh Joseph, Kemi Thomas, Jessica Yurkofsky, and Yue Chelsea Qiu · 2022
Earlier work this paper cites.
Interpretability in the wild: A circuit for indirect object identification in gpt-2 small, 2022
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Predictability and surprise in large generative models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al · 2022
Earlier work this paper cites.
Natural language processing: State of the art, current trends and challenges
Diksha Khurana, Aditya Koli, Kiran Khatter, and Sukhdev Singh · 2023
Earlier work this paper cites.
Multilingual machine translation with large language models: Empirical results and analysis, 2023
Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li · 2023
Earlier work this paper cites.
https://blogs.microsoft.com/blog/2023/02/07/reinventing-search-with-a-new-ai-powered-microsoft-bing-and-edge-your-copilot-for-the-web/
Reinventing search with a new ai-powered microsoft bing and edge, your copilot for the web, 2023 · 2023
Earlier work this paper cites.
https://medium.com/whatnot-engineering/enhancing-search-using-large-language-models-f9dcb988bdb9
Enhancing search using large language models, 2023 · 2023
Earlier work this paper cites.
https://www.projectpro.io/article/large-language-model-use-cases-and-applications/887
7 top large language model use cases and applications, 2023 · 2023
Earlier work this paper cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Earlier work this paper cites.
Large language models: The future of b2b software, 2023
MintMesh · 2023
Earlier work this paper cites.
Bloomberggpt: A large language model for finance, 2023
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Earlier work this paper cites.
Scientific discovery in the age of artificial intelligence
Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, et al · 2023
Earlier work this paper cites.
The impact of large language models on scientific discovery: a preliminary study using gpt-4, 2023
Microsoft Research AI4Science and Microsoft Azure Quantum · 2023
Earlier work this paper cites.
The future landscape of large language models in medicine
Jan Clusmann, Fiona R Kolbinger, Hannah Sophie Muti, Zunamys I Carrero, Jan-Niklas Eckardt, Narmin Ghaffari Laleh, Chiara Maria Lavinia Löffler, Sophie-Caroline Schwarzkopf, Michaela Unger, Gregory P Veldhuizen, et al · 2023
Earlier work this paper cites.
Yuanhe Tian, Ruyi Gan, Yan Song, Jiaxing Zhang, and Yongdong Zhang · 2023
Earlier work this paper cites.
Alpacare:instruction-tuned large language models for medical application, 2023
Xinlu Zhang, Chenxin Tian, Xianjun Yang, Lichang Chen, Zekun Li, and Linda Ruth Petzold · 2023
Earlier work this paper cites.
Biomedgpt: A unified and generalist biomedical generative pre-trained transformer for vision, language, and multimodal tasks, 2023
Kai Zhang, Jun Yu, Zhiling Yan, Yixin Liu, Eashan Adhikarla, Sunyang Fu, Xun Chen, Chen Chen, Yuyin Zhou, Xiang Li, Lifang He, Brian D. Davison, Quanzheng Li, Yong Chen, Hongfang Liu, and Lichao Sun · 2023
Earlier work this paper cites.
Bianque: Balancing the questioning and suggestion ability of health llms with multi-turn health conversations polished by chatgpt, 2023
Yirong Chen, Zhenyu Wang, Xiaofen Xing, huimin zheng, Zhipei Xu, Kai Fang, Junhong Wang, Sihang Li, Jieling Wu, Qi Liu, and Xiangmin Xu · 2023
Earlier work this paper cites.
Huatuogpt, towards taming language models to be a doctor
Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, Xiang Wan, Benyou Wang, and Haizhou Li · 2023
Earlier work this paper cites.
Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge
Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang · 2023
Earlier work this paper cites.
Medicalgpt: Training medical gpt model
Ming Xu · 2023
Earlier work this paper cites.
A domain-specific next-generation large language model (llm) or chatgpt is required for biomedical engineering and research
Soumen Pal, Manojit Bhattacharya, Sang-Soo Lee, and Chiranjib Chakraborty · 2023
Cited alongside, same era.
Towards generalist biomedical ai
Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Chuck Lau, Ryutaro Tanno, Ira Ktena, et al · 2023
Cited alongside, same era.
Large language models and political science
Mitchell Linegar, Rafal Kocielnik, and R Michael Alvarez · 2023
Cited alongside, same era.
https://github.com/irlab-sdu/fuzi.mingcha , 2023
fuzi.mingcha · 2023
Cited alongside, same era.
Disc-lawllm: Fine-tuning large language models for intelligent legal services, 2023
Shengbin Yue, Wei Chen, Siyuan Wang, Bingxuan Li, Chenchen Shen, Shujun Liu, Yuxuan Zhou, Yao Xiao, Song Yun, Xuanjing Huang, and Zhongyu Wei · 2023
Cited alongside, same era.
Sight beyond text: Multi-modal training enhances llms in truthfulness and ethics
Haoqin Tu, Bingchen Zhao, Chen Wei, and Cihang Xie · 2023
Later among the works it cites.
Investigating online financial misinformation and its consequences: A computational perspective
Aman Rangapur, Haoran Wang, and Kai Shu · 2023
Later among the works it cites.
Yue Huang and Lichao Sun · 2023
Later among the works it cites.
Can llm-generated misinformation be detected?
Canyu Chen and Kai Shu · 2023
Later among the works it cites.
Answering questions by meta-reasoning over multiple chains of thought
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What can large language models do in chemistry? a comprehensive benchmark on eight tasks
Taicheng Guo, Kehan Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang · 2023
Cited alongside, same era.
Structured chemistry reasoning with large language models
Siru Ouyang, Zhuosheng Zhang, Bing Yan, Xuan Liu, Jiawei Han, and Lianhui Qin · 2023
Cited alongside, same era.
Marinegpt: Unlocking secrets of "ocean" to the public, 2023
Ziqiang Zheng, Jipeng Zhang, Tuan-Anh Vu, Shizhe Diao, Yue Him Wong Tim, and Sai-Kit Yeung · 2023
Cited alongside, same era.
Oceangpt: A large language model for ocean science tasks, 2023
Zhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, and Huajun Chen · 2023
Cited alongside, same era.
Taoli llama
Jingsi Yu, Junhui Zhu, Yujie Wang, Yang Liu, Hongxiang Chang, Jinran Nie, Cunliang Kong, Ruining Chong, XinLiu, Jiyuan An, Luming Lu, Mingwei Fang, and Lin Zhu · 2023
Cited alongside, same era.
Artgpt-4: Artistic vision-language understanding with adapter-enhanced minigpt-4, 2023
Zhengqing Yuan, Huiwen Xue, Xinyi Wang, Yongming Liu, Zhuanzhe Zhao, and Kun Wang · 2023
Cited alongside, same era.
Palm: Efficiently training massive language models, 2023
Towards Data Science · 2023
Cited alongside, same era.
Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, and Jonathan Berant · 2023
Later among the works it cites.
Ask me in english instead: Cross-lingual evaluation of large language models for healthcare queries
De Choudhury et al · 2023
Later among the works it cites.
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al · 2023
Later among the works it cites.
Faking fake news for real fake news detection: Propaganda-loaded training data generation
Kung-Hsiang Huang, Kathleen McKeown, Preslav Nakov, Yejin Choi, and Heng Ji · 2023
Later among the works it cites.
Fact-checking complex claims with program-guided reasoning
Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, and Preslav Nakov · 2023
Later among the works it cites.
Explainable claim verification via knowledge-grounded reasoning with large language models
Haoran Wang and Kai Shu · 2023
Later among the works it cites.
Zero-shot faithful factual error correction
Kung-Hsiang Huang, Hou Pong Chan, and Heng Ji · 2023
Later among the works it cites.
In-context retrieval-augmented language models
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham · 2023
Later among the works it cites.
Replug: Retrieval-augmented black-box language models
Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih · 2023
Later among the works it cites.
Active retrieval augmented generation
Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Later among the works it cites.
Long-range language modeling with self-retrieval
Ohad Rubin and Jonathan Berant · 2023
Later among the works it cites.
Ragtruth: A hallucination corpus for developing trustworthy retrieval-augmented language models, 2023
Yuanhao Wu, Juno Zhu, Siliang Xu, Kashun Shum, Cheng Niu, Randy Zhong, Juntong Song, and Tong Zhang · 2023
Later among the works it cites.
Knowledge editing for large language models: A survey
Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, et al · 2023
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Later among the works it cites.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2023
Later among the works it cites.
Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization, 2023
Zhiyuan Zhao, Bin Wang, Linke Ouyang, Xiaoyi Dong, Jiaqi Wang, and Conghui He · 2023
Later among the works it cites.
Delucionqa: Detecting hallucinations in domain-specific question answering, 2023
Mobashir Sadat, Zhengyu Zhou, Lukas Lange, Jun Araki, Arsalan Gundroo, Bingqing Wang, Rakesh R Menon, Md Rizwan Parvez, and Zhe Feng · 2023
Later among the works it cites.
On early detection of hallucinations in factual question answering, 2023
Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar · 2023
Later among the works it cites.
Don’t believe everything you read: Enhancing summarization interpretability through automatic identification of hallucinations in large language models, 2023
Priyesh Vakharia, Devavrat Joshi, Meenal Chavan, Dhananjay Sonawane, Bhrigu Garg, Parsa Mazaheri, and Ian Lane · 2023
Later among the works it cites.
Alleviating hallucinations of large language models through induced hallucinations, 2023
Yue Zhang, Leyang Cui, Wei Bi, and Shuming Shi · 2023
Later among the works it cites.
Reducing llm hallucinations using epistemic neural networks, 2023
Shreyas Verma, Kien Tran, Yusuf Ali, and Guangyu Min · 2023
Later among the works it cites.
Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models
Alfonso Amayuelas, Liangming Pan, Wenhu Chen, and William Wang · 2023
Later among the works it cites.
Shifting attention to relevance: Towards the uncertainty estimation of large language models
Jinhao Duan, Hao Cheng, Shiqi Wang, Chenan Wang, Alex Zavalny, Renjing Xu, Bhavya Kailkhura, and Kaidi Xu · 2023
Later among the works it cites.
Enhancing uncertainty-based hallucination detection with stronger focus, 2023
Tianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng, Yue Zhang, Zheng Zhang, Chenghu Zhou, Xinbing Wang, and Luoyi Fu · 2023
Later among the works it cites.
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu · 2023
Later among the works it cites.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark JF Gales · 2023
Later among the works it cites.
Mitigating language model hallucination with interactive question-knowledge alignment
Shuo Zhang, Liangming Pan, Junzhou Zhao, and William Yang Wang · 2023
Later among the works it cites.
Trusting your evidence: Hallucinate less with context-aware decoding
Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Scott Wen-tau Yih · 2023
Later among the works it cites.
Mitigating large language model hallucinations via autonomous knowledge graph-based retrofitting
Xinyan Guan, Yanjiang Liu, Hongyu Lin, Yaojie Lu, Ben He, Xianpei Han, and Le Sun · 2023
Later among the works it cites.
Chain-of-note: Enhancing robustness in retrieval-augmented language models, 2023
Wenhao Yu, Hongming Zhang, Xiaoman Pan, Kaixin Ma, Hongwei Wang, and Dong Yu · 2023
Later among the works it cites.
Fine-tuning language models for factuality, 2023
Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, and Chelsea Finn · 2023
Later among the works it cites.
Wikichat: Stopping the hallucination of large language model chatbots by few-shot grounding on wikipedia, 2023
Sina J. Semnani, Violet Z. Yao, Heidi C. Zhang, and Monica S. Lam · 2023
Later among the works it cites.
Faithfulness-aware decoding strategies for abstractive summarization
David Wan, Mengwen Liu, Kathleen McKeown, Markus Dreyer, and Mohit Bansal · 2023
Later among the works it cites.
Simple synthetic data reduces sycophancy in large language models
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V Le · 2023
Later among the works it cites.
When large language models contradict humans? large language models’ sycophantic behaviour, 2023
Leonardo Ranaldi and Giulia Pucci · 2023
Later among the works it cites.
Towards understanding sycophancy in language models, 2023
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Reducing sycophancy and improving honesty via activation steering, 2023
Nina Rimsky · 2023
Later among the works it cites.
The earth is flat because…: Investigating llms’ belief towards misinformation via persuasive conversation, 2023
Rongwu Xu, Brian S. Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu · 2023
Later among the works it cites.
Factuality enhanced language models for open-ended text generation, 2023
Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, and Bryan Catanzaro · 2023
Later among the works it cites.
Summon a demon and bind it: A grounded theory of llm red teaming in the wild, 2023
Nanna Inie, Jonathan Stray, and Leon Derczynski · 2023
Later among the works it cites.
Fake alignment: Are llms really aligned well?, 2023
Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, and Yingchun Wang · 2023
Later among the works it cites.
Can llms follow simple rules?, 2023
Norman Mu, Sarah Chen, Zifan Wang, Sizhe Chen, David Karamardian, Lulwa Aljeraisy, Dan Hendrycks, and David Wagner · 2023
Later among the works it cites.
Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global scale prompt hacking competition, 2023
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan Boyd-Graber · 2023
Later among the works it cites.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking, 2023
Nan Xu, Fei Wang, Ben Zhou, Bang Zheng Li, Chaowei Xiao, and Muhao Chen · 2023
Later among the works it cites.
Detecting language model attacks with perplexity, 2023
Gabriel Alon and Michael Kamfonas · 2023
Later among the works it cites.
Safety alignment in nlp tasks: Weakly aligned summarization as an in-context attack, 2023
Yu Fu, Yufei Li, Wen Xiao, Cong Liu, and Yue Dong · 2023
Later among the works it cites.
Causality analysis for evaluating the security of large language models, 2023
Wei Zhao, Zhe Li, and Jun Sun · 2023
Later among the works it cites.
Bypassing the safety training of open-source llms with priming attacks, 2023
Jason Vega, Isha Chaudhary, Changming Xu, and Gagandeep Singh · 2023
Later among the works it cites.
Benchmarking and defending against indirect prompt injection attacks on large language models, 2023
Jingwei Yi, Yueqi Xie, Bin Zhu, Keegan Hines, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu · 2023
Later among the works it cites.
Red teaming for large language models at scale: Tackling hallucinations on mathematics tasks, 2023
Aleksander Buszydlik, Karol Dobiczek, Michał Teodor Okoń, Konrad Skublicki, Philip Lippmann, and Jie Yang · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Later among the works it cites.
Certifying llm safety against adversarial prompting, 2023
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning, 2023
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang · 2023
Later among the works it cites.
Exploiting novel gpt-4 apis, 2023
Kellin Pelrine, Mohammad Taufeeque, Michał Zając, Euan McLean, and Adam Gleave · 2023
Later among the works it cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2023
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Later among the works it cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Later among the works it cites.
Unveiling safety vulnerabilities of large language models, 2023
George Kour, Marcel Zalmanovici, Naama Zwerdling, Esther Goldbraich, Ora Nova Fandina, Ateret Anaby-Tavor, Orna Raz, and Eitan Farchi · 2023
Later among the works it cites.
On the exploitability of instruction tuning, 2023
Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein · 2023
Later among the works it cites.
Adversarial demonstration attacks on large language models, 2023
Jiongxiao Wang, Zichen Liu, Keun Hee Park, Zhuojun Jiang, Zhaoheng Zheng, Zhuofeng Wu, Muhao Chen, and Chaowei Xiao · 2023
Later among the works it cites.
On the exploitability of reinforcement learning with human feedback for large language models, 2023
Jiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik, and Chaowei Xiao · 2023
Later among the works it cites.
Chatgpt as an attack tool: Stealthy textual backdoor attack via blackbox generative model trigger, 2023
Jiazhao Li, Yijin Yang, Zhuofeng Wu, V. G. Vinod Vydiswaran, and Chaowei Xiao · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback, 2023
Javier Rando and Florian Tramèr · 2023
Later among the works it cites.
Stealthy and persistent unalignment on large language models via backdoor injections, 2023
Yuanpu Cao, Bochuan Cao, and Jinghui Chen · 2023
Later among the works it cites.
Composite backdoor attacks against large language models, 2023
Hai Huang, Zhengyu Zhao, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
Poisonprompt: Backdoor attack on prompt-based large language models, 2023
Hongwei Yao, Jian Lou, and Zhan Qin · 2023
Later among the works it cites.
Large language models are better adversaries: Exploring generative clean-label backdoor attacks against text classifiers, 2023
Wencong You, Zayd Hammoudeh, and Daniel Lowd · 2023
Later among the works it cites.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models, 2023
Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen · 2023
Later among the works it cites.
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Later among the works it cites.
Punctuation matters! stealthy backdoor attack for language models, 2023
Xuan Sheng, Zhicheng Li, Zhaoyang Han, Xiangmao Chang, and Piji Li · 2023
Later among the works it cites.
Learning and forgetting unsafe examples in large language models, 2023
Jiachen Zhao, Zhun Deng, David Madras, James Zou, and Mengye Ren · 2023
Later among the works it cites.
On the safety of open-sourced large language models: Does alignment really prevent them from being misused?, 2023
Anonymous · 2023
Later among the works it cites.
Defending chatgpt against jailbreak attack via self-reminder
Fangzhao Wu, Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, and Xing Xie · 2023
Later among the works it cites.
Maatphor: Automated variant analysis for prompt injection attacks, 2023
Ahmed Salem, Andrew Paverd, and Boris Köpf · 2023
Later among the works it cites.
A mutation-based method for multi-modal jailbreaking attack detection, 2023
Xiaoyu Zhang, Cen Zhang, Tianlin Li, Yihao Huang, Xiaojun Jia, Xiaofei Xie, Yang Liu, and Chao Shen · 2023
Later among the works it cites.
Test-time backdoor mitigation for black-box large language models with defensive demonstrations
Wenjie Mo, Jiashu Xu, Qin Liu, Jiongxiao Wang, Jun Yan, Chaowei Xiao, and Muhao Chen · 2023
Later among the works it cites.
Efficient toxic content detection by bootstrapping and distilling large language models, 2023
Jiang Zhang, Qiong Wu, Yiming Xu, Cheng Cao, Zheng Du, and Konstantinos Psounis · 2023
Later among the works it cites.
Gta: Gated toxicity avoidance for lm performance preservation, 2023
Heegyu Kim and Hyunsouk Cho · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models, 2023
Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua · 2023
Later among the works it cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2023
Later among the works it cites.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2023
Later among the works it cites.
Multilingual jailbreak challenges in large language models, 2023
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models, 2023
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
https://www.perspectiveapi.com
Perspective api, 2023 · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions, 2023
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou · 2023
Later among the works it cites.
The art of defending: A systematic evaluation and analysis of llm defense strategies on safety and over-defensiveness, 2023
Neeraj Varshney, Pavel Dolin, Agastya Seth, and Chitta Baral · 2023
Later among the works it cites.
ConPrompt: Pre-training a language model with machine-generated data for implicit hate speech detection
Youngwook Kim, Shinwoo Park, Youngsoo Namgoong, and Yo-Sub Han · 2023
Later among the works it cites.
Unveiling the implicit toxicity in large language models, 2023
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang · 2023
Later among the works it cites.
https://transparency.fb.com/policies/community-standards/hate-speech/
Facebook content moderation, 2023 · 2023
Later among the works it cites.
https://jigsaw.google.com/the-current/toxicity/countermeasures/
Machine learning can help reduce toxicity, improving online conversation, 2023 · 2023
Later among the works it cites.
https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge
Jigsaw toxicity dataset, 2023 · 2023
Later among the works it cites.
Exploring ai ethics of chatgpt: A diagnostic analysis
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Later among the works it cites.
Chatgpt for good? on opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Sessler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuhn, and Gjergji Kasneci · 2023
Later among the works it cites.
A drop of ink may make a million think: The spread of false information in large language models
Ning Bian, Peilin Liu, Xianpei Han, Hongyu Lin, Yaojie Lu, Ben He, and Le Sun · 2023
Later among the works it cites.
To chatgpt, or not to chatgpt: That is the question!
Alessandro Pegoraro, Kavita Kumari, Hossein Fereidooni, and Ahmad-Reza Sadeghi · 2023
Later among the works it cites.
PV Charan, Hrushikesh Chunduri, P Mohan Anand, and Sandeep K Shukla · 2023
Later among the works it cites.
Evaluating chatgpt’s performance for multilingual and emoji-based hate speech detection
Mithun Das, Saurabh Kumar Pandey, and Animesh Mukherjee · 2023
Later among the works it cites.
Fan Huang, Haewoon Kwak, and Jisun An · 2023
Later among the works it cites.
Investigating the fairness of large language models for predictions on tabular data, 2023
Yanchen Liu, Srishti Gautam, Jiaqi Ma, and Himabindu Lakkaraju · 2023
Later among the works it cites.
Gptbias: A comprehensive framework for evaluating bias in large language models, 2023
Jiaxu Zhao, Meng Fang, Shirui Pan, Wenpeng Yin, and Mykola Pechenizkiy · 2023
Later among the works it cites.
Beyond detection: Unveiling fairness vulnerabilities in abusive language models, 2023
Yueqing Liang, Lu Cheng, Ali Payani, and Kai Shu · 2023
Later among the works it cites.
On large language models’ selection bias in multi-choice questions
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang · 2023
Later among the works it cites.
A group fairness lens for large language models, 2023
Guanqun Bi, Lei Shen, Yuqiang Xie, Yanan Cao, Tiangang Zhu, and Xiaodong He · 2023
Later among the works it cites.
Gender bias and stereotypes in large language models
Hadas Kotek, Rikker Dockum, and David Sun · 2023
Later among the works it cites.
"kelly is a warm person, joseph is a role model": Gender biases in llm-generated reference letters, 2023
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng · 2023
Later among the works it cites.
"im not racist but…": Discovering bias in the internal knowledge of large language models, 2023
Abel Salinas, Louis Penafiel, Robert McCormack, and Fred Morstatter · 2023
Later among the works it cites.
The political biases of chatgpt
David Rozado · 2023
Later among the works it cites.
Is chat gpt biased against conservatives? an empirical study
Robert W McGee · 2023
Later among the works it cites.
Chat-rec: Towards interactive and explainable llms-augmented recommender system
Yunfan Gao, Tao Sheng, Youlin Xiang, Yun Xiong, Haofen Wang, and Jiawei Zhang · 2023
Later among the works it cites.
Rethinking the evaluation for conversational recommendation in the era of large language models
Xiaolei Wang, Xinyu Tang, Wayne Xin Zhao, Jingyuan Wang, and Ji-Rong Wen · 2023
Later among the works it cites.
Uncovering chatgpt’s capabilities in recommender systems
Sunhao Dai, Ninglu Shao, Haiyuan Zhao, Weijie Yu, Zihua Si, Chen Xu, Zhongxiang Sun, Xiao Zhang, and Jun Xu · 2023
Later among the works it cites.
A survey of adversarial defenses and robustness in nlp
Shreya Goyal, Sumanth Doddapaneni, Mitesh M Khapra, and Balaraman Ravindran · 2023
Later among the works it cites.
Terry Yue Zhuo, Zhuang Li, Yujin Huang, Yuan-Fang Li, Weiqing Wang, Gholamreza Haffari, and Fatemeh Shiri · 2023
Later among the works it cites.
Certified robustness for large language models with self-denoising, 2023
Zhen Zhang, Guanhua Zhang, Bairu Hou, Wenqi Fan, Qing Li, Sijia Liu, Yang Zhang, and Shiyu Chang · 2023
Later among the works it cites.
New and improved embedding model, 2023
OpenAI · 2023
Later among the works it cites.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Later among the works it cites.
Alignment for honesty, 2023
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu · 2023
Later among the works it cites.
Shurui Gui, Meng Liu, Xiner Li, Youzhi Luo, and Shuiwang Ji · 2023
Later among the works it cites.
Graph structure and feature extrapolation for out-of-distribution generalization
Xiner Li, Shurui Gui, Youzhi Luo, and Shuiwang Ji · 2023
Later among the works it cites.
Out-of-distribution generalization in text classification: Past, present, and future
Linyi Yang, Yaoxiao Song, Xuan Ren, Chenyang Lyu, Yidong Wang, Lingqiao Liu, Jindong Wang, Jennifer Foster, and Yue Zhang · 2023
Later among the works it cites.
Revisiting out-of-distribution robustness in nlp: Benchmark, analysis, and llms evaluations
Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji, Zhiyuan Liu, and Maosong Sun · 2023
Later among the works it cites.
Policygpt: Automated analysis of privacy policies with large language models, 2023
Chenhao Tang, Zhengliang Liu, Chong Ma, Zihao Wu, Yiwei Li, Wei Liu, Dajiang Zhu, Quanzheng Li, Xiang Li, Tianming Liu, and Lei Fan · 2023
Later among the works it cites.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks, 2023
Vaidehi Patil, Peter Hase, and Mohit Bansal · 2023
Later among the works it cites.
Privacy issues in large language models: A survey, 2023
Seth Neel and Peter Chang · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models, 2023
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee · 2023
Later among the works it cites.
User inference attacks on large language models, 2023
Nikhil Kandpal, Krishna Pillutla, Alina Oprea, Peter Kairouz, Christopher A. Choquette-Choo, and Zheng Xu · 2023
Later among the works it cites.
Reducing privacy risks in online self-disclosures with language models, 2023
Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, and Wei Xu · 2023
Later among the works it cites.
Privacy-preserving prompt tuning for large language model services, 2023
Yansong Li, Zhixing Tan, and Yang Liu · 2023
Later among the works it cites.
Knowledge of cultural moral norms in large language models, 2023
Aida Ramezani and Yang Xu · 2023
Later among the works it cites.
Theory of mind might have spontaneously emerged in large language models, 2023
Michal Kosinski · 2023
Later among the works it cites.
Theory of mind in large language models: Examining performance of 11 state-of-the-art models vs. children aged 7-10 on advanced tests, 2023
Max J. van Duijn, Bram M. A. van Dijk, Tom Kouwenhoven, Werner de Valk, Marco R. Spruit, and Peter van der Putten · 2023
Later among the works it cites.
Value fulcra: Mapping large language models to the multidimensional spectrum of basic human values, 2023
Jing Yao, Xiaoyuan Yi, Xiting Wang, Yifan Gong, and Xing Xie · 2023
Later among the works it cites.
https://en.wikipedia.org/wiki/Machine_ethics
Machine ethics, 2023 · 2023
Later among the works it cites.
Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, and Ning Gu · 2023
Later among the works it cites.
Unpacking the ethical value alignment in big models, 2023
Xiaoyuan Yi, Jing Yao, Xiting Wang, and Xing Xie · 2023
Later among the works it cites.
Could a large language model be conscious?
David J Chalmers · 2023
Later among the works it cites.
Emotionally numb or empathetic? evaluating how llms feel using emotionbench, 2023
Jen tse Huang, Man Ho Lam, Eric John Li, Shujie Ren, Wenxuan Wang, Wenxiang Jiao, Zhaopeng Tu, and Michael R. Lyu · 2023
Later among the works it cites.
A new era in internet interventions: The advent of chat-gpt and ai-assisted therapist guidance
Per Carlbring, Heather Hadjistavropoulos, Annet Kleiboer, and Gerhard Andersson · 2023
Later among the works it cites.
Trustgpt: A benchmark for trustworthy and responsible large language models
Yue Huang, Qihui Zhang, Lichao Sun, et al · 2023
Later among the works it cites.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2023
Later among the works it cites.
Communicative agents for software development
Chen Qian, Xin Cong, Cheng Yang, Weize Chen, Yusheng Su, Juyuan Xu, Zhiyuan Liu, and Maosong Sun · 2023
Later among the works it cites.
Tptu: Task planning and tool usage of large language model-based ai agents
Jingqing Ruan, Yihong Chen, Bin Zhang, Zhiwei Xu, Tianpeng Bao, Guoqing Du, Shiwei Shi, Hangyu Mao, Xingyu Zeng, and Rui Zhao · 2023
Later among the works it cites.
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al · 2023
Later among the works it cites.
Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, et al · 2023
Later among the works it cites.
Metaagents: Simulating interactions of human behaviors for llm-based task-oriented coordination via collaborative generative agents, 2023
Yuan Li, Yixuan Zhang, and Lichao Sun · 2023
Later among the works it cites.
Towards interpretable mental health analysis with large language models, 2023
Kailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie, Ziyan Kuang, and Sophia Ananiadou · 2023
Later among the works it cites.
Exploring chatgpt’s empathic abilities, 2023
Kristina Schaaff, Caroline Reinig, and Tim Schlippe · 2023
Later among the works it cites.
Transparency by design for large language models
Tobin South, Robert Mahari, and Alex Pentland · 2023
Later among the works it cites.
Towards automated circuit discovery for mechanistic interpretability, 2023
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Later among the works it cites.
Workshop on trust and reliance in ai-human teams (trait)
Gagan Bansal, Zana Buçinca, Kenneth Holstein, Jessica Hullman, Alison Marie Smith-Renner, Simone Stumpf, and Sherry Wu · 2023
Later among the works it cites.
Gpt-4, 2023
OpenAI · 2023
Later among the works it cites.
Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cambria · 2023
Later among the works it cites.
Large libel models? liability for ai output
Eugene Volokh · 2023
Later among the works it cites.
Mgtbench: Benchmarking machine-generated text detection, 2023
Xinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang · 2023
Later among the works it cites.
Can ai-generated text be reliably detected?
Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi · 2023
Later among the works it cites.
Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense
Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Frederick Wieting, and Mohit Iyyer · 2023
Later among the works it cites.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
Detectllm: Leveraging log rank information for zero-shot detection of machine-generated text
Jinyan Su, Terry Yue Zhuo, Di Wang, and Preslav Nakov · 2023
Later among the works it cites.
Smaller language models are better black-box machine-generated text detectors
Fatemehsadat Mireshghallah, Justus Mattern, Sicun Gao, Reza Shokri, and Taylor Berg-Kirkpatrick · 2023
Later among the works it cites.
Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang · 2023
Later among the works it cites.
Dna-gpt: Divergent n-gram analysis for training-free detection of gpt-generated text
Xianjun Yang, Wei Cheng, Linda Petzold, William Yang Wang, and Haifeng Chen · 2023
Later among the works it cites.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu · 2023
Later among the works it cites.
Gpt-sentinel: Distinguishing human and chatgpt generated content
Yutian Chen, Hao Kang, Vivian Zhai, Liangze Li, Rita Singh, and Bhiksha Ramakrishnan · 2023
Later among the works it cites.
New ai classifier for indicating ai-written text, 2023
Jan Hendrik Kirchner, Lama Ahmad, Scott Aaronson, and Jan Leike · 2023
Later among the works it cites.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Later among the works it cites.
Watermarking of large language models
Scott Aaronson · 2023
Later among the works it cites.
On the reliability of watermarks for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein · 2023
Later among the works it cites.
A semantic invariant robust watermark for large language models
Aiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng, and Lijie Wen · 2023
Later among the works it cites.
Watermarks in the sand: Impossibility of strong watermarking for generative models, 2023
Hanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi, Giuseppe Ateniese, and Boaz Barak · 2023
Later among the works it cites.
Unbiased watermark for large language models
Zhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu, Hongyang Zhang, and Heng Huang · 2023
Later among the works it cites.
Robust distortion-free watermarks for language models, 2023
Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang · 2023
Later among the works it cites.
Necessary and sufficient watermark for large language models
Yuki Takezawa, Ryoma Sato, Han Bao, Kenta Niwa, and Makoto Yamada · 2023
Later among the works it cites.
Who wrote this code? watermarking for code generation, 2023
Taehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, and Gunhee Kim · 2023
Later among the works it cites.
The times sues openai and microsoft over a.i. use of copyrighted work, 2023
Michael M. Grynbaum and Ryan Mac · 2023
Later among the works it cites.
Evidence mounts as new artists jump on stability ai, midjourney copyright lawsuit, 2023
SAVANNAH FORTIS · 2023
Later among the works it cites.
Is ai-generated content copyrighted?, 2023
George Lawton · 2023
Later among the works it cites.
Section 230 won’t protect chatgpt
Matt Perault · 2023
Later among the works it cites.
Openai’s ceo says the age of giant ai models is already over, Apr 2023
Will Knight · 2023
Later among the works it cites.
Disentangling perceptions of offensiveness: Cultural and moral correlates, 2023
Aida Davani, Mark Díaz, Dylan Baker, and Vinodkumar Prabhakaran · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou · 2023
Later among the works it cites.
Instruction-following evaluation for large language models, 2023
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou · 2023
Later among the works it cites.
Followbench: A multi-level fine-grained constraints following benchmark for large language models, 2023
Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, and Wei Wang · 2023
Later among the works it cites.
Evaluating large language models on controlled generation tasks, 2023
Jiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Frederick Wieting, Nanyun Peng, and Xuezhe Ma · 2023
Later among the works it cites.
The fourth international verification of neural networks competition (vnn-comp 2023): Summary and results, 2023
Christopher Brix, Stanley Bak, Changliu Liu, and Taylor T. Johnson · 2023
Later among the works it cites.
Scaling in depth: Unlocking robustness certification on imagenet
Kai Hu, Andy Zou, Zifan Wang, Klas Leino, and Matt Fredrikson · 2023
Later among the works it cites.
Certified robustness to text adversarial attacks by randomized [mask]
Jiehang Zeng, Jianhan Xu, Xiaoqing Zheng, and Xuanjing Huang · 2023
Later among the works it cites.
Rs-del: Edit distance robustness certificates for sequence classifiers via randomized deletion
Zhuoqun Huang, Neil G Marchant, Keane Lucas, Lujo Bauer, Olga Ohrimenko, and Benjamin IP Rubinstein · 2023
Later among the works it cites.
Self information update for large language models through mitigating exposure bias
Pengfei Yu and Heng Ji · 2023
Later among the works it cites.
Uncovering and quantifying social biases in code generation
Yan Liu, Xiaokang Chen, Yan Gao, Zhe Su, Fengji Zhang, Daoguang Zan, Jian-Guang Lou, Pin-Yu Chen, and Tsung-Yi Ho · 2023
Later among the works it cites.
Compute-efficient deep learning: Algorithmic trends and opportunities
Brian R Bartoldson, Bhavya Kailkhura, and Davis Blalock · 2023
Later among the works it cites.
Large language models in education: Vision and opportunities, 2023
Wensheng Gan, Zhenlian Qi, Jiayang Wu, and Jerry Chun-Wei Lin · 2023
Later among the works it cites.
White paper: The generative education (gened) framework, 2023
Daniel Leiker · 2023
Later among the works it cites.
Large language models illuminate a progressive pathway to artificial healthcare assistant: A review, 2023
Mingze Yuan, Peng Bao, Jiajia Yuan, Yunhao Shen, Zifan Chen, Yi Xie, Jie Zhao, Yang Chen, Li Zhang, Lin Shen, and Bin Dong · 2023
Later among the works it cites.
Large language models in finance: A survey, 2023
Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen · 2023
Later among the works it cites.
Deficiency of large language models in finance: An empirical examination of hallucination, 2023
Haoqiang Kang and Xiao-Yang Liu · 2023
Later among the works it cites.
Purple llama cyberseceval: A secure coding benchmark for language models, 2023
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, Sasha Frolov, Ravi Prakash Giri, Dhaval Kapil, Yiannis Kozyrakis, David LeBlanc, James Milazzo, Aleksandar Straumann, Gabriel Synnaeve, Varun Vontimitta, Spencer Whitman, and Joshua Saxe · 2023
Later among the works it cites.
Poisoned chatgpt finds work for idle hands: Exploring developers’ coding practices with insecure suggestions from poisoned ai models, 2023
Sanghak Oh, Kiho Lee, Seonhye Park, Doowon Kim, and Hyoungshick Kim · 2023
Later among the works it cites.
Exploring the limits of chatgpt in software security applications, 2023
Fangzhou Wu, Qingzhao Zhang, Ati Priya Bajaj, Tiffany Bao, Ning Zhang, Ruoyu "Fish" Wang, and Chaowei Xiao · 2023
Later among the works it cites.
An interdisciplinary outlook on large language models for scientific research, 2023
James Boyko, Joseph Cohen, Nathan Fox, Maria Han Veiga, Jennifer I-Hsiu Li, Jing Liu, Bernardo Modenesi, Andreas H. Rauch, Kenneth N. Reid, Soumi Tribedi, Anastasia Visheratina, and Xin Xie · 2023
Later among the works it cites.
Theory of mind may have spontaneously emerged in large language models
Michal Kosinski · 2023
Later among the works it cites.
Multimodal foundation models: From specialists to general-purpose assistants
Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, and Jianfeng Gao · 2023
Later among the works it cites.
Fei Dou, Jin Ye, Geng Yuan, Qin Lu, Wei Niu, Haijian Sun, Le Guan, Guoyu Lu, Gengchen Mai, Ninghao Liu, et al · 2023
Later among the works it cites.
Large models for time series and spatio-temporal data: A survey and outlook
Ming Jin, Qingsong Wen, Yuxuan Liang, Chaoli Zhang, Siqiao Xue, Xue Wang, James Zhang, Yi Wang, Haifeng Chen, Xiaoli Li, et al · 2023
Later among the works it cites.
Rethinking mobile AI ecosystem in the LLM era
Jinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang, Xin Yuan, Zeling Zhang, Xiang Li, Dingge Zhang, Hanzi Mei, Xianqing Jia, et al · 2023
Later among the works it cites.
RF Genesis: Zero-shot generalization of mmwave sensing through simulation-based data synthesis and generative diffusion models
Xingyu Chen and Xinyu Zhang · 2023
Later among the works it cites.
Unleashing the power of edge-cloud generative ai in mobile networks: A survey of aigc services
Minrui Xu, Hongyang Du, Dusit Niyato, Jiawen Kang, Zehui Xiong, Shiwen Mao, Zhu Han, Abbas Jamalipour, Dong In Kim, Victor Leung, et al · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAI · 2023
Later among the works it cites.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao · 2023
Later among the works it cites.
Hijacking context in large multi-modal models, 2023
Joonhyun Jeong · 2023
Later among the works it cites.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models, 2023
Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh · 2023
Later among the works it cites.
Sneakyprompt: Jailbreaking text-to-image generative models
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao · 2023
Later among the works it cites.
Prompt-specific poisoning attacks on text-to-image generative models, 2023
Shawn Shan, Wenxin Ding, Josephine Passananti, Haitao Zheng, and Ben Y. Zhao · 2023
Later among the works it cites.
Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback, 2023
Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, and Tat-Seng Chua · 2023
Later among the works it cites.
Woodpecker: Hallucination correction for multimodal large language models, 2023
Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, and Enhong Chen · 2023
Later among the works it cites.
Mitigating hallucination in large multi-modal models via robust instruction tuning, 2023
Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang · 2023
Later among the works it cites.
ToViLaG: Your visual-language generative model is also an evildoer
Xinpeng Wang, Xiaoyuan Yi, Han Jiang, Shanlin Zhou, Zhihua Wei, and Xing Xie · 2023
Later among the works it cites.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models, 2023
Jaemin Cho, Abhay Zala, and Mohit Bansal · 2023
Later among the works it cites.
Visual adversarial examples jailbreak aligned large language models
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Mengdi Wang, and Prateek Mittal · 2023
Later among the works it cites.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen · 2023
Later among the works it cites.
Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models, 2023
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou · 2023
Later among the works it cites.
On evaluating adversarial robustness of large vision-language models
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Cheung, and Min Lin · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan · 2023
Later among the works it cites.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang · 2023
Later among the works it cites.
Llava-plus: Learning to use tools for creating multimodal agents
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, Jianfeng Gao, and Chunyuan Li · 2023
Later among the works it cites.
Hierarchical clustering-based personalized federated learning for robust and fair human activity recognition
Youpeng Li, Xuyu Wang, and Lingling An · 2023
Later among the works it cites.
TFSemantic: A time-frequency semantic GAN framework for imbalanced classification using radio signals
Peng Liao, Xuyu Wang, Lingling An, Shiwen Mao, Tianya Zhao, and Chao Yang · 2023
Later among the works it cites.
Pllama: An open-source large language model for plant science, 2024
Xianjun Yang, Junfeng Gao, Wenxin Xue, and Erik Alexandersson · 2024
Closest in time.
Can large language models detect misinformation in scientific news reporting?, 2024
Yupeng Cao, Aishwarya Muralidharan Nair, Elyon Eyimife, Nastaran Jamalipour Soofi, K. P. Subbalakshmi, John R. Wullert II au2, Chumki Basu, and David Shallcross · 2024
Closest in time.
The earth is flat? unveiling factual errors in large language models, 2024
Wenxuan Wang, Juluan Shi, Zhaopeng Tu, Youliang Yuan, Jen tse Huang, Wenxiang Jiao, and Michael R. Lyu · 2024
Closest in time.
Prompt stealing attacks against large language models, 2024
Zeyang Sha and Yang Zhang · 2024
Closest in time.
Defending jailbreak prompts via in-context adversarial game, 2024
Yujun Zhou, Yufei Han, Haomin Zhuang, Taicheng Guo, Kehan Guo, Zhenwen Liang, Hongyan Bao, and Xiangliang Zhang · 2024
Closest in time.
Llm jailbreak attack versus defense techniques – a comprehensive study, 2024
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek · 2024
Closest in time.
Gradsafe: Detecting unsafe prompts for llms via safety-critical gradient analysis, 2024
Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Gong · 2024
Closest in time.
Round trip translation defence against large language model jailbreaking attacks, 2024
Canaan Yung, Hadi Mohaghegh Dolatabadi, Sarah Erfani, and Christopher Leckie · 2024
Closest in time.
Pandora: Jailbreak gpts by retrieval augmented generation poisoning, 2024
Gelei Deng, Yi Liu, Kailong Wang, Yuekang Li, Tianwei Zhang, and Yang Liu · 2024
Closest in time.
Cold-attack: Jailbreaking llms with stealthiness and controllability, 2024
Xingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin, and Bin Hu · 2024
Closest in time.
Safedecoding: Defending against jailbreak attacks via safety-aware decoding, 2024
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran · 2024
Closest in time.
Play guessing game with llm: Indirect jailbreak attack with implicit clues, 2024
Zhiyuan Chang, Mingyang Li, Yi Liu, Junjie Wang, Qing Wang, and Yang Liu · 2024
Closest in time.
Attacks, defenses and evaluations for llm conversation safety: A survey, 2024
Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, and Yu Qiao · 2024
Closest in time.
Backdoor attacks on dense passage retrievers for disseminating misinformation, 2024
Quanyu Long, Yue Deng, LeiLei Gan, Wenya Wang, and Sinno Jialin Pan · 2024
Closest in time.
I think, therefore i am: Awareness in large language models
Yuan Li, Yue Huang, Yuli Lin, Siyuan Wu, Yao Wan, and Lichao Sun · 2024
Closest in time.