Fetching the paper…
Reading the bibliography…
AI Safety is an emerging area of critical importance to the safe adoption and deployment of AI systems.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Martín Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019 · 1907
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Asynchronous Methods for Deep Reinforcement Learning. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016 (JMLR Workshop and Conference Proceedings, Vol. 48) , Maria-Florina Balcan and Kilian Q. Weinberger (Eds.). 1928–1937
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Towards understanding and enhancing robustness of deep learning models against malicious unlearning attacks. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 1932–1942
Wei Qian, Chenxu Zhao, Wei Le, Meiyi Ma, and Mengdi Huai. 2023 · 1942
Earlier work this paper cites.
Speculations Concerning the First Ultraintelligent Machine
Irving John Good. 1965 · 1965
Earlier work this paper cites.
A Note on the Confinement Problem
Butler W. Lampson. 1973 · 1973
Earlier work this paper cites.
Label-only membership inference attacks. In International conference on machine learning . PMLR, 1964–1974
Christopher A Choquette-Choo, Florian Tramer, Nicholas Carlini, and Nicolas Papernot. 2021 · 1974
Earlier work this paper cites.
The contribution of latent human failures to the breakdown of complex systems
James Reason. 1990 · 1990
Earlier work this paper cites.
Principles of Risk Minimization for Learning Theory. In Advances in Neural Information Processing Systems 4, [NIPS Conference, Denver, Colorado, USA, December 2-5, 1991] , John E. Moody, Stephen Jose Hanson, and Richard Lippmann (Eds.). 831–838
Vladimir Vapnik. 1991 · 1991
Earlier work this paper cites.
The coming technological singularity: How to survive in the post-human era
Vernor Vinge. 1993 · 1993
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Balancing Multiple Sources of Reward in Reinforcement Learning. In Advances in Neural Information Processing Systems 13, Papers from Neural Information Processing Systems (NIPS) 2000, Denver, CO, USA , Todd K. Leen, Thomas G. Dietterich, and Volker Tresp (Eds.). 1082–1088
Christian R. Shelton. 2000 · 2000
Earlier work this paper cites.
Watching, from the Edge of Extinction
B.P. Stearns and S.C. Stearns. 2000 · 2000
Earlier work this paper cites.
Multiagent Systems: A Survey from a Machine Learning Perspective
Peter Stone and Manuela M. Veloso. 2000 · 2000
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Existential risks: Analyzing human extinction scenarios and related hazards
Nick Bostrom. 2002 · 2002
Earlier work this paper cites.
The homograph attack
Evgeniy Gabrilovich and Alex Gontmakher. 2002 · 2002
Earlier work this paper cites.
REALM: Retrieval-Augmented Language Model Pre-Training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
The AI-Box Experiment
Eliezer S. Yudkowsky. 2002 · 2002
Earlier work this paper cites.
Causality: models, reasoning, and inference, by judea pearl, cambridge university press, 2000
Leland Gerson Neuberg. 2003 · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning. In Machine Learning, Proceedings of the Twenty-first International Conference (ICML 2004), Banff, Alberta, Canada, July 4-8, 2004 (ACM International Conference Proceeding Series, Vol. 69) , Carla E. Brodley (Ed.)
Pieter Abbeel and Andrew Y. Ng. 2004 · 2004
Earlier work this paper cites.
SQLrand: Preventing SQL Injection Attacks. In Applied Cryptography and Network Security, Second International Conference, ACNS 2004, Yellow Mountain, China, June 8-11, 2004, Proceedings (Lecture Notes in Computer Science, Vol. 3089) , Markus Jakobsson, Moti Yung, and Jianying Zhou (Eds.). 292–302
Stephen W. Boyd and Angelos D. Keromytis. 2004 · 2004
Earlier work this paper cites.
Extinction: past and present
David Jablonski. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber. 2020 · 2004
Earlier work this paper cites.
The Singularity is Near: When Humans Transcend Biology
Ray Kurzweil. 2005 · 2005
Earlier work this paper cites.
A Deep Learning Approach for Automatic Detection of Fake News
Tanik Saikh, Arkadipta De, Asif Ekbal, and Pushpak Bhattacharyya. 2020 · 2005
Earlier work this paper cites.
AI Research Considerations for Human Existential Safety (ARCHES)
Andrew Critch and David Krueger. 2020 · 2006
Earlier work this paper cites.
A Classification of SQL Injection Attacks and Countermeasures. In 2006 IEEE International Symposium on Secure Software Engineering, ISSSE 2006, Arlington, VA, USA, March 16 -17, 2006 , Samuel T. Redwine Jr. (Ed.)
William G. J. Halfond, Jeremy Viegas, and Alessandro Orso. 2006 · 2006
Earlier work this paper cites.
Eliminating SQL Injection Attacks - A Transparent Defense Mechanism. In Eighth IEEE International Workshop on Web Site Evolution (WSE 2006), 22-24 September 2006, Philadelphia, Pennsylvania, USA . 22–32
Muthusrinivasan Muthuprasanna, Ke Wei, and Suraj Kothari. 2006 · 2006
Earlier work this paper cites.
Common Crawl maintains a free, open repository of web crawl data that can be used by anyone
Common Crawl. 2007 · 2007
Earlier work this paper cites.
Robust de-anonymization of large sparse datasets. In 2008 IEEE Symposium on Security and Privacy (sp 2008) . IEEE, 111–125
Arvind Narayanan and Vitaly Shmatikov. 2008 · 2008
Earlier work this paper cites.
The Basic AI Drives. In Artificial General Intelligence 2008, Proceedings of the First AGI Conference, AGI 2008, March 1-3, 2008, University of Memphis, Memphis, TN, USA (Frontiers in Artificial Intelligence and Applications, Vol. 171) , Pei Wang, Ben Goertzel, and Stan Franklin (Eds.). 483–492
Stephen M. Omohundro. 2008 · 2008
Earlier work this paper cites.
Roman V. Yampolskiy. 2020a · 2008
Earlier work this paper cites.
Hidden Incentives for Auto-Induced Distributional Shift
David Krueger, Tegan Maharaj, and Jan Leike. 2020 · 2009
Earlier work this paper cites.
Utility indifference
Stuart Armstrong. 2010 · 2010
Earlier work this paper cites.
Unmasking Contextual Stereotypes: Measuring and Mitigating BERT’s Gender Bias
Marion Bartl, Malvina Nissim, and Albert Gatt. 2020 · 2010
Earlier work this paper cites.
Measuring and Reducing Gendered Correlations in Pre-trained Models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, and Slav Petrov. 2020a · 2010
Earlier work this paper cites.
Measuring and Reducing Gendered Correlations in Pre-trained Models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, and Slav Petrov. 2020b · 2010
Earlier work this paper cites.
Recipes for Safety in Open-domain Chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2010
Earlier work this paper cites.
The Problem with ‘Friendly’ Artificial Intelligence
Adam Keiper and Ari Schulman. 2011 · 2011
Earlier work this paper cites.
Multimodal Deep Learning. In Proceedings of the 28th International Conference on Machine Learning, ICML 2011, Bellevue, Washington, USA, June 28 - July 2, 2011 , Lise Getoor and Tobias Scheffer (Eds.). 689–696
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng. 2011 · 2011
Earlier work this paper cites.
A Systematic Analysis of XSS Sanitization in Web Application Frameworks. In Computer Security - ESORICS 2011 - 16th European Symposium on Research in Computer Security, Leuven, Belgium, September 12-14, 2011. Proceedings (Lecture Notes in Computer Science, Vol. 6879) , Vijay Atluri and Claudia Díaz (Eds.). 150–171
Joel Weinberger, Prateek Saxena, Devdatta Akhawe, Matthew Finifter, Eui Chul Richard Shin, and Dawn Song. 2011 · 2011
Earlier work this paper cites.
Thinking Inside the Box: Controlling and Using an Oracle AI
Stuart Armstrong, Anders Sandberg, and Nick Bostrom. 2012 · 2012
Earlier work this paper cites.
Open Problems in Cooperative AI
Allan Dafoe, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R. McKee, Joel Z. Leibo, Kate Larson, and Thore Graepel. 2020 · 2012
Earlier work this paper cites.
An overview of 11 proposals for building safe advanced AI
Evan Hubinger. 2020 · 2012
Earlier work this paper cites.
Marcus Hutter. 2012 · 2012
Earlier work this paper cites.
Shadow attacks: automatically evading system-call-behavior based malware detection
Weiqin Ma, Pu Duan, Sanmin Liu, Guofei Gu, and Jyh-Charn Liu. 2012 · 2012
Earlier work this paper cites.
Leakproofing the singularity artificial intelligence confinement problem
Roman V Yampolskiy. 2012 · 2012
Earlier work this paper cites.
Unique in the crowd: The privacy bounds of human mobility
Yves-Alexandre De Montjoye, César A Hidalgo, Michel Verleysen, and Vincent D Blondel. 2013 · 2013
Earlier work this paper cites.
Hallucination: Philosophy and psychology
Fiona Macpherson and Dimitris Platchias. 2013 · 2013
Earlier work this paper cites.
Intelligence explosion: Evidence and import
Luke Muehlhauser and Anna Salamon. 2013 · 2013
Earlier work this paper cites.
Ideal Advisor Theories and Personal CEV
Luke Muehlhauser and Chris Williamson. 2013 · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom. 2014 · 2014
Earlier work this paper cites.
Data privacy law: an international perspective
Lee Andrew Bygrave. 2014 · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL , Alessandro Moschitti, Bo Pang, and Walter Daelemans (Eds.). 1724–1734
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Privacy in pharmacogenetics: An { \{ End-to-End } \} case study of personalized warfarin dosing. In 23rd USENIX security symposium (USENIX Security 14) . 17–32
Matthew Fredrikson, Eric Lantz, Somesh Jha, Simon Lin, David Page, and Thomas Ristenpart. 2014 · 2014
Earlier work this paper cites.
Artificial General Intelligence: Concept, State of the Art, and Future Prospects
Ben Goertzel. 2014 · 2014
Earlier work this paper cites.
Generative Adversarial Nets. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada , Zoubin Ghahramani, Max Welling, Corinna Cortes, Neil D. Lawrence, and Kilian Q. Weinberger (Eds.). 2672–2680
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Current state of research on cross-site scripting (XSS) - A systematic literature review
Isatou Hydara, Abu Bakar Md Sultan, Hazura Zulzalil, and Novia Admodisastro. 2015 · 2014
Earlier work this paper cites.
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Glove: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL , Alessandro Moschitti, Bo Pang, and Walter Daelemans (Eds.). 1532–1543
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks. In Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada , Zoubin Ghahramani, Max Welling, Corinna Cortes, Neil D. Lawrence, and Kilian Q. Weinberger (Eds.). 3104–3112
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. 2014a · 2014
Earlier work this paper cites.
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. 2014b · 2014
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandecic and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Motivated Value Selection for Artificial Agents. In Artificial Intelligence and Ethics, Papers from the 2015 AAAI Workshop, Austin, Texas, USA, January 25, 2015 (AAAI Technical Report, Vol. WS-15-02) , Toby Walsh (Ed.)
Stuart Armstrong. 2015 · 2015
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures. In Proceedings of the 22nd ACM SIGSAC conference on computer and communications security . 1322–1333
Matt Fredrikson, Somesh Jha, and Thomas Ristenpart. 2015 · 2015
Earlier work this paper cites.
Artificial General Intelligence
Ben Goertzel. 2015 · 2015
Earlier work this paper cites.
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Corrigibility. In Artificial Intelligence and Ethics, Papers from the 2015 AAAI Workshop, Austin, Texas, USA, January 25, 2015 (AAAI Technical Report, Vol. WS-15-02) , Toby Walsh (Ed.)
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky. 2015 · 2015
Earlier work this paper cites.
Analysis of Types of Self-Improving Software. In Artificial General Intelligence - 8th International Conference, AGI 2015, AGI 2015, Berlin, Germany, July 22-25, 2015, Proceedings (Lecture Notes in Computer Science, Vol. 9205) , Jordi Bieger, Ben Goertzel, and Alexey Potapov (Eds.). 384–393
Roman V. Yampolskiy. 2015a · 2015
Earlier work this paper cites.
From Seed AI to Technological Singularity via Recursively Self-Improving Software
Roman V. Yampolskiy. 2015b · 2015
Earlier work this paper cites.
On the Limits of Recursively Self-Improving AGI. In Artificial General Intelligence - 8th International Conference, AGI 2015, AGI 2015, Berlin, Germany, July 22-25, 2015, Proceedings (Lecture Notes in Computer Science, Vol. 9205) , Jordi Bieger, Ben Goertzel, and Alexey Potapov (Eds.). 394–403
Roman V. Yampolskiy. 2015c · 2015
Earlier work this paper cites.
Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security . 308–318
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016 · 2016
Earlier work this paper cites.
Concrete Problems in AI Safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané. 2016 · 2016
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
The AGI Containment Problem. In Artificial General Intelligence - 9th International Conference, AGI 2016, New York, NY, USA, July 16-19, 2016, Proceedings (Lecture Notes in Computer Science, Vol. 9782) , Bas R. Steunebrink, Pei Wang, and Ben Goertzel (Eds.). 53–63
James Babcock, János Kramár, and Roman Yampolskiy. 2016 · 2016
Earlier work this paper cites.
Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain , Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett (Eds.). 4349–4357
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai. 2016 · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
The singularity: A philosophical analysis
David J Chalmers. 2016 · 2016
Earlier work this paper cites.
Multi-objective optimization
Kalyanmoy Deb, Karthik Sindhya, and Jussi Hakanen. 2016 · 2016
Earlier work this paper cites.
Self-modification of policy and utility function in rational agents. In Artificial General Intelligence: 9th International Conference, AGI 2016, New York, NY, USA, July 16-19, 2016, Proceedings 9 . Springer, 1–11
Tom Everitt, Daniel Filan, Mayank Daswani, and Marcus Hutter. 2016 · 2016
Earlier work this paper cites.
Cooperative Inverse Reinforcement Learning. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain , Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett (Eds.). 3909–3917
Dylan Hadfield-Menell, Stuart Russell, Pieter Abbeel, and Anca D. Dragan. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora necessarily contain human biases
Aylin Caliskan Islam, Joanna J. Bryson, and Arvind Narayanan. 2016 · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2016 · 2016
Earlier work this paper cites.
Visualizing and Understanding Neural Models in NLP. In NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016 , Kevin Knight, Ani Nenkova, and Owen Rambow (Eds.). 681–691
Jiwei Li, Xinlei Chen, Eduard H. Hovy, and Dan Jurafsky. 2016a · 2016
Earlier work this paper cites.
Understanding Neural Networks through Representation Erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016b · 2016
Earlier work this paper cites.
Death and Suicide in Universal Artificial Intelligence. In Artificial General Intelligence - 9th International Conference, AGI 2016, New York, NY, USA, July 16-19, 2016, Proceedings (Lecture Notes in Computer Science, Vol. 9782) , Bas R. Steunebrink, Pei Wang, and Ben Goertzel (Eds.). 23–32
Jarryd Martin, Tom Everitt, and Marcus Hutter. 2016 · 2016
Earlier work this paper cites.
Explaining nonlinear classification decisions with deep Taylor decomposition
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. 2017 · 2016
Earlier work this paper cites.
Anh Mai Nguyen, Jason Yosinski, and Jeff Clune. 2016 · 2016
Earlier work this paper cites.
Safely Interruptible Agents. In Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence, UAI 2016, June 25-29, 2016, New York City, NY, USA , Alexander Ihler and Dominik Janzing (Eds.)
Laurent Orseau and Stuart Armstrong. 2016 · 2016
Earlier work this paper cites.
Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples
Nicolas Papernot, Patrick D. McDaniel, and Ian J. Goodfellow. 2016 · 2016
Earlier work this paper cites.
"Why Should I Trust You?": Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 13-17, 2016 , Balaji Krishnapuram, Mohak Shah, Alexander J. Smola, Charu C. Aggarwal, Dou Shen, and Rajeev Rastogi (Eds.). 1135–1144
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models
Ashwin K. Vijayakumar, Michael Cogswell, Ramprasaath R. Selvaraju, Qing Sun, Stefan Lee, David J. Crandall, and Dhruv Batra. 2016 · 2016
Earlier work this paper cites.
Social media and fake news in the 2016 election
Hunt Allcott and Matthew Gentzkow. 2017 · 2017
Earlier work this paper cites.
Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks. In Proceedings of the Eighth International Joint Conference on Natural Language Processing, IJCNLP 2017, Taipei, Taiwan, November 27 - December 1, 2017 - Volume 1: Long Papers , Greg Kondrak and Taro Watanabe (Eds.). 1–10
Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James R. Glass. 2017 · 2017
Earlier work this paper cites.
Evasion Attacks against Machine Learning at Test Time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. 2017 · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA , Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.). 4299–4307
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Summoning the Demon: Why superintelligence is humanity’s biggest threat
Richard Clarke and R.P. Eddy. 2017 · 2017
Earlier work this paper cites.
Deep Biaffine Attention for Neural Dependency Parsing. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings
Timothy Dozat and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
Beam Search Strategies for Neural Machine Translation. In Proceedings of the First Workshop on Neural Machine Translation, NMT@ACL 2017, Vancouver, Canada, August 4, 2017 , Thang Luong, Alexandra Birch, Graham Neubig, and Andrew M. Finch (Eds.). 56–60
Markus Freitag and Yaser Al-Onaizan. 2017 · 2017
Earlier work this paper cites.
Anthropogenic processes, natural hazards, and interactions in a multi-hazard framework
Joel C. Gill and Bruce D. Malamud. 2017 · 2017
Earlier work this paper cites.
Cross-Site Scripting (XSS) attacks and defense mechanisms: classification and state-of-the-art
Shashank Gupta and Brij Bhooshan Gupta. 2017 · 2017
Earlier work this paper cites.
Annotation Artifacts in Natural Language Inference Data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 2 (Short Papers) , Marilyn A. Walker, Heng Ji, and Amanda Stent (Eds.). 107–112
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018 · 2017
Earlier work this paper cites.
The Off-Switch Game. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017, Melbourne, Australia, August 19-25, 2017 , Carles Sierra (Ed.). 220–227
Dylan Hadfield-Menell, Anca D. Dragan, Pieter Abbeel, and Stuart Russell. 2017 · 2017
Earlier work this paper cites.
Deceiving Google’s Perspective API Built for Detecting Toxic Comments
Hossein Hosseini, Sreeram Kannan, Baosen Zhang, and Radha Poovendran. 2017 · 2017
Earlier work this paper cites.
Categorical Reparameterization with Gumbel-Softmax. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA , Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.). 4765–4774
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statistics . 1273–1282
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017 · 2017
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2017 · 2017
Earlier work this paper cites.
Curiosity-driven Exploration by Self-supervised Prediction. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 (Proceedings of Machine Learning Research, Vol. 70) , Doina Precup and Yee Whye Teh (Eds.). 2778–2787
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. 2017 · 2017
Earlier work this paper cites.
Towards Crafting Text Adversarial Samples
Suranjana Samanta and Sameep Mehta. 2017 · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 . 618–626
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP) . IEEE, 3–18
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Fake News Detection on Social Media: A Data Mining Perspective
Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu. 2017 · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 (Proceedings of Machine Learning Research, Vol. 70) , Doina Precup and Yee Whye Teh (Eds.). 3319–3328
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA , Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.). 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
A Game-Theoretic Analysis of the Off-Switch Game. In Artificial General Intelligence - 10th International Conference, AGI 2017, Melbourne, VIC, Australia, August 15-18, 2017, Proceedings (Lecture Notes in Computer Science, Vol. 10414) , Tom Everitt, Ben Goertzel, and Alexey Potapov (Eds.). 167–177
Tobias Wängberg, Mikael Böörs, Elliot Catt, Tom Everitt, and Marcus Hutter. 2017 · 2017
Earlier work this paper cites.
CN-DBpedia: A never-ending Chinese knowledge extraction system. In International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems . Springer, 428–438
Bo Xu, Yong Xu, Jiaqing Liang, Chenhao Xie, Bin Liang, Wanyun Cui, and Yanghua Xiao. 2017 · 2017
Earlier work this paper cites.
Generating Natural Language Adversarial Examples. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). 2890–2896
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani B. Srivastava, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Multimodal Machine Learning: A Survey and Taxonomy
Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2019 · 2018
Earlier work this paper cites.
Why I expect successful (narrow) alignment
Tobias Baumann. 2018 · 2018
Earlier work this paper cites.
Distributed Statistical Machine Learning in Adversarial Settings: Byzantine Gradient Descent. In Abstracts of the 2018 ACM International Conference on Measurement and Modeling of Computer Systems, SIGMETRICS 2018, Irvine, CA, USA, June 18-22, 2018 , Konstantinos Psounis, Aditya Akella, and Adam Wierman (Eds.). 96
Yudong Chen, Lili Su, and Jiaming Xu. 2018 · 2018
Earlier work this paper cites.
OpenAI Wants to Make Safe AI, but That May Be an Impossible Task
Jolene Creighton. 2018 · 2018
Earlier work this paper cites.
AI governance: a research agenda
Allan Dafoe. 2018 · 2018
Earlier work this paper cites.
Essentially No Barriers in Neural Network Energy Landscape. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018 (Proceedings of Machine Learning Research, Vol. 80) , Jennifer G. Dy and Andreas Krause (Eds.). 1308–1317
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A. Hamprecht. 2018 · 2018
Earlier work this paper cites.
The alignment problem for Bayesian history-based reinforcement learners
Tom Everitt and Marcus Hutter. 2018 · 2018
Earlier work this paper cites.
Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers. In 2018 IEEE Security and Privacy Workshops, SP Workshops 2018, San Francisco, CA, USA, May 24, 2018 . 50–56
Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018 · 2018
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018 · 2018
Earlier work this paper cites.
Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada , Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicolò Cesa-Bianchi, and Roman Garnett (Eds.). 8803–8812
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P. Vetrov, and Andrew Gordon Wilson. 2018 · 2018
Earlier work this paper cites.
PipeDream: Fast and Efficient Pipeline Parallel DNN Training
Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil R. Devanur, Gregory R. Ganger, and Phillip B. Gibbons. 2018 · 2018
Earlier work this paper cites.
Adversarial Example Generation with Syntactically Controlled Paraphrase Networks. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers) , Marilyn A. Walker, Heng Ji, and Amanda Stent (Eds.). 1875–1885
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV). In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018 (Proceedings of Machine Learning Research, Vol. 80) , Jennifer G. Dy and Andreas Krause (Eds.). 2673–2682
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie J. Cai, James Wexler, Fernanda B. Viégas, and Rory Sayres. 2018 · 2018
Earlier work this paper cites.
Hallucinations in neural machine translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2018 · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018 · 2018
Earlier work this paper cites.
Deep Text Classification Can be Fooled. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden , Jérôme Lang (Ed.). 4208–4215
Bin Liang, Hongcheng Li, Miaoqiang Su, Pan Bian, Xirong Li, and Wenchang Shi. 2018 · 2018
Earlier work this paper cites.
Towards Deep Learning Models Resistant to Adversarial Attacks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018 · 2018
Earlier work this paper cites.
Targeted Syntactic Evaluation of Language Models. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). 1192–1202
Rebecca Marvin and Tal Linzen. 2018 · 2018
Earlier work this paper cites.
Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. 2019 · 2018
Earlier work this paper cites.
Stress Test Evaluation for Natural Language Inference. In Proceedings of the 27th International Conference on Computational Linguistics, COLING 2018, Santa Fe, New Mexico, USA, August 20-26, 2018 , Emily M. Bender, Leon Derczynski, and Pierre Isabelle (Eds.). 2340–2353
Aakanksha Naik, Abhilasha Ravichander, Norman M. Sadeh, Carolyn P. Rosé, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
A Deep Reinforced Model for Abstractive Summarization. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings
Romain Paulus, Caiming Xiong, and Richard Socher. 2018 · 2018
Earlier work this paper cites.
DeClarE: Debunking Fake News and False Claims using Evidence-Aware Deep Learning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). 22–32
Kashyap Popat, Subhabrata Mukherjee, Andrew Yates, and Gerhard Weikum. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Greedy Search with Probabilistic N-gram Matching for Neural Machine Translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). 4778–4784
Chenze Shao, Xilin Chen, and Yang Feng. 2018 · 2018
Earlier work this paper cites.
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018 · 2018
Earlier work this paper cites.
Democratizing ai
Bibb Allen, Sheela Agarwal, Jayashree Kalpathy-Cramer, and Keith Dreyer. 2019 · 2019
Earlier work this paper cites.
Guidelines for artificial intelligence containment
James Babcock, Janos Kramar, and Roman V Yampolskiy. 2019 · 2019
Earlier work this paper cites.
Identifying and Controlling Important Neurons in Neural Machine Translation. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James R. Glass. 2019 · 2019
Earlier work this paper cites.
Exploration by random network distillation. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov. 2019 · 2019
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks. In 28th USENIX security symposium (USENIX security 19) . 267–284
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019 · 2019
Earlier work this paper cites.
Unlabeled Data Improves Adversarial Robustness. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 11190–11201
Yair Carmon, Aditi Raghunathan, Ludwig Schmidt, John C. Duchi, and Percy Liang. 2019 · 2019
Earlier work this paper cites.
Robust Neural Machine Translation with Doubly Adversarial Inputs. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , Anna Korhonen, David R. Traum, and Lluís Màrquez (Eds.). 4324–4333
Yong Cheng, Lu Jiang, and Wolfgang Macherey. 2019 · 2019
Earlier work this paper cites.
Towards XMAS: eXplainable and trustworthy Multi-Agent Systems. In Proceedings of the Proceedings of the 1st Workshop on Artificial Intelligence and Internet of Things co-located with the 18th International Conference of the Italian Association for Artificial Intelligence (AIxIA 2019) . 22 November 2019
Giovanni Ciatto, Roberta Calegari, Andrea Omicini, and Davide Calvaresi. 2019 · 2019
Earlier work this paper cites.
What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019 . 6309–6317
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James R. Glass. 2019 · 2019
Earlier work this paper cites.
Attenuating Bias in Word vectors. In The 22nd International Conference on Artificial Intelligence and Statistics, AISTATS 2019, 16-18 April 2019, Naha, Okinawa, Japan (Proceedings of Machine Learning Research, Vol. 89) , Kamalika Chaudhuri and Masashi Sugiyama (Eds.). 879–887
Sunipa Dev and Jeff M. Phillips. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Addressing Age-Related Bias in Sentiment Analysis. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019 , Sarit Kraus (Ed.). 6146–6150
Mark Díaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. 2019 · 2019
Earlier work this paper cites.
ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. 2019 · 2019
Earlier work this paper cites.
Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 15637–15648
Dan Hendrycks, Mantas Mazeika, Saurav Kadavath, and Dawn Song. 2019 · 2019
Earlier work this paper cites.
A Structural Probe for Finding Syntax in Word Representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). 4129–4138
John Hewitt and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 103–112
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. 2019 · 2019
Earlier work this paper cites.
Attention is not Explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). 3543–3556
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
The global landscape of AI ethics guidelines
Anna Jobin, Marcello Ienca, and Effy Vayena. 2019 · 2019
Earlier work this paper cites.
The (Un)reliability of Saliency Methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2019 · 2019
Earlier work this paper cites.
TextBugger: Generating Adversarial Text Against Real-world Applications. In 26th Annual Network and Distributed System Security Symposium, NDSS 2019, San Diego, California, USA, February 24-27, 2019
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2019 · 2019
Earlier work this paper cites.
Black is to Criminal as Caucasian is to Police: Detecting and Removing Multiclass Bias in Word Embeddings. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). 615–621
Thomas Manzini, Yao Chong Lim, Alan W. Black, and Yulia Tsvetkov. 2019 · 2019
Earlier work this paper cites.
On Measuring Social Biases in Sentence Encoders. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). 622–628
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
Layer-Wise Relevance Propagation: An Overview
Grégoire Montavon, Alexander Binder, Sebastian Lapuschkin, Wojciech Samek, and Klaus-Robert Müller. 2019 · 2019
Earlier work this paper cites.
Distributed, decentralized, and democratized artificial intelligence
Gabriel Axel Montes and Ben Goertzel. 2019 · 2019
Earlier work this paper cites.
SANVis: Visual Analytics for Understanding Self-Attention Networks. In 30th IEEE Visualization Conference, IEEE VIS 2019 - Short Papers, Vancouver, BC, Canada, October 20-25, 2019 . 146–150
Cheonbok Park, Jaegul Choo, Inyoup Na, Yongjang Jo, Sungbok Shin, Jaehyo Yoo, Bum Chul Kwon, Jian Zhao, Hyungjong Noh, and Yeonsoo Lee. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019 , Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). 2463–2473
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Is Attention Interpretable?. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , Anna Korhonen, David R. Traum, and Lluís Màrquez (Eds.). 2931–2951
Sofia Serrano and Noah A. Smith. 2019 · 2019
Earlier work this paper cites.
A brief review on search engine optimization. In 2019 9th international conference on cloud computing, data science & engineering (confluence) . IEEE, 687–692
Dushyant Sharma, Rishabh Shukla, Anil Kumar Giri, and Sumit Kumar. 2019 · 2019
Earlier work this paper cites.
dEFEND: Explainable Fake News Detection. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019 , Ankur Teredesai, Vipin Kumar, Ying Li, Rómer Rosales, Evimaria Terzi, and George Karypis (Eds.). 395–405
Kai Shu, Limeng Cui, Suhang Wang, Dongwon Lee, and Huan Liu. 2019 · 2019
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al · 2019
Earlier work this paper cites.
Generative Modeling by Estimating Gradients of the Data Distribution. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 11895–11907
Yang Song and Stefano Ermon. 2019 · 2019
Earlier work this paper cites.
Assessing Social and Intersectional Biases in Contextualized Word Representations. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 13209–13220
Yi Chern Tan and L. Elisa Celis. 2019 · 2019
Earlier work this paper cites.
What do you learn from context? Probing for sentence structure in contextualized word representations. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
How Does BERT Answer Questions?: A Layer-Wise Analysis of Transformer Representations. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management, CIKM 2019, Beijing, China, November 3-7, 2019 , Wenwu Zhu, Dacheng Tao, Xueqi Cheng, Peng Cui, Elke A. Rundensteiner, David Carmel, Qi He, and Jeffrey Xu Yu (Eds.). 1823–1832
Betty van Aken, Benjamin Winter, Alexander Löser, and Felix A. Gers. 2019 · 2019
Earlier work this paper cites.
A Multiscale Visualization of Attention in the Transformer Model. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28 - August 2, 2019, Volume 3: System Demonstrations , Marta R. Costa-jussà and Enrique Alfonseca (Eds.). 37–42
Jesse Vig. 2019 · 2019
Earlier work this paper cites.
Trick Me If You Can: Human-in-the-Loop Generation of Adversarial Examples for Question Answering
Eric Wallace, Pedro Rodriguez, Shi Feng, Ikuya Yamada, and Jordan Boyd-Graber. 2019 · 2019
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Defending Against Neural Fake News. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 9051–9062
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
PAWS: Paraphrase Adversaries from Word Scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). 1298–1308
Yuan Zhang, Jason Baldridge, and Luheng He. 2019 · 2019
Earlier work this paper cites.
Deep leakage from gradients
Ligeng Zhu, Zhijian Liu, and Song Han. 2019 · 2019
Earlier work this paper cites.
Model extraction from counterfactual explanations
Ulrich Aïvodji, Alexandre Bolot, and Sébastien Gambs. 2020 · 2020
Earlier work this paper cites.
The Pushshift Reddit Dataset. In Proceedings of the Fourteenth International AAAI Conference on Web and Social Media, ICWSM 2020, Held Virtually, Original Venue: Atlanta, Georgia, USA, June 8-11, 2020 , Munmun De Choudhury, Rumi Chunara, Aron Culotta, and Brooke Foucault Welles (Eds.). 830–839
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020a · 2020
Earlier work this paper cites.
The Pushshift Reddit Dataset. In Proceedings of the Fourteenth International AAAI Conference on Web and Social Media, ICWSM 2020, Held Virtually, Original Venue: Atlanta, Georgia, USA, June 8-11, 2020 , Munmun De Choudhury, Rumi Chunara, Aron Culotta, and Brooke Foucault Welles (Eds.). 830–839
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020b · 2020
Earlier work this paper cites.
Secure single-server aggregation with (poly) logarithmic overhead. In ACM SIGSAC Conference on Computer and Communications Security . 1253–1269
James Henry Bell, Kallista A Bonawitz, Adrià Gascón, Tancrède Lepoint, and Mariana Raykova. 2020 · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020a · 2020
Earlier work this paper cites.
Explainable and Ethical AI: A Perspective on Argumentation and Logic Programming. In AIxIA 2020 - Advances in Artificial Intelligence - XIXth International Conference of the Italian Association for Artificial Intelligence, Virtual Event, November 25-27, 2020, Revised Selected Papers (Lecture Notes in Computer Science, Vol. 12414) , Matteo Baldoni and Stefania Bandini (Eds.). 19–36
Roberta Calegari, Andrea Omicini, and Giovanni Sartor. 2020 · 2020
Earlier work this paper cites.
Toward Gender-Inclusive Coreference Resolution. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 4568–4595
Yang Trista Cao and Hal Daumé III. 2020 · 2020
Earlier work this paper cites.
Backdoor attacks and defenses for deep neural networks in outsourced cloud environments
Yanjiao Chen, Xueluan Gong, Qian Wang, Xing Di, and Huayang Huang. 2020 · 2020
Earlier work this paper cites.
Statistics of Robust Optimization: A Generalized Empirical Likelihood Approach
John C. Duchi, Peter W. Glynn, and Hongseok Namkoong. 2021 · 2020
Earlier work this paper cites.
Local Model Poisoning Attacks to Byzantine-Robust Federated Learning. In 29th USENIX Security Symposium, USENIX Security 2020, August 12-14, 2020 , Srdjan Capkun and Franziska Roesner (Eds.). 1605–1622
Minghong Fang, Xiaoyu Cao, Jinyuan Jia, and Neil Zhenqiang Gong. 2020 · 2020
Earlier work this paper cites.
BAE: BERT-based Adversarial Examples for Text Classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 6174–6181
Siddhant Garg and Goutham Ramakrishnan. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL, Vol. EMNLP 2020) , Trevor Cohn, Yulan He, and Yang Liu (Eds.). 3356–3369
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020a · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL, Vol. EMNLP 2020) , Trevor Cohn, Yulan He, and Yang Liu (Eds.). 3356–3369
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020b · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard S. Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020 · 2020
Earlier work this paper cites.
Model Extraction Attacks and Defenses on Cloud-Based Machine Learning Models
Xueluan Gong, Qian Wang, Yanjiao Chen, Wang Yang, and Xinchang Jiang. 2020 · 2020
Earlier work this paper cites.
Secure weighted aggregation for federated learning
Jiale Guo, Ziyao Liu, Kwok-Yan Lam, Jun Zhao, Yiqiang Chen, and Chaoping Xing. 2020 · 2020
Earlier work this paper cites.
Trustworthy AI - Integrating Learning, Optimization and Reasoning - First International Workshop, TAILOR 2020, Virtual Event, September 4-5, 2020, Revised Selected Papers . Lecture Notes in Computer Science, Vol. 12641. Springer
Fredrik Heintz, Michela Milano, and Barry O’Sullivan (Eds.). 2021 · 2020
Earlier work this paper cites.
Intrinsic Probing through Dimension Selection. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 197–216
Lucas Torroba Hennigen, Adina Williams, and Ryan Cotterell. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
The Curious Case of Neural Text Degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
"Other-Play" for Zero-Shot Coordination. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . 4399–4410
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob N. Foerster. 2020 · 2020
Earlier work this paper cites.
Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020 . 8018–8025
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020a · 2020
Earlier work this paper cites.
Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020 . 8018–8025
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020b · 2020
Earlier work this paper cites.
Artificial intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry
Nils C. Köbis and Luca Mossink. 2021 · 2020
Earlier work this paper cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 7871–7880
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020b · 2020
Earlier work this paper cites.
BERT-ATTACK: Adversarial Attack Against BERT Using BERT. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 6193–6202
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020b · 2020
Earlier work this paper cites.
UNQOVERing Stereotypical Biases via Underspecified Questions. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL, Vol. EMNLP 2020) , Trevor Cohn, Yulan He, and Yang Liu (Eds.). 3475–3489
Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar. 2020a · 2020
Earlier work this paper cites.
MPC-enabled privacy-preserving neural network training against malicious attack
Ziyao Liu, Ivan Tjuawinata, Chaoping Xing, and Kwok-Yan Lam. 2020 · 2020
Earlier work this paper cites.
PowerTransformer: Unsupervised Controllable Revision for Biased Language Correction. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 7426–7441
Xinyao Ma, Maarten Sap, Hannah Rashkin, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
A Survey on Computational Propaganda Detection. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI 2020 , Christian Bessiere (Ed.). 4826–4832
Giovanni Da San Martino, Stefano Cresci, Alberto Barrón-Cedeño, Seunghak Yu, Roberto Di Pietro, and Preslav Nakov. 2020 · 2020
Earlier work this paper cites.
On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 1906–1919
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan T. McDonald. 2020 · 2020
Earlier work this paper cites.
Compositional Explanations of Neurons. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Jesse Mu and Jacob Andreas. 2020 · 2020
Earlier work this paper cites.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). 1953–1967
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 2020
Earlier work this paper cites.
What is being transferred in transfer learning?. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang. 2020 · 2020
Earlier work this paper cites.
Adversarial NLI: A New Benchmark for Natural Language Understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 4885–4901
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020a · 2020
Earlier work this paper cites.
Adversarial NLI: A New Benchmark for Natural Language Understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 4885–4901
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020b · 2020
Earlier work this paper cites.
interpreting GPT: the logit lens
nostalgebraist. 2020 · 2020
Earlier work this paper cites.
ActiveThief: Model Extraction Using Active Learning and Unannotated Public Data. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020 . 865–872
Soham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade, Shirish K. Shevade, and Vinod Ganapathy. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
ZeRO: memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2020, Virtual Event / Atlanta, Georgia, USA, November 9-19, 2020 , Christine Cuicchi, Irene Qualters, and William T. Kramer (Eds.). 20
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 4902–4912
Marco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
Towards probabilistic verification of machine unlearning
David Marco Sommer, Liwei Song, Sameer Wagh, and Prateek Mittal. 2020 · 2020
Earlier work this paper cites.
Improved Techniques for Training Score-Based Generative Models. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Yang Song and Stefano Ermon. 2020 · 2020
Earlier work this paper cites.
Classification of global catastrophic risks connected with artificial intelligence
Alexey Turchin and David Denkenberger. 2020 · 2020
Earlier work this paper cites.
Deepfakes and disinformation: Exploring the impact of synthetic political video on deception, uncertainty, and trust in news
Cristian Vaccari and Andrew Chadwick. 2020 · 2020
Earlier work this paper cites.
On Exposure Bias, Hallucination and Domain Shift in Neural Machine Translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 3544–3552
Chaojun Wang and Rico Sennrich. 2020 · 2020
Earlier work this paper cites.
Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERT. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 4166–4176
Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu. 2020 · 2020
Earlier work this paper cites.
Uncontrollability of AI
Roman V Yampolskiy. 2020b · 2020
Earlier work this paper cites.
US public opinion on the governance of artificial intelligence. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society . 187–193
Baobao Zhang and Allan Dafoe. 2020 · 2020
Earlier work this paper cites.
Attacks Which Do Not Kill Training Make Adversarial Learning Stronger. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . 11278–11287
Jingfeng Zhang, Xilie Xu, Bo Han, Gang Niu, Lizhen Cui, Masashi Sugiyama, and Mohan S. Kankanhalli. 2020 · 2020
Earlier work this paper cites.
Consequences of Misaligned AI. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Simon Zhuang and Dylan Hadfield-Menell. 2020 · 2020
Earlier work this paper cites.
Persistent Anti-Muslim Bias in Large Language Models. In AIES ’21: AAAI/ACM Conference on AI, Ethics, and Society, Virtual Event, USA, May 19-21, 2021 , Marion Fourcade, Benjamin Kuipers, Seth Lazar, and Deirdre K. Mulligan (Eds.). 298–306
Abubakar Abid, Maheen Farooqi, and James Zou. 2021 · 2021
Earlier work this paper cites.
Mitigating Language-Dependent Ethnic Bias in BERT. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). 533–549
Jaimeen Ahn and Alice Oh. 2021 · 2021
Earlier work this paper cites.
Search engine optimization: a review
Firas Almukhtar, Nawzad Mahmoodd, and Shahab Kareem. 2021 · 2021
Earlier work this paper cites.
ALL Dolphins Are Intelligent and SOME Are Friendly: Probing BERT for Nouns’ Semantic Properties and their Prototypicality. In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@EMNLP 2021, Punta Cana, Dominican Republic, November 11, 2021 , Jasmijn Bastings, Yonatan Belinkov, Emmanuel Dupoux, Mario Giulianelli, Dieuwke Hupkes, Yuval Pinter, and Hassan Sajjad (Eds.). 79–94
Marianna Apidianaki and Aina Garí Soler. 2021 · 2021
Earlier work this paper cites.
Do Data Breaches Damage Reputation? Evidence from 45 Cases
Anupam Arora, Rahul Telang, and Hong Xu. 2021 · 2021
Earlier work this paper cites.
Program Synthesis with Large Language Models
Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021 · 2021
Earlier work this paper cites.
Recent Advances in Adversarial Training for Adversarial Robustness. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-27 August 2021 , Zhi-Hua Zhou (Ed.). 4312–4321
Tao Bai, Jinqi Luo, Jun Zhao, Bihan Wen, and Qian Wang. 2021 · 2021
Earlier work this paper cites.
RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). 1941–1955
Soumya Barikeri, Anne Lauscher, Ivan Vulic, and Goran Glavas. 2021 · 2021
Earlier work this paper cites.
Machine unlearning. In IEEE Symposium on Security and Privacy . 141–159
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. 2021 · 2021
Earlier work this paper cites.
AI and the Future of Disinformation Campaigns
CSET Policy Brief. 2021 · 2021
Earlier work this paper cites.
Competition in pricing algorithms
Zach Y Brown and Alexander MacKay. 2021 · 2021
Earlier work this paper cites.
Truth, Lies, and Automation: How Language Models Could Change Disinformation
B Buchanan, A Lohn, M Musser, and K Sedova. 2021 · 2021
Earlier work this paper cites.
Logic-based technologies for multi-agent systems: a systematic literature review
Roberta Calegari, Giovanni Ciatto, Viviana Mascardi, and Andrea Omicini. 2021 · 2021
Earlier work this paper cites.
Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21) . 2633–2650
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
The dark side of AI-powered service interactions: Exploring the process of co-destruction from the customer perspective
Daniela Castillo, Ana Isabel Canhoto, and Emanuel Said. 2021 · 2021
Earlier work this paper cites.
Social network behavior and public opinion manipulation
Long Chen, Jianguo Chen, and Chunhe Xia. 2022a · 2021
Earlier work this paper cites.
When machine unlearning jeopardizes privacy. In Proceedings of the 2021 ACM SIGSAC conference on computer and communications security . 896–911
Min Chen, Zhikun Zhang, Tianhao Wang, Michael Backes, Mathias Humbert, and Yang Zhang. 2021b · 2021
Earlier work this paper cites.
De-Confounded Variational Encoder-Decoder for Logical Table-to-Text Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). 5532–5542
Wenqing Chen, Jidong Tian, Yitian Li, Hao He, and Yaohui Jin. 2021a · 2021
Earlier work this paper cites.
Detecting Hate Speech with GPT-3
Ke-Li Chiu and Rohan Alexander. 2021 · 2021
Earlier work this paper cites.
Cooperative AI: machines must learn to find common ground
Allan Dafoe, Yoram Bachrach, Gillian Hadfield, Eric Horvitz, Kate Larson, and Thore Graepel. 2021 · 2021
Earlier work this paper cites.
Stereotype and Skew: Quantifying Gender Bias in Pre-trained and Fine-tuned Language Models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, April 19 - 23, 2021 , Paola Merlo, Jörg Tiedemann, and Reut Tsarfaty (Eds.). 2232–2242
Daniel de Vassimon Manela, David Errington, Thomas Fisher, Boris van Breugel, and Pasquale Minervini. 2021 · 2021
Earlier work this paper cites.
Stakeholder Participation in AI: Beyond" Add Diverse Stakeholders and Stir"
Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2021 · 2021
Earlier work this paper cites.
Assessing the Reliability of Word Embedding Gender Bias Measures. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). 10012–10034
Yupei Du, Qixiang Fang, and Dong Nguyen. 2021 · 2021
Earlier work this paper cites.
Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). 2197–2214
Nouha Dziri, Andrea Madotto, Osmar Zaïane, and Avishek Joey Bose. 2021 · 2021
Earlier work this paper cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2021 · 2021
Earlier work this paper cites.
He is very intelligent, she is very beautiful? On Mitigating Social Biases in Language Modelling and Generation. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021 (Findings of ACL, Vol. ACL/IJCNLP 2021) , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). 4534–4545
Aparna Garimella, Akhash Amarnath, Kiran Kumar, Akash Pramod Yalla, Anandhavelu Natarajan, Niyati Chhaya, and Balaji Vasan Srinivasan. 2021 · 2021
Earlier work this paper cites.
TUW-Inf at GermEval2021: Rule-based and Hybrid Methods for Detecting Toxic, Engaging, and Fact-Claiming Comments. In Proceedings of the GermEval 2021 Shared Task on the Identification of Toxic, Engaging, and Fact-Claiming Comments, GermEval@KONVENS 2021, Düsseldorf, Germany, September 6, 2021 , Julian Risch, Anke Stoll, Lena Wilms, and Michael Wiegand (Eds.). 69–75
Kinga Gémes and Gábor Recski. 2021 · 2021
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). 5484–5495
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Earlier work this paper cites.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. 2021 · 2021
Earlier work this paper cites.
Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases. In AIES ’21: AAAI/ACM Conference on AI, Ethics, and Society, Virtual Event, USA, May 19-21, 2021 , Marion Fourcade, Benjamin Kuipers, Seth Lazar, and Deirdre K. Mulligan (Eds.). 122–133
Wei Guo and Aylin Caliskan. 2021 · 2021
Earlier work this paper cites.
Self-Attention Attribution: Interpreting Information Interactions Inside Transformer. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021 . 12963–12971
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2021 · 2021
Earlier work this paper cites.
Deberta: decoding-Enhanced Bert with Disentangled Attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
DYPLOC: Dynamic Planning of Content Using Mixed Language Models for Text Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). 6408–6423
Xinyu Hua, Ashwin Sreevatsa, and Lu Wang. 2021 · 2021
Earlier work this paper cites.
Engaged to a robot? The role of AI in service
Ming-Hui Huang and Roland T Rust. 2021 · 2021
Cited alongside, same era.
VisQA: X-raying Vision and Language Reasoning in Transformers
Theo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov, Moez Baccouche, and Christian Wolf. 2022 · 2021
Cited alongside, same era.
Advances and open problems in federated learning
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2021
Cited alongside, same era.
Dynabench: Rethinking Benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021 , Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tür, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (Eds.). 4110–4124
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
A Survey of Text Watermarking in the Era of Large Language Models
Aiwei Liu, Leyi Pan, Yijian Lu, Jingjing Li, Xuming Hu, Lijie Wen, Irwin King, and Philip S. Yu. 2023j · 2023
Later among the works it cites.
Exposing Attention Glitches with Flip-Flop Language Modeling. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023a · 2023
Later among the works it cites.
Visual Instruction Tuning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023h · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Objective Robustness in Deep Reinforcement Learning
Jack Koch, Lauro Langosco, Jacob Pfau, James Le, and Lee Sharkey. 2021 · 2021
Cited alongside, same era.
BERT meets Shapley: Extending SHAP Explanations to Transformer-based Classifiers. In Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation, EACL 2021, Online, April 19, 2021 , Hannu Toivonen and Michele Boggia (Eds.). 16–21
Enja Kokalj, Blaz Skrlj, Nada Lavrac, Senja Pollak, and Marko Robnik-Sikonja. 2021 · 2021
Cited alongside, same era.
Out-of-Distribution Generalization via Risk Extrapolation (REx). In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). 5815–5826
David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Rémi Le Priol, and Aaron C. Courville. 2021 · 2021
Cited alongside, same era.
Probing Across Time: What Does RoBERTa Know and When?. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021 , Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.). 820–842
Zeyu Liu, Yizhong Wang, Jungo Kasai, Hannaneh Hajishirzi, and Noah A. Smith. 2021 · 2021
Cited alongside, same era.
Generation-Augmented Retrieval for Open-Domain Question Answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). 4089–4100
Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Looking Beyond Sentence-Level Natural Language Inference for Question Answering and Text Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021 , Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tür, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (Eds.). 1322–1336
Anshuman Mishra, Dhruvesh Patel, Aparna Vijayakumar, Xiang Lorraine Li, Pavan Kapanipathi, and Kartik Talamadupula. 2021 · 2021
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). 5356–5371
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Cited alongside, same era.
Distributed and democratized learning: Philosophy and research challenges
Minh NH Nguyen, Shashi Raj Pandey, Kyi Thar, Nguyen H Tran, Mingzhe Chen, Walid Saad Bradley, and Choong Seon Hong. 2021 · 2021
Cited alongside, same era.
Hao Liu, Carmelo Sferrazza, and Pieter Abbeel. 2023k · 2023
Later among the works it cites.
Second Thoughts are Best: Learning to Re-Align With Human Values from Text Edits
Ruibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang, Tony X. Liu, and Soroush Vosoughi. 2023f · 2023
Later among the works it cites.
Prompt Injection attack against LLM-integrated Applications
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu. 2023b · 2023
Later among the works it cites.
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023c · 2023
Later among the works it cites.
Watermarking Text Data on Large Language Models for Dataset Copyright Protection
Yixin Liu, Hongsheng Hu, Xuyun Zhang, and Lichao Sun. 2023e · 2023
Later among the works it cites.
Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models’ Alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023l · 2023
Later among the works it cites.
Privacy-Enhanced Knowledge Transfer with Collaborative Split Learning over Teacher Ensembles. In Proceedings of the 2023 Secure and Trustworthy Deep Learning Systems Workshop . 1–13
Ziyao Liu, Jiale Guo, Mengmeng Yang, Wenzhuo Yang, Jiani Fan, and Kwok-Yan Lam. 2023d · 2023
Later among the works it cites.
A Survey on Federated Unlearning: Challenges, Methods, and Future Directions
Ziyao Liu, Yu Jiang, Jiyuan Shen, Minyi Peng, Kwok-Yan Lam, Xingliang Yuan, and Xiaoning Liu. 2023g · 2023
Later among the works it cites.
Long-Term Privacy-Preserving Aggregation With User-Dynamics for Federated Learning
Ziyao Liu, Hsiao-Ying Lin, and Yamin Liu. 2023i · 2023
Later among the works it cites.
Adapt in Contexts: Retrieval-Augmented Domain Adaptation via In-Context Learning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 6525–6542
Quanyu Long, Wenya Wang, and Sinno Jialin Pan. 2023 · 2023
Later among the works it cites.
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, and Daphne Ippolito. 2023 · 2023
Later among the works it cites.
Mechanistic Mode Connectivity. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). 22965–23004
Ekdeep Singh Lubana, Eric J. Bigelow, Robert P. Dick, David Scott Krueger, and Hidenori Tanaka. 2023 · 2023
Later among the works it cites.
An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang. 2023c · 2023
Later among the works it cites.
ChatGPT as a Factual Inconsistency Evaluator for Abstractive Text Summarization
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023a · 2023
Later among the works it cites.
Augmented Large Language Models with Parametric Knowledge Guiding
Ziyang Luo, Can Xu, Pu Zhao, Xiubo Geng, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023b · 2023
Later among the works it cites.
The Misinformation Susceptibility Test (MIST): A psychometrically validated measure of news veracity discernment
Rakoen Maertens, Friedrich M Götz, Hudson F Golino, Jon Roozenbeek, Claudia R Schneider, Yara Kyrychenko, John R Kerr, Stefan Stieger, William P McClanahan, Karly Drabot, et al · 2023
Later among the works it cites.
A geometric framework for fairness. In Proceedings of the 1st Workshop on Fairness and Bias in AI co-located with 26th European Conference on Artificial Intelligence (ECAI 2023), Kraków, Poland, October 1st, 2023 (CEUR Workshop Proceedings, Vol. 3523) , Roberta Calegari, Andrea Aler Tubella, Gabriel González-Castañé, Virginia Dignum, and Michela Milano (Eds.)
Alessandro Maggio, Luca Giuliani, Roberta Calegari, Michele Lombardi, and Michela Milano. 2023 · 2023
Later among the works it cites.
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 9004–9017
Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023 · 2023
Later among the works it cites.
Online public discourse on artificial intelligence and ethics in China: context, content, and implications
Yishu Mao and Kristin Shi-Kupfer. 2023 · 2023
Later among the works it cites.
A Holistic Approach to Undesired Content Detection in the Real World. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023 , Brian Williams, Yiling Chen, and Jennifer Neville (Eds.). 15009–15018
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023 · 2023
Later among the works it cites.
ChatGPT banned in Italy over privacy concerns
Shiona McCallum. 2023 · 2023
Later among the works it cites.
Using In-Context Learning to Improve Dialogue Safety. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 11882–11910
Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tur. 2023 · 2023
Later among the works it cites.
MaNtLE: Model-agnostic Natural Language Explainer. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 13493–13511
Rakesh R. Menon, Kerem Zaman, and Shashank Srivastava. 2023 · 2023
Later among the works it cites.
Responsible Use Guide: your resource for building responsibly
Meta. 2023 · 2023
Later among the works it cites.
SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning
Ning Miao, Yee Whye Teh, and Tom Rainforth. 2023 · 2023
Later among the works it cites.
The Quantization Model of Neural Scaling. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Eric J. Michaud, Ziming Liu, Uzay Girit, and Max Tegmark. 2023 · 2023
Later among the works it cites.
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2023 · 2023
Later among the works it cites.
Operationalising AI governance through ethics-based auditing: an industry case study
Jakob Mökander and Luciano Floridi. 2023 · 2023
Later among the works it cites.
Scalable Extraction of Training Data from (Production) Language Models
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. 2023 · 2023
Later among the works it cites.
Exploring the Differences Between Narrow AI, General AI, and Superintelligent AI
Institute of data. 2023 · 2023
Later among the works it cites.
OpenAI. 2023a · 2023
Later among the works it cites.
Med-HALT: Medical Domain Hallucination Test for Large Language Models. In Proceedings of the 27th Conference on Computational Natural Language Learning, CoNLL 2023, Singapore, December 6-7, 2023 , Jing Jiang, David Reitter, and Shumin Deng (Eds.). 314–334
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2023b · 2023
Later among the works it cites.
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State. In Proceedings of the 27th Conference on Computational Natural Language Learning, CoNLL 2023, Singapore, December 6-7, 2023 , Jing Jiang, David Reitter, and Shumin Deng (Eds.). 548–560
Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C. Wallace, and David Bau. 2023a · 2023
Later among the works it cites.
On the Risk of Misinformation Pollution with Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 1389–1403
Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Yang Wang. 2023 · 2023
Later among the works it cites.
Toward AI governance: Identifying best practices and potential barriers and outcomes
Emmanouil Papagiannidis, Ida Merete Enholm, Chirstian Dremel, Patrick Mikalef, and John Krogstie. 2023 · 2023
Later among the works it cites.
AI Deception: A Survey of Examples, Risks, and Potential Solutions
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. 2023 · 2023
Later among the works it cites.
Rodrigo Pedro, Daniel Castro, Paulo Carreira, and Nuno Santos. 2023 · 2023
Later among the works it cites.
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, and Jianfeng Gao. 2023a · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023b · 2023
Later among the works it cites.
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
Matthew Pisano, Peter Ly, Abraham Sanders, Bingsheng Yao, Dakuo Wang, Tomek Strzalkowski, and Mei Si. 2023 · 2023
Later among the works it cites.
Reasoning with Language Model Prompting: A Survey. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 5368–5393
Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2023 · 2023
Later among the works it cites.
FedCIO: Efficient Exact Federated Unlearning with Clustering, Isolation, and One-shot Aggregation. In 2023 IEEE International Conference on Big Data (BigData) . IEEE, 5559–5568
Hongyu Qiu, Yongwei Wang, Yonghui Xu, Lizhen Cui, and Zhiqi Shen. 2023 · 2023
Later among the works it cites.
Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, CCS 2023, Copenhagen, Denmark, November 26-30, 2023 , Weizhi Meng, Christian Damsgaard Jensen, Cas Cremers, and Engin Kirda (Eds.). 3403–3417
Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, and Yang Zhang. 2023 · 2023
Later among the works it cites.
Understanding Addition in Transformers
Philip Quirke and Fazl Barez. 2023 · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Conformal Nucleus Sampling. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 27–34
Shauli Ravfogel, Yoav Goldberg, and Jacob Goldberger. 2023 · 2023
Later among the works it cites.
A Survey of Hallucination in Large Foundation Models
Vipula Rawte, Amit P. Sheth, and Amitava Das. 2023 · 2023
Later among the works it cites.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails
Traian Rebedea, Razvan Dinu, Makesh Sreedhar, Christopher Parisien, and Jonathan Cohen. 2023 · 2023
Later among the works it cites.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2023 · 2023
Later among the works it cites.
Can AI-Generated Text be Reliably Detected?
Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi. 2023 · 2023
Later among the works it cites.
Raising the Cost of Malicious AI-Powered Image Editing. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). 29894–29918
Hadi Salman, Alaa Khaddaj, Guillaume Leclerc, Andrew Ilyas, and Aleksander Madry. 2023 · 2023
Later among the works it cites.
Governance of superintelligence
Ilya Sutskever Sam Altman, Greg Brockman. 2023 · 2023
Later among the works it cites.
Embarrassingly Simple Text Watermarks
Ryoma Sato, Yuki Takezawa, Han Bao, Kenta Niwa, and Makoto Yamada. 2023 · 2023
Later among the works it cites.
Artificial intelligence governance for businesses
Johannes Schneider, Rene Abraham, Christian Meske, and Jan Vom Brocke. 2023 · 2023
Later among the works it cites.
Towards best practices in AGI safety and governance: A survey of expert opinion
Jonas Schuett, Noemi Dreksler, Markus Anderljung, David McCaffary, Lennart Heim, Emma Bluemke, and Ben Garfinkel. 2023 · 2023
Later among the works it cites.
Elizabeth Seger, Noemi Dreksler, Richard Moulange, Emily Dardaman, Jonas Schuett, K Wei, Christoph Winter, Mackenzie Arnold, Seán Ó hÉigeartaigh, Anton Korinek, et al · 2023
Later among the works it cites.
AI Imitating Artist "Style" Drives Call to Rethink Copyright Law
Riddhi Setty. 2023 · 2023
Later among the works it cites.
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Rusheb Shah, Quentin Feuillade-Montixi, Soroush Pour, Arush Tagade, Stephen Casper, and Javier Rando. 2023 · 2023
Later among the works it cites.
Towards Understanding Sycophancy in Language Models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2023 · 2023
Later among the works it cites.
Impact of big data analytics and ChatGPT on cybersecurity. In 2023 4th International Conference on Computing and Communication Systems (I3CS) . IEEE, 1–6
Pawankumar Sharma and Bibhu Dash. 2023 · 2023
Later among the works it cites.
Practices for governing agentic ai systems
Yonadav Shavit, Sandhini Agarwal, Miles Brundage, Steven Adler, Cullen O’Keefe, Rosie Campbell, Teddy Lee, Pamela Mishkin, Tyna Eloundou, Alan Hickey, et al · 2023
Later among the works it cites.
Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael B. Abu-Ghazaleh. 2023c · 2023
Later among the works it cites.
Shaping the Emerging Norms of Using Large Language Models in Social Computing Research. In Computer Supported Cooperative Work and Social Computing, CSCW 2023, Minneapolis, MN, USA, October 14-18, 2023 , Casey Fiesler, Loren G. Terveen, Morgan Ames, Susan R. Fussell, Eric Gilbert, Vera Liao, Xiaojuan Ma, Xinru Page, Mark Rouncefield, Vivek Singh, and Pamela J. Wisniewski (Eds.). 569–571
Hong Shen, Tianshi Li, Toby Jia-Jun Li, Joon Sung Park, and Diyi Yang. 2023b · 2023
Later among the works it cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023a · 2023
Later among the works it cites.
Large Language Models Can Be Easily Distracted by Irrelevant Context. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). 31210–31227
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. 2023b · 2023
Later among the works it cites.
Safer-Instruct: Aligning Language Models with Automated Preference Data
Taiwei Shi, Kai Chen, and Jieyu Zhao. 2023a · 2023
Later among the works it cites.
Human-centricity in AI governance: A systemic approach
Anton Sigfrids, Jaana Leikas, Henrikki Salo-Pöntinen, and Emmi Koskimies. 2023 · 2023
Later among the works it cites.
Tree Prompting: Efficient Task Adaptation without Fine-Tuning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 6253–6267
Chandan Singh, John X. Morris, Alexander M. Rush, Jianfeng Gao, and Yuntian Deng. 2023 · 2023
Later among the works it cites.
AI model GPT-3 (dis)informs us better than humans
Giovanni Spitale, Nikola Biller-Andorno, and Federico Germani. 2023 · 2023
Later among the works it cites.
A new ChatGPT Zero Day Attack is Undetectable Malware That Steals Data
QS Study. 2023 · 2023
Later among the works it cites.
Adapting Fake News Detection to the Era of Large Language Models
Jinyan Su, Claire Cardie, and Preslav Nakov. 2023 · 2023
Later among the works it cites.
EVA-CLIP: Improved Training Techniques for CLIP at Scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023a · 2023
Later among the works it cites.
Yanshen Sun, Jianfeng He, Shuo Lei, Limeng Cui, and Chang-Tien Lu. 2023b · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Later among the works it cites.
Top French university bans use of ChatGPT to prevent plagiarism
The Straits Times. 2023 · 2023
Later among the works it cites.
Individual (non) resilience of university students to digital media manipulation after COVID-19 (case study of Slovak initiatives)
Hedviga Tkácová, Martina Pavlíková, Eva Stranovská, and Roman Králik. 2023 · 2023
Later among the works it cites.
Announcing OpenChatKit
together.ai. 2023 · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a · 2023
Later among the works it cites.
CREST: A Joint Framework for Rationalization and Counterfactual Text Generation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 15109–15126
Marcos V. Treviso, Alexis Ross, Nuno Miguel Guerreiro, and André F. T. Martins. 2023 · 2023
Later among the works it cites.
How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs
Haoqin Tu, Chenhang Cui, Zijun Wang, Yiyang Zhou, Bingchen Zhao, Junlin Han, Wangchunshu Zhou, Huaxiu Yao, and Cihang Xie. 2023 · 2023
Later among the works it cites.
AI and global governance: Modalities, rationales, tensions
Michael Veale, Kira Matus, and Robert Gorwa. 2023 · 2023
Later among the works it cites.
Unmasking Nationality Bias: A Study of Human Perception of Nationalities in AI-Generated Articles. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, AIES 2023, Montréal, QC, Canada, August 8-10, 2023 , Francesca Rossi, Sanmay Das, Jenny Davis, Kay Firth-Butterfield, and Alex John (Eds.). 554–565
Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Kenneth Huang, and Shomir Wilson. 2023a · 2023
Later among the works it cites.
Pranav Narayanan Venkit, Mukund Srinath, and Shomir Wilson. 2023b · 2023
Later among the works it cites.
FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry W. Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc V. Le, and Thang Luong. 2023 · 2023
Later among the works it cites.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023a · 2023
Later among the works it cites.
Can ChatGPT Defend its Belief in Truth? Evaluating LLM Reasoning via Debate. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 11865–11881
Boshi Wang, Xiang Yue, and Huan Sun. 2023i · 2023
Later among the works it cites.
Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity
Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Jiayang Cheng, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, Yidong Wang, Linyi Yang, Jindong Wang, Xing Xie, Zheng Zhang, and Yue Zhang. 2023f · 2023
Later among the works it cites.
On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, Binxing Jiao, Yue Zhang, and Xing Xie. 2023d · 2023
Later among the works it cites.
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023g · 2023
Later among the works it cites.
Robust Backdoor Attack with Visible, Semantic, Sample-Specific, and Compatible Triggers
Ruotong Wang, Hongrui Chen, Zihao Zhu, Li Liu, Yong Zhang, Yanbo Fan, and Baoyuan Wu. 2023b · 2023
Later among the works it cites.
Knowledge Editing for Large Language Models: A Survey
Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, and Jundong Li. 2023k · 2023
Later among the works it cites.
Machine Unlearning via Representation Forgetting With Parameter Self-Sharing
Weiqi Wang, Chenhan Zhang, Zhiyi Tian, and Shui Yu. 2023j · 2023
Later among the works it cites.
KnowledGPT: Enhancing Large Language Models with Retrieval and Storage Access on Knowledge Bases
Xintao Wang, Qianwen Yang, Yongting Qiu, Jiaqing Liang, Qianyu He, Zhouhong Gu, Yanghua Xiao, and Wei Wang. 2023h · 2023
Later among the works it cites.
Self-Instruct: Aligning Language Models with Self-Generated Instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 13484–13508
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023e · 2023
Later among the works it cites.
Implementing BERT and fine-tuned RobertA to detect AI generated news by ChatGPT
Zecong Wang, Jiaxi Cheng, Chen Cui, and Chenhao Yu. 2023c · 2023
Later among the works it cites.
Jailbroken: How Does LLM Safety Training Fail?. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023a · 2023
Later among the works it cites.
Simple synthetic data reduces sycophancy in large language models
Jerry W. Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V. Le. 2023b · 2023
Later among the works it cites.
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Zeming Wei, Yifei Wang, and Yisen Wang. 2023c · 2023
Later among the works it cites.
Large Language Models are Better Reasoners with Self-Verification. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 2550–2575
Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao. 2023 · 2023
Later among the works it cites.
How Strangers Got My Email Address From ChatGPT’s Model
Jeremy White. 2023 · 2023
Later among the works it cites.
Lessons learned from translating AI from development to deployment in healthcare
Kasumi Widner, Sunny Virmani, Jonathan Krause, Jay Nayar, Richa Tiwari, Elin Rønby Pedersen, Divleen Jeji, Naama Hammel, Yossi Matias, Greg S Corrado, et al · 2023
Later among the works it cites.
Cheating Aussie student fails uni exam after being caught using artificial intelligence chatbot to write essay - now Australia’s top universities are considering a bizarre solution to stop it happening again
Hannah Wilcox. 2023 · 2023
Later among the works it cites.
Trustworthy AI: Deciding What to Decide
Caesar Wu, Yuan-Fang Li, Jian Li, Jingjing Xu, and Pascal Bouvry. 2023a · 2023
Later among the works it cites.
Fake News in Sheep’s Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks
Jiaying Wu and Bryan Hooi. 2023 · 2023
Later among the works it cites.
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
Yuanwei Wu, Xiang Li, Yixin Liu, Pan Zhou, and Lichao Sun. 2023b · 2023
Later among the works it cites.
DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy Liang, Quoc V. Le, Tengyu Ma, and Adams Wei Yu. 2023 · 2023
Later among the works it cites.
Contrastive Post-training Large Language Models on Data Curriculum
Canwen Xu, Corby Rosset, Luciano Del Corro, Shweti Mahajan, Julian J. McAuley, Jennifer Neville, Ahmed Hassan Awadallah, and Nikhil Rao. 2023b · 2023
Later among the works it cites.
An LLM can Fool Itself: A Prompt-Based Adversarial Attack
Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, and Mohan S. Kankanhalli. 2023a · 2023
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection. In NeurIPS 2023 Workshop on Backdoors in Deep Learning-The Good, the Bad, and the Ugly
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin. 2023 · 2023
Later among the works it cites.
Large Language Models as Optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. 2023e · 2023
Later among the works it cites.
Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions
Hui Yang, Sifu Yue, and Yunzhong He. 2023f · 2023
Later among the works it cites.
Local differential privacy and its applications: A comprehensive survey
Mengmeng Yang, Taolin Guo, Tianqing Zhu, Ivan Tjuawinata, Jun Zhao, and Kwok-Yan Lam. 2023b · 2023
Later among the works it cites.
Watermarking Text Generated by Black-Box Language Models
Xi Yang, Kejiang Chen, Weiming Zhang, Chang Liu, Yuang Qi, Jie Zhang, Han Fang, and Nenghai Yu. 2023a · 2023
Later among the works it cites.
Data Poisoning Attacks Against Multimodal Encoders. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). 39299–39313
Ziqing Yang, Xinlei He, Zheng Li, Michael Backes, Mathias Humbert, Pascal Berrang, and Yang Zhang. 2023c · 2023
Later among the works it cites.
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. 2023d · 2023
Later among the works it cites.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023b · 2023
Later among the works it cites.
A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Eric Sun, and Yue Zhang. 2023a · 2023
Later among the works it cites.
Robust Multi-bit Natural Language Watermarking through Invariant Features. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 2092–2115
KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak. 2023 · 2023
Later among the works it cites.
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2023 · 2023
Later among the works it cites.
GLM-130B: An Open Bilingual Pre-trained Model. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023 · 2023
Later among the works it cites.
Transfer Attacks and Defenses for Large Language Models on Coding Tasks
Chi Zhang, Zifan Wang, Ravi Mangal, Matt Fredrikson, Limin Jia, and Corina S. Pasareanu. 2023e · 2023
Later among the works it cites.
Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models
Hanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi, Giuseppe Ateniese, and Boaz Barak. 2023b · 2023
Later among the works it cites.
An Efficient FHE-Enabled Secure Cloud-Edge Computing Architecture for IoMT Data Protection With its Application to Pandemic Modeling
Linru Zhang, Xiangning Wang, Jiabo Wang, Rachael Pung, Huaxiong Wang, and Kwok-Yan Lam. 2024g · 2023
Later among the works it cites.
FedRecovery: Differentially Private Machine Unlearning for Federated Learning Frameworks
Lefeng Zhang, Tianqing Zhu, Haibin Zhang, Ping Xiong, and Wanlei Zhou. 2023f · 2023
Later among the works it cites.
AgrEvader: Poisoning membership inference against Byzantine-robust federated learning. In Proceedings of the ACM Web Conference 2023 . 2371–2382
Yanjun Zhang, Guangdong Bai, Mahawaga Arachchige Pathum Chamikara, Mengyao Ma, Liyue Shen, Jingwei Wang, Surya Nepal, Minhui Xue, Long Wang, and Joseph Liu. 2023a · 2023
Later among the works it cites.
Prompts Should not be Seen as Secrets: Systematically Measuring Prompt Extraction Attack Success
Yiming Zhang and Daphne Ippolito. 2023 · 2023
Later among the works it cites.
Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023c · 2023
Later among the works it cites.
Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023d · 2023
Later among the works it cites.
Explainability for Large Language Models: A Survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2023c · 2023
Later among the works it cites.
Retrieving Multimodal Information for Augmented Generation: A Survey. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 4736–4756
Ruochen Zhao, Hailin Chen, Weishi Wang, Fangkai Jiao, Do Xuan Long, Chengwei Qin, Bosheng Ding, Xiaobao Guo, Minzhi Li, Xingxuan Li, and Shafiq Joty. 2023b · 2023
Later among the works it cites.
Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 5823–5840
Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin, and Lidong Bing. 2023e · 2023
Later among the works it cites.
Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 5823–5840
Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin, and Lidong Bing. 2023f · 2023
Later among the works it cites.
Provable Robust Watermarking for AI-Generated Text
Xuandong Zhao, Prabhanjan Ananth, Lei Li, and Yu-Xiang Wang. 2023a · 2023
Later among the works it cites.
Calibrating Sequence likelihood Improves Conditional Language Generation. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023
Yao Zhao, Misha Khalman, Rishabh Joshi, Shashi Narayan, Mohammad Saleh, and Peter J. Liu. 2023d · 2023
Later among the works it cites.
On Evaluating Adversarial Robustness of Large Vision-Language Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Cheung, and Min Lin. 2023g · 2023
Later among the works it cites.
Understanding and Improving Adversarial Attacks on Latent Diffusion Model
Boyang Zheng, Chumeng Liang, Xiaoyu Wu, and Yan Liu. 2023b · 2023
Later among the works it cites.
Why does chatgpt fall short in providing truthful answers
Shen Zheng, Jie Huang, and Kevin Chen-Chuan Chang. 2023a · 2023
Later among the works it cites.
LIMA: Less Is More for Alignment. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.)
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023b · 2023
Later among the works it cites.
A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). 11082–11100
Yi Zhou, José Camacho-Collados, and Danushka Bollegala. 2023a · 2023
Later among the works it cites.
Beyond one-preference-for-all: Multi-objective direct preference optimization
Zhanhui Zhou, Jie Liu, Chao Yang, Jing Shao, Yu Liu, Xiangyu Yue, Wanli Ouyang, and Yu Qiao. 2023c · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023a · 2023
Later among the works it cites.
PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, and Xing Xie. 2023b · 2023
Later among the works it cites.
A Pilot Study of Query-Free Adversarial Attack against Stable Diffusion. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023 - Workshops, Vancouver, BC, Canada, June 17-24, 2023 . 2385–2392
Haomin Zhuang, Yihua Zhang, and Sijia Liu. 2023 · 2023
Later among the works it cites.
On Robustness of Prompt-based Semantic Parsing with Large Pre-trained Language Model: An Empirical Study on Codex. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2-6, 2023 , Andreas Vlachos and Isabelle Augenstein (Eds.). 1090–1102
Terry Yue Zhuo, Zhuang Li, Yujin Huang, Fatemeh Shiri, Weiqing Wang, Gholamreza Haffari, and Yuan-Fang Li. 2023 · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic. 2024 · 2024
Closest in time.
Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models
Guangji Bai, Zheng Chai, Chen Ling, Shiyu Wang, Jiaying Lu, Nan Zhang, Tingwei Shi, Ziyang Yu, Mengdan Zhu, Yifei Zhang, Carl Yang, Yue Cheng, and Liang Zhao. 2024 · 2024
Closest in time.
Measuring Political Bias in Large Language Models: What Is Said and How It Is Said
Yejin Bang, Delong Chen, Nayeon Lee, and Pascale Fung. 2024 · 2024
Closest in time.
The Dark Side of Language Models: Exploring the Potential of LLMs in Multimedia Disinformation Generation and Dissemination
Dipto Barman, Ziyi Guo, and Owen Conlan. 2024 · 2024
Closest in time.
Mechanistic Interpretability for AI Safety - A Review
Leonard Bereska and Efstratios Gavves. 2024 · 2024
Closest in time.
Graph of Thoughts: Solving Elaborate Problems with Large Language Models. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada , Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan (Eds.). 17682–17690
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2024 · 2024
Closest in time.
Creating video from text
Clarence Ng David Schnurr Eric Luhman Joe Taylor Li Jing Natalie Summers Ricky Wang Rohan Sahai Ryan O’Rourke Troy Luhman Will DePue Yufei Guo Connor Holmes Bill Peebles Tim Brooks. 2024 · 2024
Closest in time.
Locating and Mitigating Gender Bias in Large Language Models
Yuchen Cai, Ding Cao, Rongxi Guo, Yaqin Wen, Guiquan Liu, and Enhong Chen. 2024 · 2024
Closest in time.
Introduction to Special Issue on Trustworthy Artificial Intelligence
Roberta Calegari, Fosca Giannotti, Francesca Pratesi, and Michela Milano. 2024 · 2024
Closest in time.
Can LLMs’ Tuning Methods Work in Medical Multimodal Domain?
Jiawei Chen, Yue Jiang, Dingkang Yang, Mingcheng Li, Jinjie Wei, Ziyun Qian, and Lihua Zhang. 2024a · 2024
Closest in time.
Evaluating large language models in medical applications: a survey
Xiaolan Chen, Jiayang Xiang, Shanfu Lu, Yexin Liu, Mingguang He, and Danli Shi. 2024b · 2024
Closest in time.
Typographic Attacks in Large Multimodal Models Can be Alleviated by More Informative Prompts
Hao Cheng, Erjia Xiao, and Renjing Xu. 2024 · 2024
Closest in time.
Safety Cases: How to Justify the Safety of Advanced AI Systems
Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. 2024 · 2024
Closest in time.
SaulLM-7B: A pioneering Large Language Model for Law
Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Dominic Culver, Rui Melo, Caio Corro, André F. T. Martins, Fabrizio Esposito, Vera Lúcia Raposo, Sofia Morgado, and Michael Desa. 2024 · 2024
Closest in time.
Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, Zhixing Tan, Junwu Xiong, Xinyu Kong, Zujie Wen, Ke Xu, and Qi Li. 2024 · 2024
Closest in time.
Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E. Ho. 2024b · 2024
Closest in time.
Security and Privacy Challenges of Large Language Models: A Survey
Badhan Chandra Das, M. Hadi Amini, and Yanzhao Wu. 2024 · 2024
Closest in time.
MASTERKEY: Automated jailbreaking of large language model chatbots. In Proc. ISOC NDSS
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2024 · 2024
Closest in time.
Strategic Data Revocation in Federated Unlearning. In IEEE INFOCOM 2024-IEEE Conference on Computer Communications . IEEE
Ningning Ding, Ermin Wei, and Randall Berry. 2024 · 2024
Closest in time.
Avoiding Copyright Infringement via Machine Unlearning
Guangyao Dou, Zheyuan Liu, Qing Lyu, Kaize Ding, and Eric Wong. 2024 · 2024
Closest in time.
An Interactive Agent Foundation Model
Zane Durante, Bidipta Sarkar, Ran Gong, Rohan Taori, Yusuke Noda, Paul Tang, Ehsan Adeli, Shrinidhi Kowshika Lakshmikanth, Kevin A. Schulman, Arnold Milstein, Demetri Terzopoulos, Ade Famoti, Noboru Kuno, Ashley J. Llorens, Hoi Vo, Katsushi Ikeuchi, Li Fei-Fei, Jianfeng Gao, Naoki Wake, and Qiuyuan Huang. 2024 · 2024
Closest in time.
Risks and Opportunities of Open-Source Generative AI
Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schröder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Aaron Purewal, Botos Csaba, Fabro Steibel, Fazel Keshtkar, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan Arturo Nolazco, Lori Landay, Matthew Jackson, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, and Jakob N. Foerster. 2024 · 2024
Closest in time.
Yihe Fan, Yuxin Cao, Ziyu Zhao, Ziyao Liu, and Shaofeng Li. 2024 · 2024
Closest in time.
Algorithmic Collusion by Large Language Models
Sara Fish, Yannai A Gonczarowski, and Ran I Shorrer. 2024 · 2024
Closest in time.
Generative AI’s privacy problem
Ina Fried. 2024 · 2024
Closest in time.
Damián Ariel Furman, Juan Junqueras, Z Burçe Gümüslü, Edgar Altszyler, Joaquin Navajas, Ophelia Deroy, and Justin Sulik. 2024 · 2024
Closest in time.
Confucius: Iterative Tool Learning from Introspection Feedback by Easy-to-Difficult Curriculum. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada , Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan (Eds.). 18030–18038
Shen Gao, Zhengliang Shi, Minghang Zhu, Bowen Fang, Xin Xin, Pengjie Ren, Zhumin Chen, Jun Ma, and Zhaochun Ren. 2024 · 2024
Closest in time.
Responsible Artificial Intelligence: A Structured Literature Review
Sabrina Göllner, Marina Tropmann-Frick, and Bostjan Brumen. 2024 · 2024
Closest in time.
Adversarial Prompting in LLMs
Prompt Engineering Guide. 2024 · 2024
Closest in time.
Uri Hacohen, Adi Haviv, Shahar Sarfaty, Bruria Friedman, Niva Elkin-Koren, Roi Livni, and Amit H. Bermano. 2024 · 2024
Closest in time.
Towards Independence Criterion in Machine Unlearning of Features and Labels
Ling Han, Nanqing Luo, Hao Huang, Jing Chen, and Mary-Anne Hartley. 2024a · 2024
Closest in time.
LLM Multi-Agent Systems: Challenges and Open Problems
Shanshan Han, Qifan Zhang, Yuhang Yao, Weizhao Jin, Zhaozhuo Xu, and Chaoyang He. 2024b · 2024
Closest in time.
Machine-Made Media: Monitoring the Mobilization of Machine-Generated Articles on Misinformation and Mainstream News Websites. In Proceedings of the Eighteenth International AAAI Conference on Web and Social Media, ICWSM 2024, Buffalo, New York, USA, June 3-6, 2024 , Yu-Ru Lin, Yelena Mejova, and Meeyoung Cha (Eds.). 542–556
Hans W. A. Hanley and Zakir Durumeric. 2024 · 2024
Closest in time.
Vision-fused Jailbreak: A Multi-modal Collaborative Jailbreak Attack
Haojie Hao, Jiakai Wang, Hainan Li, and Zhilei Zhu. 2024 · 2024
Closest in time.
Curiosity-driven Red-teaming for Large Language Models
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang, Aldo Pareja, James R. Glass, Akash Srivastava, and Pulkit Agrawal. 2024 · 2024
Closest in time.
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning. In 2024 IEEE Symposium on Security and Privacy (SP)
Hongsheng Hu, Shuo Wang, Tian Dong, and Minhui Xue. 2024b · 2024
Closest in time.
Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models
Xijie Huang, Xinyuan Wang, Hantao Zhang, Jiawen Xi, Jingkun An, Hao Wang, and Chengwei Pan. 2024 · 2024
Closest in time.
Fast-FedUL: A Training-Free Federated Unlearning with Provable Skew Resilience
Thanh Trung Huynh, Trong Bang Nguyen, Phi Le Nguyen, Thanh Tam Nguyen, Matthias Weidlich, Quoc Viet Hung Nguyen, and Karl Aberer. 2024 · 2024
Closest in time.
Model AI Governance Framework for Generative AI
Singapore IMDA. 2024 · 2024
Closest in time.
Towards Efficient and Certified Recovery from Poisoning Attacks in Federated Learning
Yu Jiang, Jiyuan Shen, Ziyao Liu, Chee Wei Tan, and Kwok-Yan Lam. 2024 · 2024
Closest in time.
People cannot distinguish GPT-4 from a human in a Turing test
Cameron R. Jones and Benjamin K. Bergen. 2024 · 2024
Closest in time.
Watermark Stealing in Large Language Models
Nikola Jovanovic, Robin Staab, and Martin T. Vechev. 2024 · 2024
Closest in time.
Singapore Gum: A Look at the Country’s Chewing Gum Ban
kaizenaire. 2024 · 2024
Closest in time.
The Ethics of Interaction: Mitigating Security Threats in LLMs
Ashutosh Kumar, Sagarika Singh, Shiv Vignesh Murty, and Swathy Ragupathy. 2024 · 2024
Closest in time.
A Survey of Large Language Models in Finance (FinLLMs)
Jean Lee, Nicholas Stevens, Soyeon Caren Han, and Minseok Song. 2024 · 2024
Closest in time.
FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs
Jinmin Li, Kuofeng Gao, Yang Bai, Jingyun Zhang, Shutao Xia, and Yisen Wang. 2024a · 2024
Closest in time.
Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. 2024b · 2024
Closest in time.
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
Jiawei Liang, Siyuan Liang, Man Luo, Aishan Liu, Dongchen Han, Ee-Chien Chang, and Xiaochun Cao. 2024a · 2024
Closest in time.
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
Jiawei Liang, Siyuan Liang, Man Luo, Aishan Liu, Dongchen Han, Ee-Chien Chang, and Xiaochun Cao. 2024b · 2024
Closest in time.
Against The Achilles’ Heel: A Survey on Red Teaming for Generative Models
Lizhi Lin, Honglin Mu, Zenan Zhai, Minghan Wang, Yuxia Wang, Renxi Wang, Junjie Gao, Yixuan Zhang, Wanxiang Che, Timothy Baldwin, Xudong Han, and Haonan Li. 2024b · 2024
Closest in time.
Incentive and Dynamic Client Selection for Federated Unlearning
Yijing Lin, Zhipeng Gao, Hongyang Du, Dusit Niyato, Jiawen Kang, and Xiaoyuan Liu. 2024a · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024d · 2024
Closest in time.
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
Xiaoze Liu, Ting Sun, Tianyang Xu, Feijie Wu, Cunxiang Wang, Xiaoqian Wang, and Jing Gao. 2024e · 2024
Closest in time.
Safety of Multimodal Large Language Models on Images and Text
Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang, and Yu Qiao. 2024j · 2024
Closest in time.
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, and Lichao Sun. 2024i · 2024
Closest in time.
Towards Safer Large Language Models through Machine Unlearning. In Findings of the Association for Computational Linguistics: ACL 2024
Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang. 2024a · 2024
Closest in time.
Breaking the Trilemma of Privacy, Utility, and Efficiency via Controllable Machine Unlearning
Zheyuan Liu, Guangyao Dou, Yijun Tian, Chunhui Zhang, Eli Chien, and Ziwei Zhu. 2024b · 2024
Closest in time.
Guaranteeing Data Privacy in Federated Unlearning with Dynamic User Participation
Ziyao Liu, Yu Jiang, Weifeng Jiang, Jiale Guo, Jun Zhao, and Kwok-Yan Lam. 2024c · 2024
Closest in time.
Backdoor Attacks via Machine Unlearning. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada , Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan (Eds.). 14115–14123
Zihao Liu, Tianhao Wang, Mengdi Huai, and Chenglin Miao. 2024f · 2024
Closest in time.
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
Ziyao Liu, Huanyi Ye, Chen Chen, and Kwok-Yan Lam. 2024g · 2024
Closest in time.
Privacy-Preserving Federated Unlearning with Certified Client Removal
Ziyao Liu, Huanyi Ye, Yu Jiang, Jiyuan Shen, Jiale Guo, Ivan Tjuawinata, and Kwok-Yan Lam. 2024h · 2024
Closest in time.
Test-Time Backdoor Attacks on Multimodal Large Language Models
Dong Lu, Tianyu Pang, Chao Du, Qian Liu, Xianjun Yang, and Min Lin. 2024a · 2024
Closest in time.
Responsible AI pattern catalogue: A collection of best practices for AI governance and engineering
Qinghua Lu, Liming Zhu, Xiwei Xu, Jon Whittle, Didar Zowghi, and Aurelie Jacquet. 2024b · 2024
Closest in time.
From Understanding to Utilization: A Survey on Explainability for Large Language Models
Haoyan Luo and Lucia Specia. 2024 · 2024
Closest in time.
Secformer: Towards fast and accurate privacy-preserving inference for large language models
Jinglong Luo, Yehong Zhang, Jiaqi Zhang, Xin Mu, Hui Wang, Yue Yu, and Zenglin Xu. 2024 · 2024
Closest in time.
Large Language Models are Geographically Biased
Rohin Manvi, Samar Khanna, Marshall Burke, David B. Lobell, and Stefano Ermon. 2024 · 2024
Closest in time.
The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
Tula Masterman, Sandi Besen, Mason Sawtell, and Alex Chao. 2024 · 2024
Closest in time.
Cybersecurity statistics in 2024
Sierra Campbell Mehdi Punjwani. 2024 · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
AI Meta. 2024 · 2024
Closest in time.
Large Language Models: A Survey
Shervin Minaee, Tomás Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024 · 2024
Closest in time.
More human than human: Measuring ChatGPT political bias
Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. 2024 · 2024
Closest in time.
Secret Collusion Among Generative AI Agents
Sumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina, Philip H. S. Torr, Lewis Hammond, and Christian Schröder de Witt. 2024 · 2024
Closest in time.
Terrence Neumann, Sooyong Lee, Maria De-Arteaga, Sina Fazelpour, and Matthew Lease. 2024 · 2024
Closest in time.
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
Zhenyang Ni, Rui Ye, Yuxi Wei, Zhen Xiang, Yanfeng Wang, and Siheng Chen. 2024 · 2024
Closest in time.
Jailbreaking Attack against Multimodal Large Language Model
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin. 2024 · 2024
Closest in time.
Unifying Large Language Models and Knowledge Graphs: A Roadmap
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024 · 2024
Closest in time.
Teach LLMs to Phish: Stealing Private Information from Language Models
Ashwinee Panda, Christopher A Choquette-Choo, Zhengming Zhang, Yaoqing Yang, and Prateek Mittal. 2024 · 2024
Closest in time.
No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
Qi Pang, Shengyuan Hu, Wenting Zheng, and Virginia Smith. 2024 · 2024
Closest in time.
Visual Adversarial Examples Jailbreak Aligned Large Language Models. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada , Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan (Eds.). 21527–21536
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. 2024 · 2024
Closest in time.
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
Yusu Qian, Haotian Zhang, Yinfei Yang, and Zhe Gan. 2024 · 2024
Closest in time.
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
Maan Qraitem, Nazia Tasnim, Piotr Teterwak, Kate Saenko, and Bryan A. Plummer. 2024 · 2024
Closest in time.
ChatGPT: Italy says OpenAI’s chatbot breaches data protection rules
Imran Rahman-Jones. 2024 · 2024
Closest in time.
WARM: On the Benefits of Weight Averaged Reward Models
Alexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi, Geoffrey Cideron, Olivier Bachem, and Johan Ferret. 2024 · 2024
Closest in time.
Cybercrime To Cost The World $9.5 Trillion USD Annually In 2024
2023 Annual Cybercrime Report. 2024 · 2024
Closest in time.
Federated Unlearning: A Survey on Methods, Design Guidelines, and Evaluation Metrics
Nicolò Romandini, Alessio Mora, Carlo Mazzocca, Rebecca Montanari, and Paolo Bellavista. 2024 · 2024
Closest in time.
Rethinking LLM Memorization through the Lens of Adversarial Compression
Avi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C. Lipton, and J. Zico Kolter. 2024 · 2024
Closest in time.
Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel, and Stephan Gunnemann. 2024 · 2024
Closest in time.
Shang Shang, Xinqiang Zhao, Zhongjiang Yao, Yepeng Yao, Liya Su, Zijing Fan, Xiaodan Zhang, and Zhengwei Jiang. 2024 · 2024
Closest in time.
Rethinking Interpretability in the Era of Large Language Models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao. 2024 · 2024
Closest in time.
ChatGPT hallucinating: can it get any more humanlike?
Konstantinos C Siontis, Zachi I Attia, Samuel J Asirvatham, and Paul A Friedman. 2024 · 2024
Closest in time.
Preference Ranking Optimization for Human Alignment. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada , Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan (Eds.). 18990–18998
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2024 · 2024
Closest in time.
TITANIC: Towards Production Federated Learning with Large Language Models. In IEEE INFOCOM
Ningxin Su, Chenghao Hu, Baochun Li, and Bo Li. 2024 · 2024
Closest in time.
Exploring the Deceptive Power of LLM-Generated Fake News: A Study of Real-World Detection Challenges
Yanshen Sun, Jianfeng He, Limeng Cui, Shuo Lei, and Chang-Tien Lu. 2024a · 2024
Closest in time.
Unlink to Unlearn: Simplifying Edge Unlearning in GNNs. In Companion Proceedings of the ACM on Web Conference 2024 . 489–492
Jiajun Tan, Fei Sun, Ruichen Qiu, Du Su, and Huawei Shen. 2024a · 2024
Closest in time.
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
Zhen Tan, Chengshuai Zhao, Raha Moraffah, Yifan Li, Yu Kong, Tianlong Chen, and Huan Liu. 2024b · 2024
Closest in time.
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
Xijia Tao, Shuai Zhong, Lei Li, Qi Liu, and Lingpeng Kong. 2024b · 2024
Closest in time.
Communication Efficient and Provable Federated Unlearning
Youming Tao, Cheng-Long Wang, Miao Pan, Dongxiao Yu, Xiuzhen Cheng, and Di Wang. 2024a · 2024
Closest in time.
Challenging Systematic Prejudices: An Investigation into Bias Against Women and Girls
Daniel Van Niekerk, María Peréz-Ortiz, John Shawe-Taylor, Davor Orlic, Jackie Kay, Noah Siegel, Katherine Evans, Nyalleng Moorosi, Tina Eliassi-Rad, Leonie Maria Tanczer, et al · 2024
Closest in time.
FVFL: A Flexible and Verifiable Privacy-Preserving Federated Learning Scheme
Gang Wang, Li Zhou, Qingming Li, Xiaoran Yan, Ximeng Liu, and Yuncheng Wu. 2024e · 2024
Closest in time.
Pandora’s White-Box: Increased Training Data Leakage in Open LLMs
Jeffrey G Wang, Jason Wang, Marvin Li, and Seth Neel. 2024d · 2024
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. 2024b · 2024
Closest in time.
Server-Initiated Federated Unlearning to Eliminate Impacts of Low-Quality Data
Pengfei Wang, Wei Song, Heng Qi, Changjun Zhou, Fuliang Li, Yong Wang, Peng Sun, and Qiang Zhang. 2024c · 2024
Closest in time.
Yuhao Wang, Yusheng Liao, Heyang Liu, Hongcheng Liu, Yu Wang, and Yanfeng Wang. 2024a · 2024
Closest in time.
Gradient-Based Language Model Red Teaming. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2024 - Volume 1: Long Papers, St. Julian’s, Malta, March 17-22, 2024 , Yvette Graham and Matthew Purver (Eds.). 2862–2881
Nevan Wichers, Carson Denison, and Ahmad Beirami. 2024 · 2024
Closest in time.
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick D. McDaniel, and Chaowei Xiao. 2024b · 2024
Closest in time.
Cardinality Counting in" Alcatraz": A Privacy-aware Federated Learning Approach. In Proceedings of the ACM on Web Conference 2024 . 3076–3084
Nan Wu, Xin Yuan, Shuo Wang, Hongsheng Hu, and Minhui Xue. 2024a · 2024
Closest in time.
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
Xuansheng Wu, Haiyan Zhao, Yaochen Zhu, Yucheng Shi, Fan Yang, Tianming Liu, Xiaoming Zhai, Wenlin Yao, Jundong Li, Mengnan Du, and Ninghao Liu. 2024c · 2024
Closest in time.
Towards AI Safety: A Taxonomy for AI System Evaluation
Boming Xia, Qinghua Lu, Liming Zhu, and Zhenchang Xing. 2024 · 2024
Closest in time.
Knowledge Conflicts for LLMs: A Survey
Rongwu Xu, Zehan Qi, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. 2024a · 2024
Closest in time.
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
Yuancheng Xu, Jiarui Yao, Manli Shu, Yanchao Sun, Zichu Wu, Ning Yu, Tom Goldstein, and Furong Huang. 2024c · 2024
Closest in time.
A Survey on Robotics with Foundation Models: toward Embodied AI
Zhiyuan Xu, Kun Wu, Junjie Wen, Jinming Li, Ning Liu, Zhengping Che, and Jian Tang. 2024b · 2024
Closest in time.
AI: Unexplainable, Unpredictable, Uncontrollable
Roman V Yampolskiy. 2024 · 2024
Closest in time.
On protecting the data privacy of large language models (llms): A survey
Biwei Yan, Kun Li, Minghui Xu, Yueyan Dong, Yue Zhang, Zhaochun Ren, and Xiuzheng Cheng. 2024 · 2024
Closest in time.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024 · 2024
Closest in time.
Xiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen, Xiaofeng Mao, Longtao Huang, Hui Xue, Wenhai Wang, Kui Ren, and Jingyi Wang. 2024a · 2024
Closest in time.
Towards Efficient and Robust Federated Unlearning in IoT Networks
Yanli Yuan, BingBing Wang, Chuan Zhang, Zehui Xiong, Chunhai Li, and Liehuang Zhu. 2024b · 2024
Closest in time.
General Intelligence and Seed AI - Creating Complete Minds Capable of Open-Ended Self-Improvement
Eliezer Yudkowsky. 2024 · 2024
Closest in time.
Extracting Prompts by Inverting LLM Outputs
Collin Zhang, John X Morris, and Vitaly Shmatikov. 2024d · 2024
Closest in time.
Secure Transformer Inference Made Non-interactive
Jiawen Zhang, Jian Liu, Xinpeng Yang, Yinghao Wang, Kejia Chen, Xiaoyang Hou, Kui Ren, and Xiaohu Yang. 2024c · 2024
Closest in time.
Towards building the federatedGPT: Federated instruction tuning. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6915–6919
Jianyi Zhang, Saeed Vahidian, Martin Kuo, Chunyuan Li, Ruiyi Zhang, Tong Yu, Guoyin Wang, and Yiran Chen. 2024f · 2024
Closest in time.
Device Scheduling and Assignment in Hierarchical Federated Learning for Internet of Things
Tinghao Zhang, Kwok-Yan Lam, and Jun Zhao. 2024b · 2024
Closest in time.
Intention Analysis Prompting Makes Large Language Models A Good Jailbreak Defender
Yuqi Zhang, Liang Ding, Lefei Zhang, and Dacheng Tao. 2024a · 2024
Closest in time.
Privacy-preserving and fairness-aware federated learning for critical infrastructure protection and resilience. In Proceedings of the ACM on Web Conference 2024 . 2986–2997
Yanjun Zhang, Ruoxi Sun, Liyue Shen, Guangdong Bai, Minhui Xue, Mark Huasong Meng, Xue Li, Ryan Ko, and Surya Nepal. 2024e · 2024
Closest in time.
Zaibin Zhang, Yongting Zhang, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, and Jing Shao. 2024h · 2024
Closest in time.
Static and Sequential Malicious Attacks in the Context of Selective Forgetting
Chenxu Zhao, Wei Qian, Rex Ying, and Mengdi Huai. 2024a · 2024
Closest in time.
Opening the Black Box of Large Language Models: Two Views on Holistic Interpretability
Haiyan Zhao, Fan Yang, Himabindu Lakkaraju, and Mengnan Du. 2024d · 2024
Closest in time.
LLM-based Federated Recommendation
Jujia Zhao, Wenjie Wang, Chen Xu, Zhaochun Ren, See-Kiong Ng, and Tat-Seng Chua. 2024b · 2024
Closest in time.
The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
Qinyu Zhao, Ming Xu, Kartik Gupta, Akshay Asthana, Liang Zheng, and Stephen Gould. 2024c · 2024
Closest in time.
PermLLM: Private Inference of Large Language Models within 3 Seconds under WAN
Fei Zheng, Chaochao Chen, Zhongxuan Han, and Xiaolin Zheng. 2024 · 2024
Closest in time.
NetGuard: Protecting commercial web APIs from model inversion attacks using GAN-generated fake samples. In Proceedings of the ACM Web Conference . 2045–2053
Xueluan Gong, Ziyao Wang, Yanjiao Chen, Qian Wang, Cong Wang, and Chao Shen. 2023c · 2053
Closest in time.