Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are swiftly advancing in architecture and capability, and as they integrate more deeply into complex systems, the urgency to scrutinize their security properties grows.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020a · 1901
Earlier work this paper cites.
Robust neural machine translation with doubly adversarial inputs
Yong Cheng, Lu Jiang, and Wolfgang Macherey. 2019 · 1906
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 1908
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 1908
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019a · 1908
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Foundations of statistical natural language processing
Christopher Manning and Hinrich Schutze. 1999 · 1999
Earlier work this paper cites.
A multi-agent system for natural language understanding
Mostafa M Aref. 2003 · 2003
Earlier work this paper cites.
Undersensitivity in neural reading comprehension
Johannes Welbl, Pasquale Minervini, Max Bartolo, Pontus Stenetorp, and Sebastian Riedel. 2020 · 2003
Earlier work this paper cites.
Sqlrand: Preventing sql injection attacks
Stephen W Boyd and Angelos D Keromytis. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020b · 2005
Earlier work this paper cites.
A classification of sql-injection attacks and countermeasures
William G. J. Halfond, Jeremy Viegas, and Alessandro Orso. 2006 · 2006
Earlier work this paper cites.
The radicalization risks of gpt-3 and advanced neural language models
Kris McGuffie and Alex Newhouse. 2020 · 2009
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020b · 2010
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Šrndić, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. 2013 · 2013
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013 · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Proceedings of the 2014 conference on empirical methods in natural language processing (emnlp)
Alessandro Moschitti, Bo Pang, and Walter Daelemans. 2014 · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016 · 2016
Earlier work this paper cites.
Defensive distillation is not robust to adversarial examples
Nicholas Carlini and David Wagner. 2016 · 2016
Earlier work this paper cites.
Adversarial machine learning at scale
Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2016 · 2016
Earlier work this paper cites.
Distillation as a defense to adversarial perturbations against deep neural networks
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. 2016 · 2016
Earlier work this paper cites.
Machine learning with adversaries: Byzantine tolerant gradient descent
Peva Blanchard, El Mahdi El Mhamdi, Rachid Guerraoui, and Julien Stainer. 2017 · 2017
Earlier work this paper cites.
Distributed statistical machine learning in adversarial settings: Byzantine gradient descent
Yudong Chen, Lili Su, and Jiaming Xu. 2017 · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2017 · 2017
Earlier work this paper cites.
Deceiving google’s perspective api built for detecting toxic comments
Hossein Hosseini, Sreeram Kannan, Baosen Zhang, and Radha Poovendran. 2017 · 2017
Earlier work this paper cites.
Adversarial attacks on neural network policies
Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Reluplex: An efficient smt solver for verifying deep neural networks
Guy Katz, Clark Barrett, David L Dill, Kyle Julian, and Mykel J Kochenderfer. 2017 · 2017
Earlier work this paper cites.
Rhmd: Evasion-resilient hardware malware detectors
Khaled N Khasawneh, Nael Abu-Ghazaleh, Dmitry Ponomarev, and Lei Yu. 2017 · 2017
Earlier work this paper cites.
Deep text classification can be fooled
Bin Liang, Hongcheng Li, Miaoqiang Su, Pan Bian, Xirong Li, and Wenchang Shi. 2017 · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017 · 2017
Earlier work this paper cites.
Towards crafting text adversarial samples
Suranjana Samanta and Sameep Mehta. 2017 · 2017
Earlier work this paper cites.
Certified defenses for data poisoning attacks
Jacob Steinhardt, Pang Wei W Koh, and Percy S Liang. 2017 · 2017
Earlier work this paper cites.
Feature squeezing: Detecting adversarial examples in deep neural networks
Weilin Xu, David Evans, and Yanjun Qi. 2017 · 2017
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Earlier work this paper cites.
Boosting adversarial attacks with momentum
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. 2018 · 2018
Earlier work this paper cites.
Black-box generation of adversarial text sequences to evade deep learning classifiers
Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018 · 2018
Earlier work this paper cites.
Cnn-based projected gradient descent for consistent ct image reconstruction
Harshit Gupta, Kyong Hwan Jin, Ha Q Nguyen, Michael T McCann, and Michael Unser. 2018 · 2018
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Adversarial malware binaries: Evading deep learning for malware detection in executables
Bojan Kolosnjaji, Ambra Demontis, Battista Biggio, Davide Maiorca, Giorgio Giacinto, Claudia Eckert, and Fabio Roli. 2018 · 2018
Earlier work this paper cites.
Adversarial examples for natural language classification problems
Volodymyr Kuleshov, Shantanu Thakoor, Tingfung Lau, and Stefano Ermon. 2018 · 2018
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2018 · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
Robust machine comprehension models via adversarial training
Yicheng Wang and Mohit Bansal. 2018 · 2018
Earlier work this paper cites.
Provable defenses against adversarial examples via the convex outer adversarial polytope
Eric Wong and Zico Kolter. 2018 · 2018
Earlier work this paper cites.
Generating natural adversarial examples
Zhengli Zhao, Dheeru Dua, and Sameer Singh. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019b · 2019
Earlier work this paper cites.
A repository of conversational datasets
Matthew Henderson, Paweł Budzianowski, Iñigo Casanueva, Sam Coope, Daniela Gerz, Girish Kumar, Nikola Mrkšić, Georgios Spithourakis, Pei-Hao Su, Ivan Vulić, et al. 2019 · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. 2019 · 2019
Earlier work this paper cites.
Certified robustness to adversarial examples with differential privacy
Mathias Lecuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu, and Suman Jana. 2019 · 2019
Earlier work this paper cites.
TextBugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2019 · 2019
Earlier work this paper cites.
Sensitivity of adversarial perturbation in fast gradient sign method
Yujie Liu, Shuai Mao, Xiang Mei, Tao Yang, and Xuran Zhao. 2019b · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Do NLP models know numbers? probing numeracy in embeddings
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019b · 2019
Earlier work this paper cites.
Untargeted adversarial attack via expanding the semantic gap
Aming Wu, Yahong Han, Quanxin Zhang, and Xiaohui Kuang. 2019 · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Earlier work this paper cites.
Taamr: Targeted adversarial attack against multimedia recommender systems
Tommaso Di Noia, Daniele Malitesta, and Felice Antonio Merra. 2020 · 2020
Earlier work this paper cites.
Local model poisoning attacks to { \{ Byzantine-Robust } \} federated learning
Minghong Fang, Xiaoyu Cao, Jinyuan Jia, and Neil Gong. 2020 · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020 · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Earlier work this paper cites.
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020 · 2020
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Earlier work this paper cites.
Privacy risks of general-purpose language models
Xudong Pan, Mi Zhang, Shouling Ji, and Min Yang. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Universal adversarial training
Ali Shafahi, Mahyar Najibi, Zheng Xu, John Dickerson, Larry S Davis, and Tom Goldstein. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Earlier work this paper cites.
On adaptive attacks to adversarial example defenses
Florian Tramer, Nicholas Carlini, Wieland Brendel, and Aleksander Madry. 2020 · 2020
Earlier work this paper cites.
Random erasing data augmentation
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. 2020 · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Earlier work this paper cites.
A survey on adversarial attacks and defences
Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harri Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Cited alongside, same era.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. 2021 · 2021
Cited alongside, same era.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. 2021 · 2021
Cited alongside, same era.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
The intelligent agent nlp-based customer service system
Changran Huang. 2021 · 2021
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. 2023 · 2023
Closest in time.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. 2023 · 2023
Closest in time.
Master thread of ways i have discovered to get chatgpt to output text that it’s not supposed to, including bigotry, urls and personal information, and more
Colin Fraser. 2023 · 2023
Closest in time.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Cited alongside, same era.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021 · 2021
Cited alongside, same era.
Reading isn’t believing: Adversarial attacks on multi-modal neurons
David A Noever and Samantha E Miller Noever. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021 · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021 · 2021
Cited alongside, same era.
Understanding the capabilities, limitations, and societal impact of large language models
Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli. 2021 · 2021
Cited alongside, same era.
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al. 2023 · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023 · 2023
Closest in time.
Multimodal-gpt: A vision and language model for dialogue with humans
Tao Gong, Chengqi Lyu, Shilong Zhang, Yudong Wang, Miao Zheng, Qian Zhao, Kuikun Liu, Wenwei Zhang, Ping Luo, and Kai Chen. 2023 · 2023
Closest in time.
https://blog.google/technology/ai/google-bard-updates-io-2023/
Google-Bard · 2023
Closest in time.
Indirect prompt injection threats
Kai Greshakeblog. 2023 · 2023
Closest in time.
Adversarial prompting guide
injection Guide. 2023 · 2023
Closest in time.
A two sentence jailbreak for gpt-4 and claude and why nobody knows how to fix it
Alexey Guzey. 2023 · 2023
Closest in time.
Sydney bing chat
Marvin von Hagen. 2023 · 2023
Closest in time.
Fedmlsecurity: A benchmark for attacks and defenses in federated learning and llms
Shanshan Han, Baturalp Buyukates, Zijian Hu, Han Jin, Weizhao Jin, Lichao Sun, Xiaoyang Wang, Chulin Xie, Kai Zhang, Qifan Zhang, et al. 2023 · 2023
Closest in time.
Tabllm: Few-shot classification of tabular data with large language models
Stefan Hegselmann, Alejandro Buendia, Hunter Lang, Monica Agrawal, Xiaoyi Jiang, and David Sontag. 2023 · 2023
Closest in time.
Llm self defense: By self examination, llms know they are being tricked
Alec Helbling, Mansi Phute, Matthew Hull, and Duen Horng Chau. 2023 · 2023
Closest in time.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Closest in time.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Lerf: Language embedded radiance fields
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. 2023 · 2023
Closest in time.
Adversarial attacks on tables with entity swap
Aneta Koleva, Martin Ringsquandl, and Volker Tresp. 2023 · 2023
Closest in time.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. 2023 · 2023
Closest in time.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2023 · 2023
Closest in time.
Google bard jailbreak: Prompt to bard jailbreak
Akash Kushwaha. 2023 · 2023
Closest in time.
Lakera prompt injection challenge
Gandalf Lakera. 2023 · 2023
Closest in time.
Huggingface h4 stack exchange preference dataset
Nathan Lambert, Lewis Tunstall, Nazneen Rajani, and Tristan Thrush. 2023 · 2023
Closest in time.
Langchain prompt injection webinar
PI LangchainWebinar. 2023 · 2023
Closest in time.
Analyzing leakage of personally identifiable information in language models
Nils Lukas, Ahmed Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-Béguelin. 2023 · 2023
Closest in time.
Deepmem: Ml models as storage channels and their (mis-) applications
Md Abdullah Al Mamun, Quazi Mishkatul Alam, Erfan Shaigani, Pedram Zaree, Ihsen Alouani, and Nael Abu-Ghazaleh. 2023 · 2023
Closest in time.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023 · 2023
Closest in time.
Inverse scaling: When bigger isn’t better
Ian R McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish, Aaron Mueller, Ameya Prabhu, Euan McLean, Aaron Kirtland, Alexis Ross, Alisa Liu, et al. 2023 · 2023
Closest in time.
https://blogs.bing.com/search/july-2023/Bing-Chat-Enterprise-announced,-multimodal-Visual-Search-rolling-out-to-Bing-Chat
Microsoft-Bing · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms.”
NLPTeam MosaicML. 2023 · 2023
Closest in time.
Use of llms for illicit purposes: Threats, prevention measures, and vulnerabilities
Maximilian Mozes, Xuanli He, Bennett Kleinberg, and Lewis D Griffin. 2023 · 2023
Closest in time.
A robust analysis of adversarial attacks on federated learning environments
Akarsh K Nair, Ebin Deni Raj, and Jayakrushna Sahoo. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Openai - explore what’s possible with some example applications
AI OpenAIApplications. 2023 · 2023
Closest in time.
The prompt engineering platform to experiment with different prompt versions
PI Parea. 2023 · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023 · 2023
Closest in time.
Rodrigo Pedro, Daniel Castro, Paulo Carreira, and Nuno Santos. 2023 · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Closest in time.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023 · 2023
Closest in time.
Aman Priyanshu, Supriti Vijay, Ayush Kumar, Rakshit Naidu, and Fatemehsadat Mireshghallah. 2023 · 2023
Closest in time.
Midjourney, chatgpt, dall·e, stable diffusion and more prompt marketplace
buysell PromptBase. 2023 · 2023
Closest in time.
Visual adversarial examples jailbreak large language models
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Mengdi Wang, and Prateek Mittal. 2023 · 2023
Closest in time.
Huachuan Qiu, Shuai Zhang, Anqi Li, Hongliang He, and Zhenzhong Lan. 2023 · 2023
Closest in time.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury. 2023 · 2023
Closest in time.
Image to prompt injection with google bard
Johann Rehberger. 2023 · 2023
Closest in time.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al. 2023 · 2023
Closest in time.
Bushra Sabir, M Ali Babar, and Sharif Abuadbba. 2023 · 2023
Closest in time.
New prompt injection attack on chatgpt web version. markdown images can steal your chat data
Roman Samoilenko a. 2023 · 2023
Closest in time.
New prompt injection attack on chatgpt web version. reckless copy-pasting may lead to serious privacy issues in your chat
Roman Samoilenko b. 2023 · 2023
Closest in time.
Training language models with language feedback at scale
Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez. 2023 · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Closest in time.
On the adversarial robustness of multi-modal foundation models
Christian Schlarmann and Matthias Hein. 2023 · 2023
Closest in time.
Prompt injection cheat sheet: How to manipulate ai language models
staff Seclify. 2023 · 2023
Closest in time.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Dhruv Shah, Błażej Osiński, Sergey Levine, et al. 2023 · 2023
Closest in time.
Role-play with large language models
Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023 · 2023
Closest in time.
Plug and pray: Exploiting off-the-shelf components of multi-modal models
Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. 2023 · 2023
Closest in time.
Towards expert-level medical question answering with large language models
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, et al. 2023 · 2023
Closest in time.
That doesn’t go there: Attacks on shared state in multi-user augmented reality applications
Carter Slocum, Yicheng Zhang, Erfan Shayegani, Pedram Zaree, Nael Abu-Ghazaleh, and Jiasi Chen. 2023 · 2023
Closest in time.
Pandagpt: One model to instruction-follow them all
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. 2023 · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick. 2023 · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
Creating large language model applications utilizing langchain: A primer on developing llm apps fast
Oguzhan Topsakal and Tahir Cetin Akinci. 2023 · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua. 2023 · 2023
Closest in time.
Writesonic - an ai-powered writing tool
PI Writesonic. 2023 · 2023
Closest in time.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023 · 2023
Closest in time.
Ai injections: Direct and indirect prompt injections and their implications
Red Wunderwuzzi. 2023 · 2023
Closest in time.
Multimodal learning with transformers: A survey
Peng Xu, Xiatian Zhu, and David A Clifton. 2023 · 2023
Closest in time.
Instruction in the wild: A user-based instruction dataset
F Xue, Z Zheng, and Y You. 2023 · 2023
Closest in time.
Virtual prompt injection for instruction-tuned large language models
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin. 2023 · 2023
Closest in time.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2023 · 2023
Closest in time.
Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success
Yiming Zhang and Daphne Ippolito. 2023 · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023 · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Closest in time.
Are large pre-trained language models leaking your personal information?
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. 2022 · 2047
Closest in time.