Fetching the paper…
Reading the bibliography…
Robust content moderation classifiers are essential for the safety of Generative AI systems.
Unsupervised cross-lingual representation learning at scale, 2020
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 1911
Earlier work this paper cites.
Detection of harassment on web 2.0
Dawei Yin, Zhenzhen Xue, Liangjie Hong, Brian Davison, April Edwards, and Lynne Edwards · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
Abusive language detection in online user content
Chikashi Nobata, Joel Tetreault, Achint Thomas, Yashar Mehdad, and Yi Chang · 2016
Earlier work this paper cites.
Toxic comment classification challenge, 2017
C.J. Adams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, Nithum, and Will Cukierski · 2017
Earlier work this paper cites.
Using convolutional neural networks to classify hate-speech
Björn Gambäck and Utpal Kumar Sikdar · 2017
Earlier work this paper cites.
Refining word embeddings for sentiment analysis
Liang-Chih Yu, Jin Wang, K. Robert Lai, and Xuejie Zhang · 2017
Earlier work this paper cites.
Community standards report, 2019
Mike Schroepfer · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2020
Earlier work this paper cites.
The shift to generalized ai to better identify violating content, 2021
Meta · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation, 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Earlier work this paper cites.
Precise zero-shot dense retrieval without relevance labels, 2022
Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan · 2022
Earlier work this paper cites.
Augly: Data augmentations for robustness, 2022
Zoe Papakipos and Joanna Bitton · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Claude, 2023
Anthropic · 2023
Cited alongside, same era.
Emu: Enhancing image generation models using photogenic needles in a haystack, 2023
Xiaoliang Dai, Ji Hou, Chih-Yao Ma, Sam Tsai, Jialiang Wang, Rui Wang, Peizhao Zhang, Simon Vandenhende, Xiaofang Wang, Abhimanyu Dubey, Matthew Yu, Abhishek Kadian, Filip Radenovic, Dhruv Mahajan, Kunpeng Li, Yue Zhao, Vladan Petrovic, Mitesh Kumar Singh, Simran Motwani, Yi Wen, Yiwen Song, Roshan Sumbaly, Vignesh Ramanathan, Zijian He, Peter Vajda, and Devi Parikh · 2023
Cited alongside, same era.
How AI is being abused to create child sexual abuse imagery
Internet Watch Foundation · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou · 2024
Closest in time.
The llama 3 herd of models, 2024
Abhimanyu Dubey · 2024
Closest in time.
Retrieval-augmented generation for large language models: A survey, 2024
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang · 2024
Closest in time.
Content moderation by llm: From accuracy to legitimacy, 2024
Tao Huang · 2024
Closest in time.
Latent guard: a safety framework for text-to-image generation, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Cited alongside, same era.
How to train your dragon: Diverse augmentation towards generalizable dense retrieval, 2023
Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz, Jimmy Lin, Yashar Mehdad, Wen tau Yih, and Xilun Chen · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Cited alongside, same era.
Chatgpt, 2023
OpenAI · 2023
Cited alongside, same era.
Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, and Yang Zhang · 2023
Cited alongside, same era.
Performance and risk trade-offs for multi-word text prediction at scale
Aniket Vashishtha, Shirtika S. Prasad, Payal Bajaj, Vishrav Chaudhary, Kate Cook, Sandipan Dandapat, Sunayana Sitaram, and Monojit Choudhury · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Cited alongside, same era.
Generate rather than retrieve: Large language models are strong context generators, 2023
Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang · 2023
Cited alongside, same era.
Runtao Liu, Ashkan Khakzar, Jindong Gu, Qifeng Chen, Philip Torr, and Fabio Pizzati · 2024
Closest in time.
Movie gen: A cast of media foundation models, 2024
Meta · 2024
Closest in time.
DALL-E 3 system card
OpenAI · 2024
Closest in time.
Unsafebench: Benchmarking image safety classifiers on real-world and ai-generated images, 2024
Yiting Qu, Xinyue Shen, Yixin Wu, Michael Backes, Savvas Zannettou, and Yang Zhang · 2024
Closest in time.
The language barrier: Dissecting safety challenges of LLMs in multilingual contexts
Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi · 2024
Closest in time.
Ai risk categorization decoded (air 2024): From government regulations to corporate policies, 2024
Yi Zeng, Kevin Klyman, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, and Bo Li · 2024
Closest in time.
Raft: Adapting language model to domain specific rag, 2024
Tianjun Zhang, Shishir G. Patil, Naman Jain, Sheng Shen, Matei Zaharia, Ion Stoica, and Joseph E. Gonzalez · 2024
Closest in time.
Retrieval-augmented generation for ai-generated content: A survey, 2024
Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, Jie Jiang, and Bin Cui · 2024
Closest in time.