Fetching the paper…
Reading the bibliography…
This paper introduces LlavaGuard, a suite of VLM-based vision safeguards that address the critical need for reliable guardrails in the era of large-scale data and models.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
The socio-moral image database (smid): A novel stimulus set for the study of social, moral and affective processes
Crone, D. L., Bode, S., Murawski, C., and Laham, S. M · 2018
Earlier work this paper cites.
Deep neural network for nsfw detection, 2020
Laborde, G · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
Large image datasets: A pyrrhic win for computer vision?
Birhane, A. and Prabhu, V. U · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes, 2021
Birhane, A., Prabhu, V. U., and Kahembwe, E · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Datasheets for datasets
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., III, H. D., and Crawford, K · 2021
Earlier work this paper cites.
Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation
Karkkainen, K. and Joo, J · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models, 2021
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W., Legassick, S., Irving, G., and Gabriel, I · 2021
Earlier work this paper cites.
GLIDE: towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Can machines help us answering question 16 in datasheets, and in turn reflecting on inappropriate content?
Schramowski, P., Tauchmann, C., and Kersting, K · 2022
Earlier work this paper cites.
LAION-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C. W., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S. R., Crowson, K., Schmidt, L., Kaczmarczyk, R., and Jitsev, J · 2022
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Easily accessible text-to-image generation amplifies demographic stereotypes at large scale
Bianchi, F., Kalluri, P., Durmus, E., Ladhak, F., Cheng, M., Nozza, D., Hashimoto, T., Jurafsky, D., Zou, J., and Caliskan, A · 2023
Cited alongside, same era.
Into the LAION’s den: Investigating hate in multimodal datasets
Birhane, A., vinay uday prabhu, Han, S., Boddeti, V., and Luccioni, S · 2023
Cited alongside, same era.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models
Cho, J., Zala, A., and Bansal, M · 2023
Cited alongside, same era.
An overview of catastrophic ai risks, 2023
Hendrycks, D., Mazeika, M., and Woodside, T · 2023
Cited alongside, same era.
An empirical study of metrics to measure representational harms in pre-trained language models
Hosseini, S., Palangi, H., and Awadallah, A. H · 2023
Cited alongside, same era.
Nsfw image detection
Falconsai · 2024
Closest in time.
NudeNet: Nudity Detection with Deep Learning
NotAI-tech · 2024
Closest in time.
Moderation guide
OpenAI · 2024
Closest in time.
Unsafebench: Benchmarking image safety classifiers on real-world and ai-generated images, 2024
Qu, Y., Shen, X., Wu, Y., Backes, M., Zannettou, S., and Zhang, Y · 2024
Closest in time.
Nsfw filter
Sanali209 · 2024
Closest in time.
Alert: A comprehensive benchmark for assessing large language models’ safety through red teaming, 2024
Tedeschi, S., Friedrich, F., Schramowski, P., Kersting, K., Navigli, R., Nguyen, H., and Li, B · 2024
Closest in time.
AI regulation: A pro-innovation approach
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M · 2023
Cited alongside, same era.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation, 2023
Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J · 2023
Cited alongside, same era.
Amplifying limitations, harms and risks of large language models
O’Neill, M. and Connor, M · 2023
Cited alongside, same era.
Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models
Schramowski, P., Brack, M., Deiseroth, B., and Kersting, K · 2023
Cited alongside, same era.
Identifying and eliminating csam in generative ml training data and models, 2023
Thiel, D · 2023
Cited alongside, same era.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., Truong, S. T., Arora, S., Mazeika, M., Hendrycks, D., Lin, Z., Cheng, Y., Koyejo, S., Song, D., and Li, B · 2023
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., et al · 2024
Cited alongside, same era.
UK · 2024
Closest in time.
Fact sheet: President biden issues executive order on safe, secure, and trustworthy artificial intelligence
US · 2024
Closest in time.
Introducing v0.5 of the ai safety benchmark from mlcommons, 2024
Vidgen, B., Agrawal, A., Ahmed, A. M., Akinwande, V., Al-Nuaimi, N., Alfaraj, N., Alhajjar, E., Aroyo, L., Bavalatti, T., Blili-Hamelin, B., Bollacker, K., Bomassani, R., Boston, M. F., Campos, S., Chakra, K., Chen, C., Coleman, C., Coudert, Z. D., Derczynski, L., Dutta, D., Eisenberg, I., Ezick, J., Frase, H., Fuller, B., Gandikota, R., Gangavarapu, A., Gangavarapu, A., Gealy, J., Ghosh, R., Goel, J., Gohar, U., Goswami, S., Hale, S. A., Hutiri, W., Imperial, J. M., Jandial, S., Judd, N., Juefei-Xu, F., Khomh, F., Kailkhura, B., Kirk, H. R., Klyman, K., Knotz, C., Kuchnik, M., Kumar, S. H., Lengerich, C., Li, B., Liao, Z., Long, E. P., Lu, V., Mai, Y., Mammen, P. M., Manyeki, K., McGregor, S., Mehta, V., Mohammed, S., Moss, E., Nachman, L., Naganna, D. J., Nikanjam, A., Nushi, B., Oala, L., Orr, I., Parrish, A., Patlak, C., Pietri, W., Poursabzi-Sangdeh, F., Presani, E., Puletti, F., Röttger, P., Sahay, S., Santos, T., Scherrer, N., Sebag, A. S., Schramowski, P., Shahbazi, A., Sharma, V., Shen, X., Sistla, V., Tang, L., Testuggine, D., Thangarasa, V., Watkins, E. A., Weiss, R., Welty, C., Wilbers, T., Williams, A., Wu, C.-J., Yadav, P., Yang, X., Zeng, Y., Zhang, W., Zhdanov, F., Zhu, J., Liang, P., Mattson, P., and Vanschoren, J · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., Fan, Y., Dang, K., Du, M., Ren, X., Men, R., Liu, D., Zhou, C., Zhou, J., and Lin, J · 2024
Closest in time.
Ai risk categorization decoded (air 2024): From government regulations to corporate policies, 2024
Zeng, Y., Klyman, K., Zhou, A., Yang, Y., Pan, M., Jia, R., Song, D., Liang, P., and Li, B · 2024
Closest in time.
Stylebreeder: Exploring and democratizing artistic styles through text-to-image models, 2024
Zheng, M., Simsar, E., Yesiltepe, H., Tombari, F., Simon, J., and Yanardag, P · 2024
Closest in time.
T2ISafety: Benchmark for assessing fairness, toxicity, and privacy in image generation, 2025
Li, L., Shi, Z., Hu, X., Dong, B., Qin, Y., Liu, X., Sheng, L., and Shao, J · 2025
Closest in time.
Tschannen, M., Gritsenko, A., Wang, X., Naeem, M. F., Alabdulmohsin, I., Parthasarathy, N., Evans, T., Beyer, L., Xia, Y., Mustafa, B., Hénaff, O., Harmsen, J., Steiner, A., and Zhai, X · 2025
Closest in time.