Fetching the paper…
Reading the bibliography…
We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2.
Perspective api
Google · 2017
Earlier work this paper cites.
Counterfactual fairness
M. J. Kusner, J. Loftus, C. Russell, and R. Silva · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
O. Sener and S. Savarese · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Mix-and-match tuning for self-supervised semantic segmentation
X. Zhan, Z. Liu, P. Luo, X. Tang, and C. Loy · 2018
Earlier work this paper cites.
Deep batch active learning by diverse, uncertain gradient lower bounds
J. T. Ash, C. Zhang, A. Krishnamurthy, J. Langford, and A. Agarwal · 2019
Earlier work this paper cites.
Batch active learning at scale
G. Citovsky, G. DeSalvo, C. Gentile, L. Karydas, A. Rajagopalan, A. Rostamizadeh, and S. Kumar · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Self-guided noise-free data generation for efficient zero-shot learning
J. Gao, R. Pi, Y. Lin, H. Xu, J. Ye, Z. Wu, W. Zhang, X. Liang, Z. Li, and L. Kong · 2022
Earlier work this paper cites.
Ask me what you need: Product retrieval using knowledge from gpt-3
S. Y. Kim, H. Park, K. Shin, and K.-M. Kim · 2022
Earlier work this paper cites.
Data augmentation for intent classification with off-the-shelf large language models
G. Sahu, P. Rodriguez, I. H. Laradji, P. Atighehchian, D. Vazquez, and D. Bahdanau · 2022
Earlier work this paper cites.
" i’m sorry to hear that": Finding new biases in language models with a holistic descriptor dataset
E. M. Smith, M. Hall, M. Kambadur, E. Presani, and A. Williams · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Rethinking conversational agents in the era of llms: Proactivity, non-collaborativity, and beyond
Y. Deng, W. Lei, M. Huang, and T.-S. Chua · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, et al · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset, 2023
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, C. Zhang, R. Sun, Y. Wang, and Y. Yang · 2023
Cited alongside, same era.
Harnessing large-language models to generate private synthetic text
Humans or llms as the judge? a study on judgement biases
G. H. Chen, S. Chen, Z. Liu, F. Jiang, and B. Wang · 2024
Closest in time.
Aegis: Online adaptive ai content safety moderation with ensemble of llm experts
S. Ghosh, P. Varshney, E. Galinkin, and C. Parisien · 2024
Closest in time.
Responsible generative ai toolkit: https://ai.google.dev/responsible/principles, 2024
Google · 2024
Closest in time.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
S. Han, K. Rao, A. Ettinger, L. Jiang, B. Y. Lin, N. Lambert, Y. Choi, and N. Dziri · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Kurakin, N. Ponomareva, U. Syed, L. MacDermed, and A. Terzis · 2023
Cited alongside, same era.
Z. Lin, Z. Wang, Y. Tong, Y. Wang, Y. Guo, Y. Wang, and J. Shang · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
T. Markov, C. Zhang, S. Agarwal, F. E. Nekoul, T. Lee, S. Adler, A. Jiang, and L. Weng · 2023
Cited alongside, same era.
Scalable extraction of training data from (production) language models
M. Nasr, N. Carlini, J. Hayase, M. Jagielski, A. F. Cooper, D. Ippolito, C. A. Choquette-Choo, E. Wallace, F. Tramèr, and K. Lee · 2023
Cited alongside, same era.
Aart: Ai-assisted red-teaming with diverse data generation for new llm-powered applications
B. Radharapu, K. Robinson, L. Aroyo, and P. Lahoti · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
A. Anthropic · 2024
Cited alongside, same era.
Meta llama guard 2
L. Team
Cited in the paper.
H. Huang, Y. Qu, J. Liu, M. Yang, and T. Zhao · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
L. Li, B. Dong, R. Wang, X. Hu, W. Zuo, D. Lin, Y. Qiao, and J. Shao · 2024
Closest in time.
N. Liu, L. Chen, X. Tian, W. Zou, K. Chen, and M. Cui · 2024
Closest in time.
On llms-driven synthetic data generation, curation, and evaluation: A survey
L. Long, R. Wang, R. Xiao, J. Zhao, X. Ding, G. Chen, and H. Wang · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.