Fetching the paper…
Reading the bibliography…
As Large Language Models (LLMs) become more integrated into our daily lives, it is crucial to identify and mitigate their risks, especially when the risks can have profound impacts on human users and societies.
Studies in the quality of life: Delivered by the institute of personnel management in november 1957
Trist, E. L. and Bamforth, K. W · 1957
Earlier work this paper cites.
Social psychology of creativity: A consensual assessment technique
Amabile, T. M · 1982
Earlier work this paper cites.
Safari: Versatile and efficient evaluations for robustness of interpretability
Huang, W., Zhao, X., Jin, G., and Huang, X · 1998
Earlier work this paper cites.
Managing conflicts in goal-driven requirements engineering
van Lamsweerde, A., Darimont, R., and Letier, E · 1998
Earlier work this paper cites.
Pareto multi objective optimization
Ngatchou, P., Zarei, A., and El-Sharkawi, A · 2005
Earlier work this paper cites.
Jr., D. M., Prabhakaran, V., Kuhlberg, J., Smart, A., and Isaac, W. S · 2006
Earlier work this paper cites.
A tutorial on conformal prediction
Shafer, G. and Vovk, V · 2008
Earlier work this paper cites.
The chronic care model and diabetes management in us primary care settings: A systematic review
Crabtree, B. F., Miller, W. L., and Stange, K. C · 2011
Earlier work this paper cites.
The behaviour change wheel: a new method for characterising and designing behaviour change interventions
Michie, S., Van Stralen, M. M., and West, R · 2011
Earlier work this paper cites.
Pointer networks
Vinyals, O., Fortunato, M., and Jaitly, N · 2015
Earlier work this paper cites.
Deep learning with differential privacy
Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L · 2016
Earlier work this paper cites.
Adversarial training methods for semi-supervised text classification
Miyato, T., Dai, A. M., and Goodfellow, I · 2016
Earlier work this paper cites.
Whole-system approaches to improving the health and wellbeing of healthcare workers: A systematic review
Brand, S., Thompson Coon, J., Fleming, L., Carroll, L., Bethel, A., and Wyatt, K · 2017
Earlier work this paper cites.
Safety verification of deep neural networks
Huang, X., Kwiatkowska, M., Wang, S., and Wu, M · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Li, J., Ji, S., Du, T., Li, B., and Wang, T · 2018
Earlier work this paper cites.
Reachability analysis of deep neural networks with provable guarantees
Ruan, W., Huang, X., and Kwiatkowska, M · 2018
Earlier work this paper cites.
Feature-guided black-box safety testing of deep neural networks
Wicker, M., Huang, X., and Kwiatkowska, M · 2018
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Cohen, J., Rosenfeld, E., and Kolter, Z · 2019
Earlier work this paper cites.
Privacy risks of securing machine learning models against adversarial examples
Song, L., Shokri, R., and Mittal, P · 2019
Earlier work this paper cites.
Structural test coverage criteria for deep neural networks
Sun, Y., Huang, X., Kroening, D., Sharp, J., Hill, M., and Ashmore, R · 2019
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Adversarial training for large neural language models
Liu, X., Cheng, H., He, P., Chen, W., Wang, Y., Poon, H., and Gao, J · 2020
Earlier work this paper cites.
On adaptive attacks to adversarial example defenses
Tramer, F., Carlini, N., Brendel, W., and Madry, A · 2020
Earlier work this paper cites.
Analyzing information leakage of updates to natural language models
Zanella-Béguelin, S., Wutschitz, L., Tople, S., Rühle, V., Paverd, A., Ohrimenko, O., Köpf, B., and Brockschmidt, M · 2020
Earlier work this paper cites.
The secret revealer: Generative model-inversion attacks against deep neural networks
Zhang, Y., Jia, R., Pei, H., Wang, W., Li, B., and Song, D · 2020
Earlier work this paper cites.
A safety framework for critical systems utilising deep neural networks
Zhao, X., Banks, A., Sharp, J., Robu, V., Flynn, D., Fisher, M., and Huang, X · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
Graph neural networks meet neural-symbolic computing: a survey and perspective
Lamb, L. C., d’Avila Garcez, A., Gori, M., Prates, M. O., Avelar, P. H., and Vardi, M. Y · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Scaling language model training to a trillion parameters using megatron, 2021
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., and Catanzaro, B · 2021
Earlier work this paper cites.
Variational model inversion attacks
Wang, K.-C., FU, Y., Li, K., Khisti, A. J., Zemel, R., and Makhzani, A · 2021
Earlier work this paper cites.
Challenges in detoxifying language models
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S · 2021
Earlier work this paper cites.
Reconstructing training data with informed adversaries
Balle, B., Cherubin, G., and Hayes, J · 2022
Earlier work this paper cites.
A sociotechnical view of algorithmic fairness
Dolata, M., Feuerriegel, S., and Schwabe, G · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Earlier work this paper cites.
Rethinking with retrieval: Faithful large language model inference
He, H., Zhang, H., and Roth, D · 2022
Earlier work this paper cites.
Are large pre-trained language models leaking your personal information?
Huang, J., Shao, H., and Chang, K. C.-C · 2022
Earlier work this paper cites.
All the news that’s fit to fabricate: Ai-generated text as a tool of media misinformation
Kreps, S., McCain, R. M., and Brundage, M · 2022
Earlier work this paper cites.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Kuhn, L., Gal, Y., and Farquhar, S · 2022
Earlier work this paper cites.
Large language models can be strong differentially private learners
Li, X., Tramer, F., Liang, P., and Hashimoto, T · 2022
Earlier work this paper cites.
Differentially private model compression
Mireshghallah, F., Backurs, A., Inan, H. A., Wutschitz, L., and Kulkarni, J · 2022
Earlier work this paper cites.
A survey of machine unlearning, 2022
Nguyen, T. T., Huynh, T. T., Nguyen, P. L., Liew, A. W.-C., Yin, H., and Nguyen, Q. V. H · 2022
Earlier work this paper cites.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
You are what you write: Preserving privacy in the era of large language models
Plant, R., Giuffrida, V., and Gkatzia, D · 2022
Earlier work this paper cites.
Measuring and narrowing the compositionality gap in language models
Press, O., Zhang, M., Min, S., Schmidt, L., Smith, N. A., and Lewis, M · 2022
Earlier work this paper cites.
Towards a standard for identifying and managing bias in artificial intelligence
Schwartz, R., Vassilev, A., Greene, K., Perine, L., Burt, A., and Hall, P · 2022
Earlier work this paper cites.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Shaikh, O., Zhang, H., Held, W., Bernstein, M., and Yang, D · 2022
Cited alongside, same era.
Just fine-tune twice: Selective differential privacy for large language models
Shi, W., Shea, R., Chen, S., Zhang, C., Jia, R., and Yu, Z · 2022
Cited alongside, same era.
A robust bias mitigation procedure based on the stereotype content model
Ungless, E. L., Rafferty, A., Nag, H., and Ross, B · 2022
Cited alongside, same era.
Being ’seen’ vs. ’mis-seen’: Tensions between privacy and fairness in computer vision
Xiang, A · 2022
Cited alongside, same era.
Differentially private fine-tuning of language models
Yu, D., Naik, S., Backurs, A., Gopi, S., Inan, H. A., Kamath, G., Kulkarni, J., Lee, Y. T., Manoel, A., Wutschitz, L., Yekhanin, S., and Zhang, H · 2022
Social bias probing: Fairness benchmarking for language models
Manerba, M. M., Stańczak, K., Guidotti, R., and Augenstein, I · 2023
Later among the works it cites.
Survey on ai ethics: A socio-technical perspective, 2023
Mbiazi, D., Bhange, M., Babaei, M., Sheth, I., and Kenfack, P. J · 2023
Later among the works it cites.
Mireshghallah, N., Kim, H., Zhou, X., Tsvetkov, Y., Sap, M., Shokri, R., and Choi, Y · 2023
Later among the works it cites.
More human than human: Measuring chatgpt political bias
Motoki, F., Pinho Neto, V., and Rodrigues, V · 2023
Later among the works it cites.
Use of llms for illicit purposes: Threats, prevention measures, and vulnerabilities
Mozes, M., He, X., Kleinberg, B., and Griffin, L. D · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Certified robustness against natural language attacks by causal intervention
Zhao, H., Ma, C., Dong, X., Luu, A. T., Deng, Z.-H., and Zhang, H · 2022
Cited alongside, same era.
Llm safety review: Benchmarks and analysis
ActiveFence · 2023
Cited alongside, same era.
Intentional biases in llm responses
Badyal, N., Jacoby, D., and Coady, Y · 2023
Cited alongside, same era.
Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., et al · 2023
Cited alongside, same era.
Fleek: Factual error detection and correction with evidence retrieved from external knowledge
Bayat, F. F., Qian, K., Han, B., Sang, Y., Belyi, A., Khorshidi, S., Wu, F., Ilyas, I. F., and Li, Y · 2023
Cited alongside, same era.
A group fairness lens for large language models
Bi, G., Shen, L., Xie, Y., Cao, Y., Zhu, T., and He, X · 2023
Cited alongside, same era.
Science in the age of large language models
Birhane, A., Kasirzadeh, A., Leslie, D., and Wachter, S · 2023
Cited alongside, same era.
Later among the works it cites.
Socialstigmaqa: A benchmark to uncover stigma amplification in generative language models
Nagireddy, M., Chiazor, L., Singh, M., and Baldini, I · 2023
Later among the works it cites.
Is GPT-4 getting worse over time?
Narayanan, A. and Kapoor, S · 2023
Later among the works it cites.
In-contextual bias suppression for large language models
Oba, D., Kaneko, M., and Bollegala, D · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI, R · 2023
Later among the works it cites.
Ovalle, A., Mehrabi, N., Goyal, P., Dhamala, J., Chang, K.-W., Zemel, R., Galstyan, A., Pinter, Y., and Gupta, R · 2023
Later among the works it cites.
Controlling the extraction of memorized data from large language models via prompt-tuning
Ozdayi, M. S., Peris, C., Fitzgerald, J., Dupuy, C., Majmudar, J., Khan, H., Parikh, R., and Gupta, R · 2023
Later among the works it cites.
A domain-specific next-generation large language model (llm) or chatgpt is required for biomedical engineering and research
Pal, S., Bhattacharya, M., Lee, S.-S., and Chakraborty, C · 2023
Later among the works it cites.
Emptying the ocean with a spoon: Should we edit models?
Pinter, Y. and Elhadad, M · 2023
Later among the works it cites.
Guardrails ai
Rajpal, S · 2023
Later among the works it cites.
In-context retrieval-augmented language models
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Later among the works it cites.
Knowledge of cultural moral norms in large language models
Ramezani, A. and Xu, Y · 2023
Later among the works it cites.
A trip towards fairness: Bias and de-biasing in large language models
Ranaldi, L., Ruzzetti, E. S., Venditti, D., Onorati, D., and Zanzotto, F. M · 2023
Later among the works it cites.
Razumovskaia, E., Vulić, I., Marković, P., Cichy, T., Zheng, Q., Wen, T.-H., and Budzianowski, P · 2023
Later among the works it cites.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails
Rebedea, T., Dinu, R., Sreedhar, M., Parisien, C., and Cohen, J · 2023
Later among the works it cites.
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2023
Later among the works it cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D · 2023
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Later among the works it cites.
Fairness in serving large language models
Sheng, Y., Cao, S., Li, D., Zhu, B., Li, Z., Zhuo, D., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Subtle misogyny detection and mitigation: An expert-annotated dataset
Sheppard, B., Richter, A., Cohen, A., Smith, E. A., Kneese, T., Pelletier, C., Baldini, I., and Dong, Y · 2023
Later among the works it cites.
Aligning with whom? large language models have gender and racial biases in subjective nlp tasks
Sun, H., Pei, J., Choi, M., and Jurgens, D · 2023
Later among the works it cites.
TextVerifier: Robustness verification for textual classifiers with certifiable guarantees
Sun, S. and Ruan, W · 2023
Later among the works it cites.
What do llamas really think? revealing preference biases in language model representations
Tang, R., Zhang, X., Lin, J., and Ture, F · 2023
Later among the works it cites.
Auditing and mitigating cultural bias in llms
Tao, Y., Viberg, O., Baker, R. S., and Kizilcec, R. F · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Bypassing the safety training of open-source llms with priming attacks
Vega, J., Chaudhary, I., Xu, C., and Singh, G · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Large language models can be good privacy protection learners
Xiao, Y., Jin, Y., Bai, Y., Wu, Y., Yang, X., Luo, X., Yu, W., Zhao, X., Liu, Y., Chen, H., et al · 2023
Later among the works it cites.
An empirical analysis of parameter-efficient methods for debiasing pre-trained language models
Xie, Z. and Lukasiewicz, T · 2023
Later among the works it cites.
Promptcare: Prompt copyright protection by watermark injection and verification, 2023
Yao, H., Lou, J., Ren, K., and Qin, Z · 2023
Later among the works it cites.
Evaluating interfaced llm bias
Yeh, K.-C., Chi, J.-A., Lian, D.-C., and Hsieh, S.-K · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Yong, Z.-X., Menghini, C., and Bach, S. H · 2023
Later among the works it cites.
Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success
Zhang, Y. and Ippolito, D · 2023
Later among the works it cites.
Verify-and-edit: A knowledge-enhanced chain-of-thought framework
Zhao, R., Li, X., Joty, S., Qin, C., and Bing, L · 2023
Later among the works it cites.
Public perceptions of gender bias in large language models: Cases of chatgpt and ernie
Zhou, K. Z. and Sanfilippo, M. R · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
jailbreakchat., 2024
Albert, A · 2024
Closest in time.
Improving deep neural network generalization and robustness to background bias via layer-wise relevance propagation optimization
Bassi, P. R. A. S., Dertkigil, S. S. J., and Cavalli, A · 2024
Closest in time.
Towards guaranteed safe ai: A framework for ensuring robust and reliable ai systems, 2024
”davidad” Dalrymple, D., Skalse, J., Bengio, Y., Russell, S., Tegmark, M., Seshia, S., Omohundro, S., Szegedy, C., Goldhaber, B., Ammann, N., Abate, A., Halpern, J., Barrett, C., Zhao, D., Zhi-Xuan, T., Wing, J., and Tenenbaum, J · 2024
Closest in time.
Learning to trust your feelings: Leveraging self-awareness in llms for hallucination mitigation
Liang, Y., Song, Z., Wang, H., and Zhang, J · 2024
Closest in time.
What is the v-model in software development?
Oppermann, A · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
Tonmoy, S., Zaman, S., Jain, V., Rani, A., Rawte, V., Chadha, A., and Das, A · 2024
Closest in time.
Hallucination is inevitable: An innate limitation of large language models
Xu, Z., Jain, S., and Kankanhalli, M · 2024
Closest in time.