Fetching the paper…
Reading the bibliography…
Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems.
Computer Fraud and Abuse Act
CFAA · 1986
Earlier work this paper cites.
Sdmi cracks revealed
Greene, T. C · 2001
Earlier work this paper cites.
Cosmetic compliance and the failure of negotiated governance
Krawiec, K. D · 2003
Earlier work this paper cites.
Conflicts of interest and the case of auditor independence: Moral seduction and strategic issue cycling
Moore, D. A., Tetlock, P. E., Tanlu, L., and Bazerman, M. H · 2006
Earlier work this paper cites.
Rebooting responsible disclosure: a focus on protecting end users
Blog, G · 2010
Earlier work this paper cites.
Hacking the law: Are bug bounties a true safe harbor?
Elazari, A · 2018
Earlier work this paper cites.
Coming in from the cold: A safe harbor from the cfaa and the dmca §1201 for security researchers
Etcovich, D. and van der Merwe, T · 2018
Earlier work this paper cites.
Private ordering shaping cybersecurity policy: The case of bug bounties
Elazari, A · 2019
Earlier work this paper cites.
Is tricking a robot hacking?
Evtimov, I., O’Hair, D., Fernandes, E., Calo, R., and Kohno, T · 2019
Earlier work this paper cites.
Fraudsters used ai to mimic ceo’s voice in unusual cybercrime case
Stupp, C · 2019
Earlier work this paper cites.
How to make a chatbot that isn’t racist or sexist
Douglas Heaven, W · 2020
Earlier work this paper cites.
https://github.com/disclose/research-threats , 2021
Research threats: Legal threats against security researchers · 2021
Earlier work this paper cites.
Facebook banned me for life because i help people use it less, 10 2021
Barclay, L · 2021
Earlier work this paper cites.
Commission on information disorder final report
Boyd, D., DiResta, R., Donovan, J., douek, e., Frye, E., Gleicher, N., Raji, D., Rid, T., Roth, Y., Wanless, A., and Wolf, C · 2021
Earlier work this paper cites.
Missouri threatens to sue a reporter who flagged a security flaw
Brodkin, J · 2021
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Earlier work this paper cites.
The copyright office expands your security research rights
Colannino, J · 2021
Earlier work this paper cites.
Facebook disables ad observatory; academicians and journalists fire back
DeLong, L. A · 2021
Earlier work this paper cites.
The facebook files: A wall street journal investigation
Horwitz, J., Wells, G., Seetharaman, D., Hagey, K., Scheck, J., Purnell, N., Schechner, S., and Glazer, E · 2021
Earlier work this paper cites.
A proposal for researcher access to platform data: The platform transparency and accountability act
Persily, N · 2021
Earlier work this paper cites.
America’s anti-hacking laws pose a risk to national security
Pfefferkorn, R · 2021
Earlier work this paper cites.
The steep cost of capture
Whittaker, M · 2021
Earlier work this paper cites.
“transparency-washing” in the digital age : A corporate agenda of procedural fetishism
Zalnieriute, M · 2021
Earlier work this paper cites.
A safe harbor for platform research
Abdo, A., Krishnan, R., Krent, S., Welber Falcón, E., and Woods, A. K · 2022
Earlier work this paper cites.
Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem
Costanza-Chock, S., Raji, I. D., and Buolamwini, J · 2022
Earlier work this paper cites.
Department of justice announces new policy for charging cases under the computer fraud and abuse act
Department of Justice · 2022
Earlier work this paper cites.
It’s time to open the black box of social media
DiResta, R., Edelson, L., Nyhan, B., and Zuckerman, E · 2022
Earlier work this paper cites.
Bug bounties for algorithmic harms?, 2022
Kenway, J., François, C., Costanza-Chock, S., Raji, I. D., and Buolamwini, J · 2022
Earlier work this paper cites.
The time is now to develop community norms for the release of foundation models, 2022
Liang, P., Bommasani, R., Creel, K. A., and Reich, R · 2022
Earlier work this paper cites.
Shooting the messenger: Remediation of disclosed vulnerabilities as cfaa “loss”
Pfefferkorn, R · 2022
Earlier work this paper cites.
Outsider oversight: Designing a third party audit ecosystem for ai governance
Raji, I. D., Xu, P., Honigsberg, C., and Ho, D · 2022
Earlier work this paper cites.
Red-teaming the stable diffusion safety filter
Rando, J., Paleka, D., Lindner, D., Heim, L., and Tramèr, F · 2022
Earlier work this paper cites.
Dual use of artificial-intelligence-powered drug discovery
Urbina, F., Lentzos, F., Invernizzi, C., and Ekins, S · 2022
Earlier work this paper cites.
Post-summit civil society communique, 11 2023
Ada Lovelace Institute · 2023
Earlier work this paper cites.
Bug hunters’ perspectives on the challenges and benefits of the bug bounty ecosystem
Akgul, O., Eghtesad, T., Elazari, A., Gnawali, O., Grossklags, J., Mazurek, M. L., Votipka, D., and Laszka, A · 2023
Earlier work this paper cites.
Frontier ai regulation: Managing emerging risks to public safety, 2023
Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O’Keefe, C., Whittlestone, J., Avin, S., Brundage, M., Bullock, J., Cass-Beggs, D., Chang, B., Collins, T., Fist, T., Hadfield, G., Hayes, A., Ho, L., Hooker, S., Horvitz, E., Kolt, N., Schuett, J., Shavit, Y., Siddarth, D., Trager, R., and Wolf, K · 2023
Earlier work this paper cites.
Responsible disclosure policy, December 2023
Anthropic · 2023
Earlier work this paper cites.
Sam altman — who warned ai poses ‘risk of extinction’ to humanity — is also a ‘doomsday prepper’
Barrabi, T · 2023
Earlier work this paper cites.
100+ researchers say they stopped studying x, fearing elon musk might sue them
Belanger, A · 2023
Earlier work this paper cites.
Emergent autonomous scientific research capabilities of large language models, 2023
Boiko, D. A., MacKnight, R., and Gomes, G · 2023
Earlier work this paper cites.
Structured access for third-party research on frontier ai models: Investigating researchers’ model access requirements, 2023
Bucknall, B. S. and Trager, R. F · 2023
Earlier work this paper cites.
Vulnerability disclosure policy: What is it & why is it important?
Bugcrowd · 2023
Cited alongside, same era.
Artificial influence: An analysis of ai-driven persuasion
Burtell, M. and Woodside, T · 2023
Cited alongside, same era.
Extracting training data from diffusion models
Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tramer, F., Balle, B., Ippolito, D., and Wallace, E · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
The ftc voice cloning challenge
Commission, F. T · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2023
Later among the works it cites.
Whose opinions do language models reflect?
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T · 2023
Later among the works it cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Shah, R., Montixi, Q. F., Pour, S., Tagade, A., and Rando, J · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deshpande, A., Murahari, V., Rajpurohit, T., Kalyan, A., and Narasimhan, K · 2023
Cited alongside, same era.
Safe, secure, and trustworthy development and use of artificial intelligence
Executive Office of the President · 2023
Cited alongside, same era.
Ai red-teaming is not a one-stop solution to ai harms: Recommendations for using red-teaming for ai accountability
Friedler, S., Singh, R., Blili-Hamelin, B., Metcalf, J., and Chen, B. J · 2023
Cited alongside, same era.
Mart: Improving llm safety with multi-round automatic red-teaming
Ge, S., Zhou, C., Hou, R., Khabsa, M., Wang, Y.-C., Wang, Q., Han, J., and Mao, Y · 2023
Cited alongside, same era.
Generative ai has an intellectual property problem
Gil, A., Neelbauer, J., and A. Schweidel, D · 2023
Cited alongside, same era.
Asymmetric ideological segregation in exposure to political news on facebook
González-Bailón, S., Lazer, D., Barberá, P., Zhang, M., Allcott, H., Brown, T., Crespo-Tenorio, A., Freelon, D., Gentzkow, M., Guess, A. M., Iyengar, S., Kim, Y. M., Malhotra, N., Moehler, D., Nyhan, B., Pan, J., Rivera, C. V., Settle, J., Thorson, E., Tromble, R., Wilkins, A., Wojcieszak, M., de Jonge, C. K., Franco, A., Mason, W., Stroud, N. J., and Tucker, J. A · 2023
Cited alongside, same era.
The Times Sues OpenAI and Microsoft Over A.I. Use of Copyrighted Work
Grynbaum, M. M. and Mac, R · 2023
Cited alongside, same era.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Later among the works it cites.
Can large language models democratize access to dual-use biotechnology?
Soice, E. H., Rocha, R., Cordova, K., Specter, M., and Esvelt, K. M · 2023
Later among the works it cites.
Evaluating the social impact of generative ai systems in systems and society, 2023
Solaiman, I., Talat, Z., Agnew, W., Ahmad, L., Baker, D., Blodgett, S. L., au2, H. D. I., Dodge, J., Evans, E., Hooker, S., Jernite, Y., Luccioni, A. S., Lusoli, A., Mitchell, M., Newman, J., Png, M.-T., Strait, A., and Vassilev, A · 2023
Later among the works it cites.
The Coming Wave: Technology, Power, and the Twenty-First Century’s Greatest Dilemma
Suleyman, M. and Bhaskar, M · 2023
Later among the works it cites.
Artificial intelligence risk management framework (ai rmf 1.0), 2023-01-26 05:01:00 2023
Tabassi, E · 2023
Later among the works it cites.
URL https://assets-global.website-files.com/62713397a014368302d4ddf5/6579fcd1b821fdc1e507a6d0_Hacking-Policy-Council-statement-on-AI-red-teaming-protections-20231212.pdf
The Hacking Policy Council, Dec 2023 · 2023
Later among the works it cites.
Generative ml and csam: Implications and mitigations, 2023
Thiel, D., Stroebel, M., and Portnoff, R · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Advancing governance, innovation, and risk management for agency use of artificial intelligence, October 2023
United States Office of Management and Budget · 2023
Later among the works it cites.
Towards a greater understanding of coordinated vulnerability disclosure policy documents
Walshe, T. and Simpson, A · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Sociotechnical safety evaluation of generative ai systems
Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., and Isaac, W. S · 2023
Later among the works it cites.
’he would still be here’: Man dies by suicide after talking with ai chatbot, widow says
Xiang, C · 2023
Later among the works it cites.
Xu, R., Lin, B. S., Yang, S., Zhang, T., Shi, W., Zhang, T., Fang, Z., Xu, W., and Qiu, H · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Yong, Z. X., Menghini, C., and Bach, S · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., and Xing, X · 2023
Later among the works it cites.
Removing RLHF Protections in GPT-4 via Fine-Tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Later among the works it cites.
How language model hallucinations can snowball
Zhang, M., Press, O., Merrill, W., Liu, A., and Smith, N. A · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
The age of surveillance capitalism
Zuboff, S · 2023
Later among the works it cites.
Ai auditing: The broken bus on the road to ai accountability, 2024
Birhane, A., Steed, R., Ojewale, V., Vecchione, B., and Raji, I. D · 2024
Closest in time.
OpenAI says New York Times ’hacked’ ChatGPT to build copyright lawsuit
Brittain, B · 2024
Closest in time.
Black-box access is insufficient for rigorous ai audits, 2024
Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., Sharkey, L., Krishna, S., Hagen, M. V., Alberti, S., Chan, A., Sun, Q., Gerovitch, M., Bau, D., Tegmark, M., Krueger, D., and Hadfield-Menell, D · 2024
Closest in time.
Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts, 2024
European Council · 2024
Closest in time.
Llm agents can autonomously hack websites
Fang, R., Bindu, R., Gupta, A., Zhan, Q., and Kang, D · 2024
Closest in time.
Laion and the challenges of preventing ai-generated csam
Gupta, R · 2024
Closest in time.
On the societal impact of open foundation models
Kapoor, S., Bommasani, R., Klyman, K., Longpre, S., Ramaswami, A., Cihon, P., Hopkins, A., Bankston, K., Biderman, S., Bogen, M., et al · 2024
Closest in time.
Generative ai has a visual plagiarism problem
Marcus, G. and Southen, R · 2024
Closest in time.
Democratizing the future of ai r&d: Nsf to launch national ai research resource pilot
National Science Foundation · 2024
Closest in time.
Test, evaluation & red-teaming, 2024
NIST · 2024
Closest in time.
Researcher access program application, 2024
OpenAI · 2024
Closest in time.
Detecting pretraining data from large language models
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L · 2024
Closest in time.
Petition for new exemption to section 1201 of the digital millenium copyright act: Exemption for security research pertaining to generative ai bias, June 2023
Weiss, J · 2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.
Weak-to-strong jailbreaking on large language models, 2024
Zhao, X., Yang, X., Pang, T., Du, C., Li, L., Wang, Y.-X., and Wang, W. Y · 2024
Closest in time.
Generative red team recap, Oct 2023
Sven Cattell · 2031
Closest in time.