Fetching the paper…
Reading the bibliography…
Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice.
Underwriters’ laboratories: testing for public safety
G. E. Schall · 1970
Earlier work this paper cites.
The social control of technology
D. Collinridge · 1982
Earlier work this paper cites.
Cognitive aspects of survey measurement and mismeasurement
R. Tourangeau · 2003
Earlier work this paper cites.
Structured transparency: a framework for addressing use/mis-use trade-offs when sharing information
A. Trask, E. Bluemke, B. Garfinkel, C. G. Cuervas-Mons, and A. Dafoe · 2012
Earlier work this paper cites.
A framework for responsible innovation
R. Owen, P. M. Stilgoe, and J. Bessant · 2013
Earlier work this paper cites.
Collingridge and the dilemma of control: towards responsible and accountable innovation
A. Genus and A. Stirling · 2018
Earlier work this paper cites.
XTREME: a massively multilingual multi-task benchmark for evaluating cross-lingual generalization
J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. First, and M. Johnson · 2020
Earlier work this paper cites.
Trustworthy online controlled experiments: a practical guide to A/B testing
R. Kohavi, D. Tang, and Y. Xu · 2020
Earlier work this paper cites.
Decolonial AI: decolonial theory as sociotechnical foresight in artificial intelligence
S. Mohamed, M.-T. Png, and W. Isaac · 2020
Earlier work this paper cites.
Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing
I. D. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes · 2020
Earlier work this paper cites.
On the value of out-of-distribution testing: an example of Goodhart’s law
D. Teney, K. Kafle, R. Shrestha, E. Abbasnejad, C. Kanan, and A. van den Hengel · 2020
Earlier work this paper cites.
Forty years of food safety risk assessment: a history and analysis
F. Wu and J. V. Rodricks · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: can language models be too big?
E. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Earlier work this paper cites.
Measurement and fairness
A. Z. Jacobs and H. Wallach · 2021
Earlier work this paper cites.
Visually grounded reasoning across languages and cultures
F. Liu, E. Bugliarello, E. M. Ponti, S. Reddy, N. Collier, and D. Elliott · 2021
Earlier work this paper cites.
BBQ: a hand-built bias benchmark for question answering
A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. R. Bowman · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Earlier work this paper cites.
AI and the everything in the whole wide world benchmark
I. D. Raji, E. M. Bender, A. Paullada, E. Denton, and A. Hanna · 2021
Earlier work this paper cites.
A step toward more inclusive people annotations for fairness
C. Schumann, S. Ricco, U. Prabhu, V. Ferrari, and C. Pantofaru · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
A. Glaese, N. McAleese, M. Trębacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, L. Campbell-Gillingham, J. Uesato, P.-S. Huang, R. Comanescu, F. Yang, A. See, S. Dathathri, R. Greig, C. Chen, D. Fritz, J. S. Elias, R. Green, S. Mokrá, N. Fernando, B. Wu, R. Foley, S. Young, I. Gabriel, W. Isaac, J. Mellor, D. Hassabis, K. Kavukcuoglu, L. A. Hendricks, and G. Irving · 2022
Earlier work this paper cites.
X-risk analysis for AI research
D. Hendrycks and M. Mazeika · 2022
Earlier work this paper cites.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Cited alongside, same era.
The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world
W. A. G. Rojas, S. Diamos, K. R. Kini, D. Kanter, V. J. Reddi, and C. Coleman · 2022
Cited alongside, same era.
Structured access: an emerging paradigm for safe AI deployment
T. Shevlane · 2022
Cited alongside, same era.
Large language models encode clinical knowledge
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, P. Payne, M. Seneviratne, P. Gamble, C. Kelly, N. Scharli, A. Chowdhery, P. Mansfield, B. A. y. Arcas, D. Webster, G. S. Corrado, Y. Matias, K. Chou, J. Gottweis, N. Tomasev, Y. Liu, A. Rajkomar, J. Barral, C. Semturs, A. Karthikesalingam, and V. Natarajan · 2022
Cited alongside, same era.
Co-writing with opinionated language models affects users’ views
M. Jakesch, A. Bhat, D. Buschek, L. Zalmanson, and M. Naaman · 2023
Later among the works it cites.
Auditing large language models: a three-layered approach
J. Mökander, J. Schuett, H. R. Kirk, and L. Floridi · 2023
Later among the works it cites.
Capabilities of GPT-4 on medical challenge problems
H. Nori, N. King, S. M. McKinney, D. Carignan, and E. Horvitz · 2023
Later among the works it cites.
GPT-4V(ision) system card, Sept. 2023
OpenAI · 2023
Later among the works it cites.
Exploring relationship development with social chatbots: a mixed-method study of Replika
I. Pentina, T. Hancock, and T. Xie · 2023
Later among the works it cites.
From plane crashes to algorithmic harm: applicability of safety engineering frameworks for responsible ML
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taxonomy of risks posed by language models
L. Weidinger, J. Uesato, M. Rauh, C. Griffin, P.-S. Huang, J. Mellor, A. Glaese, M. Cheng, B. Balle, A. Kasirzadeh, C. Biles, S. Brown, Z. Kenton, W. Hawkins, T. Stepleton, A. Birhane, L. A. Hendricks, L. Rimell, W. Isaac, J. Haas, S. Legassick, G. Irving, and I. Gabriel · 2022
Cited alongside, same era.
Dices dataset: Diversity in conversational ai evaluation for safety
L. Aroyo, A. S. Taylor, M. Diaz, C. M. Homan, A. Parrish, G. Serapio-Garcia, V. Prabhakaran, and D. Wang · 2023
Cited alongside, same era.
Representation in AI evaluations
A. S. Bergman, L. A. Hendricks, M. Rauh, B. Wu, W. Agnew, M. Kunesch, I. Duan, I. Gabriel, and W. Isaac · 2023
Cited alongside, same era.
Typology of risks of generative text-to-image models
C. Bird, E. Ungless, and A. Kasirzadeh · 2023
Cited alongside, same era.
Towards monosemanticity: decomposing language models with dictionary learning, Oct. 2023
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. L. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah · 2023
Cited alongside, same era.
Responsibility & safety, 2023
DeepMind · 2023
Cited alongside, same era.
Are large language models a threat to digital public goods? evidence from activity on stack overflow
M. del Rio-Chanona, N. Laurentsyeva, and J. Wachs · 2023
Cited alongside, same era.
Investigating data contamination in modern benchmarks for large language models
C. Deng, Y. Zhao, X. Tang, M. Gerstein, and A. Cohan · 2023
Cited alongside, same era.
S. Rismani, R. Shelby, A. Smart, E. Jatho, J. Kroll, A. Moon, and N. Rostamzadeh · 2023
Later among the works it cites.
Data contamination through the lens of time
M. Roberts, H. Thakur, C. Herlihy, C. White, and S. Dooley · 2023
Later among the works it cites.
Sociotechnical harms of algorithmic systems: scoping a taxonomy for harm reduction
R. Shelby, S. Rismani, K. Henne, A. Moon, N. Rostamzadeh, P. Nicholas, N. Yilla-Akbari, J. Gallegos, A. Smart, E. Garcia, and G. Virk · 2023
Later among the works it cites.
Model evaluation for extreme risks
T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt, L. Ho, D. Siddarth, S. Avin, W. Hawkins, B. Kim, I. Gabriel, V. Bolina, J. Clark, Y. Bengio, P. Christiano, and A. Dafoe · 2023
Later among the works it cites.
Evaluating the social impact of generative AI systems in systems and society
I. Solaiman, Z. Talat, W. Agnew, L. Ahmad, D. Baker, S. L. Blodgett, H. Daumé III, J. Dodge, E. Evans, S. Hooker, et al · 2023
Later among the works it cites.
‘State of the science’ report to understand capabilities and risks of frontier AI: statement by the Chair, 2 November 2023, Nov. 2023
UK Government · 2023
Later among the works it cites.
Sociotechnical safety evaluation of generative AI systems
L. Weidinger, M. Rauh, N. Marchal, A. Manzini, L. A. Hendricks, J. Mateos-Garcia, S. Bergman, J. Kay, C. Griffin, B. Bariach, I. Gabriel, V. Rieser, and W. Isaac · 2023
Later among the works it cites.
Ensuring safe, secure, and trustworthy AI, 2023
White House · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Z.-X. Yong, C. Menghini, and S. H. Bach · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson · 2023
Later among the works it cites.
Red-teaming for generative AI: Silver bullet or security theater?
M. Feffer, A. Sinha, Z. C. Lipton, and H. Heidari · 2024
Closest in time.
The ethics of advanced ai assistants
I. Gabriel, A. Manzini, G. Keeling, L. A. Hendricks, V. Rieser, H. Iqbal, N. Tomašev, I. Ktena, Z. Kenton, M. Rodriguez, S. El-Sayed, S. Brown, C. Akbulut, A. Trask, E. Hughes, A. S. Bergman, R. Shelby, N. Marchal, C. Griffin, J. Mateos-Garcia, L. Weidinger, W. Street, B. Lange, A. Ingerman, A. Lentz, R. Enger, A. Barakat, V. Krakovna, J. O. Siy, Z. Kurth-Nelson, A. McCroskery, V. Bolina, H. Law, M. Shanahan, L. Alberts, B. Balle, S. de Haas, Y. Ibitoye, A. Dafoe, B. Goldberg, S. Krier, A. Reese, S. Witherspoon, W. Hawkins, M. Rauh, D. Wallace, M. Franklin, J. A. Goldstein, J. Lehman, M. Klenk, S. Vallor, C. Biles, M. Ringel Morris, H. King, B. Agüera y Arcas, W. Isaac, and J. Manyika · 2024
Closest in time.
Gemini: a family of highly capable multimodal models; v1 update
n. Gemini Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2024
Closest in time.
How does generative artificial intelligence impact student creativity?
S. Habib, T. Vogel, X. Anli, and E. Thorne · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities
M. Phuong, M. Aitchison, E. Catt, S. Cogan, A. Kaskasoli, V. Krakovna, D. Lindner, M. Rahtz, Y. Assael, S. Hodkinson, H. Howard, T. Lieberum, R. Kumar, M. A. Raad, A. Webson, L. Ho, S. Lin, S. Farquhar, M. Hutter, G. Deletang, A. Ruoss, S. El-Sayed, S. Brown, A. Dragan, R. Shah, A. Dafoe, and T. Shevlane · 2024
Closest in time.
Gradient-based language model red teaming
N. Wichers, C. Denison, and A. Beirami · 2024
Closest in time.