Fetching the paper…
Reading the bibliography…
We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols.
PubMedQA: A dataset for biomedical research question answering, 2019
Q. Jin, B. Dhingra, Z. Liu, W. W. Cohen, and X. Lu · 1909
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021b
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2009
Earlier work this paper cites.
Tacit and Explicit Knowledge
H. Collins · 2010
Earlier work this paper cites.
Tacit knowledge and the biological weapons regime
J. Revill and C. Jefferson · 2013
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Highly accurate protein structure prediction with alphafold, 2021
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, et al · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in nlp, 2021
D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, Z. Ma, T. Thrush, S. Riedel, Z. Waseem, P. Stenetorp, R. Jia, M. Bansal, C. Potts, and A. Williams · 2021
Earlier work this paper cites.
The characteristics of pandemic pathogens, 2018
A. A. Adalja, M. Watson, E. S. Toner, A. Cicero, and T. V. Inglesby · 2022
Earlier work this paper cites.
Mapping global dynamics of benchmark creation and saturation in artificial intelligence
S. Ott, A. Barbosa-Silva, K. Blagec, J. Brauner, and M. Samwald · 2022
Earlier work this paper cites.
A. Pal, L. K. Umapathi, and M. Sankarasubbu · 2022
Earlier work this paper cites.
Safeguarding mail-order DNA synthesis in the age of artificial intelligence
S. Batalis, C. Schuerger, G. K. Gronvall, and M. E. Walsh · 2023
Earlier work this paper cites.
Biosecurity risk assessment for the use of artificial intelligence in synthetic biology, 2024
L. P. De Haro · 2023
Earlier work this paper cites.
A. Gopal, N. Helm-Burger, L. Justen, E. H. Soice, T. Tzeng, G. Jeyapragasan, S. Grimm, B. Mueller, and K. M. Esvelt · 2023
Earlier work this paper cites.
Biosecurity in the age of AI, 2023
Helena · 2023
Earlier work this paper cites.
Screening state of play: The biosecurity practices of synthetic DNA providers
A. Kane and M. T. Parker · 2023
Earlier work this paper cites.
Critical assessment of methods of protein structure prediction (CASP)—round xv
A. Kryshtafovych, T. Schwede, M. Topf, K. Fidelis, and J. Moult · 2023
Earlier work this paper cites.
Proposed biosecurity oversight framework for the future of science, 2023
K. N. C. Letts · 2023
Earlier work this paper cites.
The operational risks of AI in large-scale biological attacks: A red-team approach, 2023
C. A. Mouton, C. Lucas, and E. Guest · 2023
Earlier work this paper cites.
Preparedness framework (beta), 2023
OpenAI · 2023
Earlier work this paper cites.
GPQA: A graduate-level google-proof q&a benchmark, 2023
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2023
Earlier work this paper cites.
Understanding AI-facilitated biological weapon development, 2023
S. Rose and C. Nelson · 2023
Earlier work this paper cites.
J. B. Sandbrink · 2023
Cited alongside, same era.
Enhancing gene synthesis security: An updated framework for synthetic nucleic acid screening and the responsible use of synthetic biological materials
C. M. Sharkey, M. Lekveishvili, T. de la Rosa, and K. Danskin · 2023
Cited alongside, same era.
Large language models encode clinical knowledge, 2023
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, et al · 2023
Cited alongside, same era.
NexusRaven: A commercially-permissive language model for function calling
V. K. Srinivasan, Z. Dong, B. Zhu, B. Yu, H. Mao, D. Mosk-Aoyama, K. Keutzer, J. Jiao, and J. Zhang · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023
The WMDP benchmark: Measuring and reducing malicious use with unlearning, 2024
N. Li, A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A.-K. Dombrowski, S. Goel, L. Phan, G. Mukobi, N. Helm-Burger, R. Lababidi, L. Justen, A. B. Liu, M. Chen, I. Barrass, O. Zhang, X. Zhu, R. Tamirisa, B. Bharathi, A. Khoja, Z. Zhao, A. Herbert-Voss, C. B. Breuer, S. Marks, O. Patel, A. Zou, M. Mazeika, Z. Wang, P. Oswal, W. Lin, A. A. Hunt, J. Tienken-Harder, K. Y. Shih, K. Talley, J. Guan, R. Kaplan, I. Steneker, D. Campbell, B. Jokubaitis, A. Levinson, J. Wang, W. Qian, K. K. Karmakar, S. Basart, S. Fitz, M. Levine, P. Kumaraguru, U. Tupakula, V. Varadharajan, R. Wang, Y. Shoshitaishvili, J. Ba, K. M. Esvelt, A. Wang, and D. Hendrycks · 2024
Later among the works it cites.
MathVista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2024
Later among the works it cites.
T. R. McIntosh, T. Susnjak, N. Arachchilage, T. Liu, P. Watters, and M. N. Halgamuge · 2024
Later among the works it cites.
How predictable is language model benchmark performance?, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, A. Kluska, A. Lewkowycz, A. Agarwal, A. Power, A. Ray, A. Warstadt, A. W. Kocurek, A. Safaya, A. Tazarv, A. Xiang, A. Parrish, A. Nie, A. Hussain, A. Askell, A. Dsouza, A. Slone, A. Rahane, A. S. Iyer, A. Andreassen, A. Madotto, A. Santilli, A. Stuhlmüller, A. Dai, A. La, A. Lampinen, A. Zou, et al · 2023
Cited alongside, same era.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2023
Cited alongside, same era.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Cited alongside, same era.
AgentHarm: A benchmark for measuring harmfulness of llm agents, 2024
M. Andriushchenko, A. Souly, M. Dziemian, D. Duenas, M. Lin, J. Wang, D. Hendrycks, A. Zou, Z. Kolter, M. Fredrikson, E. Winsor, J. Wynne, Y. Gal, and X. Davies · 2024
Cited alongside, same era.
Responsible scaling policy, 2024
Anthropic · 2024
Cited alongside, same era.
Lessons from the trenches on reproducible evaluation of language models, 2024-05-29
S. Biderman, H. Schoelkopf, L. Sutawika, L. Gao, J. Tow, B. Abbasi, A. F. Aji, P. S. Ammanamanchi, S. Black, J. Clive, A. DiPofi, J. Etxaniz, B. Fattori, J. Z. Forde, C. Foster, J. Hsu, M. Jaiswal, W. Y. Lee, H. Li, C. Lovering, N. Muennighoff, E. Pavlick, J. Phang, A. Skowron, S. Tan, X. Tang, K. A. Wang, G. I. Winata, F. Yvon, and A. Zou · 2024
Cited alongside, same era.
AI and biosecurity: The need for governance, 2024
D. Bloomfield, J. Pannu, A. W. Zhu, M. Y. Ng, A. Lewis, E. Bendavid, S. M. Asch, T. Hernandez-Boussard, A. Cicero, and T. Inglesby · 2024
Cited alongside, same era.
MLE-bench: Evaluating machine learning agents on machine learning engineering, 2024
J. S. Chan, N. Chowdhury, O. Jaffe, J. Aung, D. Sherburn, E. Mays, G. Starace, K. Liu, L. Maksin, T. Patwardhan, L. Weng, and A. Mądry · 2024
Cited alongside, same era.
D. Owen · 2024
Later among the works it cites.
Leaving the barn door open for clever hans: Simple features predict LLM benchmark answers, 2024
L. Pacchiardi, M. Tesic, L. G. Cheke, and J. Hernández-Orallo · 2024
Later among the works it cites.
Building an early warning system for LLM-aided biological threat creation, 2024
T. Patwardhan, K. Liu, T. Markov, N. Chowdhury, D. Leet, N. Cone, C. Maltbie, J. Huizinga, C. Wainwright, S. F. Jackson, S. Adler, R. Casagrande, A. Madry, and OpenAI · 2024
Later among the works it cites.
The reality of AI and biorisk, 2024
A. Peppin, A. Reuel, S. Casper, E. Jones, A. Strait, U. Anwar, A. Agrawal, S. Kapoor, S. Koyejo, M. Pellat, R. Bommasani, N. Frosst, and S. Hooker · 2024
Later among the works it cites.
CRISPR-GPT: An LLM agent for automated design of gene-editing experiments, 2024
Y. Qu, K. Huang, H. Cousins, W. A. Johnson, D. Yin, M. Shah, D. Zhou, R. Altman, M. Wang, and L. Cong · 2024
Later among the works it cites.
BetterBench: Assessing AI benchmarks, uncovering issues, and establishing best practices, 2024
A. Reuel, A. Hardy, C. Smith, M. Lamparth, M. Hardy, and M. J. Kochenderfer · 2024
Later among the works it cites.
MMLU-Pro+: Evaluating higher-order reasoning and shortcut learning in llms, 2024
S. A. Taghanaki, A. Khani, and A. Khasahmadi · 2024
Later among the works it cites.
BioCoder: a benchmark for bioinformatics code generation with large language models, 2024
X. Tang, B. Qian, R. Gao, J. Chen, X. Chen, and M. B. Gerstein · 2024
Later among the works it cites.
Inspect AI: Framework for Large Language Model Evaluations, 2024
UK AI Safety Institute · 2024
Later among the works it cites.
The whack-a-mole governance challenge for AI-enabled synthetic biology: literature review and emerging frameworks, 2024
T. A. Undheim · 2024
Later among the works it cites.
Towards risk analysis of the impact of AI on the deliberate biological threat landscape, 2024
M. E. Walsh · 2024
Later among the works it cites.
Virologist opinions: An important component for the governance of the convergence of artificial intelligence and dual-use research of concern, 2025
M. E. Walsh and G. K. Gronvall · 2024
Later among the works it cites.
Berkeley function calling leaderboard
F. Yan, H. Mao, C. C.-J. Ji, T. Zhang, S. G. Patil, I. Stoica, and J. E. Gonzalez · 2024
Later among the works it cites.
τ \tau -bench: A benchmark for tool-agent-user interaction in real-world domains, 2024
S. Yao, N. Shinn, P. Razavi, and K. Narasimhan · 2024
Later among the works it cites.
Toward comprehensive benchmarking of the biological knowledge of frontier large language models, 2025
S. Dev, C. Teague, K. Brady, Y.-C. J. Lee, S. L. Gebauer, H. A. Bradley, G. Ellison, B. Persaud, J. Despanie, B. D. Castello, A. Worland, M. Miller, D. Maciorowski, A. Salas, D. K. Nguyen, J. Liu, J. Johnson, A. Sloan, W. Stonehouse, T. Merrill, T. Goode, J. Greg McKelvey, and E. Guest · 2025
Closest in time.
Foundation models may exhibit staged progression in novel cbrn threat disclosure, 2025
K. M. Esvelt · 2025
Closest in time.
L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, S. Shi, M. Choi, A. Agrawal, A. Chopra, A. Khoja, R. Kim, J. Hausenloy, O. Zhang, M. Mazeika, et al · 2025
Closest in time.