Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are playing an increasingly important role in scientific research, yet there remains a lack of comprehensive benchmarks to evaluate the breadth and depth of scientific knowledge embedded in these models.
Identification of common molecular subsequences
T. F. Smith, M. S. Waterman, et al · 1981
Earlier work this paper cites.
A revision of Bloom’s taxonomy: An overview
D. R. Krathwohl · 2002
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Predicting organic reaction outcomes with Weisfeiler-Lehman network
W. Jin, C. Coley, R. Barzilay, and T. Jaakkola · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
J. Welbl, N. F. Liu, and M. Gardner · 2017
Earlier work this paper cites.
Think you have solved question answering? Try ARC, the AI2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
MoleculeNet: A benchmark for molecular machine learning
Z. Wu, B. Ramsundar, E. N. Feinberg, J. Gomes, C. Geniesse, A. S. Pappu, K. Leswing, and V. Pande · 2018
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Q. Jin, B. Dhingra, Z. Liu, W. W. Cohen, and X. Lu · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Chromatin potential identified by shared single-cell profiling of rna and chromatin
S. Ma, B. Zhang, L. M. LaFave, A. S. Earl, Z. Chiang, Y. Hu, J. Ding, A. Brack, V. K. Kartha, T. Tay, et al · 2020
Earlier work this paper cites.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang · 2021
Earlier work this paper cites.
Text2mol: Cross-modal molecule retrieval with natural language queries
C. Edwards, C. Zhai, and H. Ji · 2021
Earlier work this paper cites.
Pubchem in 2021: new data content and improved web interfaces
S. Kim, J. Chen, T. Cheng, A. Gindulyte, J. He, S. He, Q. Li, B. A. Shoemaker, P. A. Thiessen, B. Yu, et al · 2021
Earlier work this paper cites.
Translation between molecules and natural language
C. Edwards, T. Lai, K. Ros, G. Honke, K. Cho, and H. Ji · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
A. Pal, L. K. Umapathi, and M. Sankarasubbu · 2022
Cited alongside, same era.
Galactica: A large language model for science
R. Taylor, M. Kardas, G. Cucurull, T. Scialom, A. Hartshorn, E. Saravia, A. Poulton, V. Kerkez, and R. Stojnic · 2022
Cited alongside, same era.
PEER: A comprehensive and multi-task benchmark for protein sequence understanding
M. Xu, Z. Zhang, J. Lu, Z. Zhu, Y. Zhang, M. Chang, R. Liu, and J. Tang · 2022
Cited alongside, same era.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Later among the works it cites.
The Claude 3 model family: Opus, sonnet, haiku
A. Anthropic · 2024
Closest in time.
SciAssess: Benchmarking LLM proficiency in scientific literature analysis
H. Cai, X. Cai, J. Chang, S. Li, L. Yao, C. Wang, Z. Gao, Y. Li, M. Lin, S. Yang, et al · 2024
Closest in time.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Closest in time.
L+M-24: Building a dataset for Language+Molecules @ ACL 2024
C. Edwards, Q. Wang, L. Zhao, and H. Ji · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Red-teaming large language models using chain of utterances for safety-alignment
R. Bhardwaj and S. Poria · 2023
Cited alongside, same era.
UniProt: the universal protein knowledgebase in 2023
U. Consortium · 2023
Cited alongside, same era.
Mol-Instructions: A large-scale biomolecular instruction dataset for large language models
Y. Fang, X. Liang, N. Zhang, K. Liu, R. Huang, Z. Chen, X. Fan, and H. Chen · 2023
Cited alongside, same era.
Xiezhi: An ever-updating benchmark for holistic domain knowledge evaluation
Z. Gu, X. Zhu, H. Ye, L. Zhang, J. Wang, Y. Zhu, S. Jiang, Z. Xiong, Z. Li, W. Wu, Q. He, R. Xu, W. Huang, J. Liu, Z. Wang, S. Wang, W. Zheng, H. Feng, and Y. Xiao · 2023
Cited alongside, same era.
What can large language models do in chemistry? A comprehensive benchmark on eight tasks
T. Guo, B. Nan, Z. Liang, Z. Guo, N. Chawla, O. Wiest, X. Zhang, et al · 2023
Cited alongside, same era.
Control risk for potential misuse of artificial intelligence in science
J. He, W. Feng, Y. Min, J. Yi, K. Tang, S. Li, J. Zhang, K. Chen, W. Zhou, X. Xie, et al · 2023
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with Pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica · 2023
Cited alongside, same era.
Lab-Bench: Measuring capabilities of language models for biology research
J. M. Laurent, J. D. Janizek, M. Ruzo, M. M. Hinks, M. J. Hammerling, S. Narayanan, M. Ponnapati, A. D. White, and S. G. Rodriques · 2024
Closest in time.
Are large language models superhuman chemists?
A. Mirza, N. Alampara, S. Kunchapu, B. Emoekabu, A. Krishnan, M. Wilhelmi, M. Okereke, J. Eberhardt, A. M. Elahi, M. Greiner, et al · 2024
Closest in time.
Openai o1 system card, 2024
OpenAI · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Scieval: A multi-level large language model evaluation benchmark for scientific research
L. Sun, Y. Han, Z. Zhao, D. Ma, Z. Shen, B. Chen, L. Chen, and K. Yu · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, et al · 2024
Closest in time.
B. Yu, F. N. Baker, Z. Chen, X. Ning, and H. Sun · 2024
Closest in time.
SciGLM: Training scientific language models with self-reflective instruction annotation and tuning
D. Zhang, Z. Hu, S. Zhoubian, Z. Du, K. Yang, Z. Wang, Y. Yue, Y. Dong, and J. Tang · 2024
Closest in time.
ChemLLM: A chemical large language model
D. Zhang, W. Liu, Q. Tan, J. Chen, H. Yan, Y. Yan, J. Li, W. Huang, X. Yue, D. Zhou, et al · 2024
Closest in time.
ChemDFM: Dialogue foundation model for chemistry
Z. Zhao, D. Ma, L. Chen, L. Sun, Z. Li, H. Xu, Z. Zhu, S. Zhu, S. Fan, G. Shen, et al · 2024
Closest in time.