Fetching the paper…
Reading the bibliography…
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science.
The impact factor
Eugene Garfield et al · 1994
Earlier work this paper cites.
The five environmental spheres
Stanley E Manahan · 2006
Earlier work this paper cites.
Matplotlib and seaborn
Ekaba Bisong · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
The era5 global reanalysis
Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, András Horányi, Joaquín Muñoz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, et al · 2020
Earlier work this paper cites.
Climatext: A dataset for climate change topic detection
Francesco S Varini, Jordan Boyd-Graber, Massimiliano Ciaramita, and Markus Leippold · 2020
Earlier work this paper cites.
A deep learning model based on bert and sentence transformer for semantic keyphrase extraction on big social data
R Devika, Subramaniyaswamy Vairavasundaram, C Sakthi Jay Mahenthar, Vijayakumar Varadarajan, and Ketan Kotecha · 2021
Earlier work this paper cites.
Climatebert: A pretrained language model for climate-related text
Nicolas Webersinke, Mathias Kraus, Julia Anna Bingler, and Markus Leippold · 2021
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Oceangpt: A large language model for ocean science tasks
Zhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, and Huajun Chen · 2023
Earlier work this paper cites.
Llm-based code generation method for golang compiler testing
Qiuhan Gu · 2023
Earlier work this paper cites.
Enhancing large language models with climate resources
Mathias Kraus, Julia Anna Bingler, Markus Leippold, Tobias Schimanski, Chiara Colesanti Senni, Dominik Stammbach, Saeid Ashraf Vaghefi, and Nicolas Webersinke · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Earlier work this paper cites.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting · 2023
Earlier work this paper cites.
Scibench: Evaluating college-level scientific problem-solving abilities of large language models
Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang · 2023
Earlier work this paper cites.
Towards llm-based fact verification on news claims with a hierarchical step-by-step prompting method
Xuan Zhang and Wei Gao · 2023
Cited alongside, same era.
Geogpt: Understanding and processing geospatial tasks through an autonomous gpt
Yifan Zhang, Cheng Wei, Shangyou Wu, Zhengting He, and Wenhao Yu · 2023
Cited alongside, same era.
This reference does not exist: an exploration of llm citation accuracy and relevance
Courtni Byun, Piper Vasicek, and Kevin Seppi · 2024
Cited alongside, same era.
Sciassess: Benchmarking llm proficiency in scientific literature analysis
Hengxing Cai, Xiaochen Cai, Junhan Chang, Sihang Li, Lin Yao, Changxin Wang, Zhifeng Gao, Hongshuai Wang, Yongge Li, Mujie Lin, et al · 2024
Cited alongside, same era.
From calculation to adjudication: Examining llm judges on mathematical reasoning tasks
Andreas Stephan, Dawei Zhu, Matthias Aßenmacher, Xiaoyu Shen, and Benjamin Roth · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al · 2024
Later among the works it cites.
Mineru: An open-source solution for precise document content extraction
Bin Wang, Chao Xu, Xiaomeng Zhao, Linke Ouyang, Fan Wu, Zhiyuan Zhao, Rui Xu, Kaiwen Liu, Yuan Qu, Fukai Shang, et al · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yanxi Chen, Yaliang Li, Bolin Ding, and Jingren Zhou · 2024
Cited alongside, same era.
K2: A foundation language model for geoscience knowledge understanding and utilization
Cheng Deng, Tianhang Zhang, Zhongmou He, Qiyuan Chen, Yuanyuan Shi, Yi Xu, Luoyi Fu, Weinan Zhang, Xinbing Wang, Chenghu Zhou, et al · 2024
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Cited alongside, same era.
Opendatalab: Empowering general artificial intelligence with open datasets
Conghui He, Wei Li, Zhenjiang Jin, Chao Xu, Bin Wang, and Dahua Lin · 2024
Cited alongside, same era.
Gpt-4o: The cutting-edge advancement in multimodal llm
Raisa Islam and Owana Marzia Moushi · 2024
Cited alongside, same era.
Diagnostic performances of claude 3 opus and claude 3.5 sonnet from patient history and key images in radiology’s “diagnosis please” cases
Ryo Kurokawa, Yuji Ohizumi, Jun Kanzawa, Mariko Kurokawa, Yuki Sonoda, Yuta Nakamura, Takao Kiguchi, Wataru Gonoi, and Osamu Abe · 2024
Cited alongside, same era.
Iterative large language models evolution through self-critique
Qianxi Li · 2024
Cited alongside, same era.
Llm with relation classifier for document-level relation extraction
Xingzuo Li, Kehai Chen, Yunfei Long, and Min Zhang · 2024
Cited alongside, same era.
Jiaheng Wei, Yuanshun Yao, Jean-Francois Ton, Hongyi Guo, Andrew Estornell, and Yang Liu · 2024
Later among the works it cites.
Generate-on-graph: Treat llm as both agent and kg in incomplete knowledge graph question answering
Yao Xu, Shizhu He, Jiabei Chen, Zihao Wang, Yangqiu Song, Hanghang Tong, Guang Liu, Kang Liu, and Jun Zhao · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Moose-chem: Large language models for rediscovering unseen chemistry scientific hypotheses
Zonglin Yang, Wanhao Liu, Ben Gao, Tong Xie, Yuqiang Li, Wanli Ouyang, Soujanya Poria, Erik Cambria, and Dongzhan Zhou · 2024
Later among the works it cites.
Easytool: Enhancing llm-based agents with concise tool instruction
Siyu Yuan, Kaitao Song, Jiangjie Chen, Xu Tan, Yongliang Shen, Ren Kan, Dongsheng Li, and Deqing Yang · 2024
Later among the works it cites.
Chemllm: A chemical large language model
Di Zhang, Wei Liu, Qian Tan, Jingdan Chen, Hang Yan, Yuliang Yan, Jiatong Li, Weiran Huang, Xiangyu Yue, Wanli Ouyang, et al · 2024
Later among the works it cites.
Grok, gemini, chatgpt and deepseek: Comparison and applications in conversational artificial intelligence
Murillo Edson de Carvalho Souza and Li Weigang · 2025
Closest in time.
Supergpqa: Scaling llm evaluation across 285 graduate disciplines
Xinrun Du, Yifan Yao, Kaijing Ma, Bingli Wang, Tianyu Zheng, King Zhu, Minghao Liu, Yiming Liang, Xiaolong Jin, Zhenlin Wei, et al · 2025
Closest in time.
The accuracy, robustness, and readability of llm-generated sustainability-related word definitions
Alice Heiman · 2025
Closest in time.
Perception, reason, think, and plan: A survey on large multimodal reasoning models
Yunxin Li, Zhenyu Liu, Zitao Li, Xuanyu Zhang, Zhenran Xu, Xinyu Chen, Haoyuan Shi, Shenyuan Jiang, Xintong Wang, Jifang Wang, et al · 2025
Closest in time.
Evaluating the efficacy of large language models in generating medical documentation: A comparative study of chatgpt-4, chatgpt-4o, and claude
Bryan Lim, Ishith Seth, Molly Maxwell, Roberto Cuomo, Richard J Ross, and Warren M Rozen · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al · 2025
Closest in time.
The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha · 2025
Closest in time.