Fetching the paper…
Reading the bibliography…
The National Environment Policy Act (NEPA) stands as a foundational piece of environmental legislation in the United States, requiring federal agencies to consider the environmental impacts of their proposed actions.
Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel
J Peter Kincaid, Robert P Fishburne Jr, Richard L Rogers, and Brad S Chissom. 1975 · 1975
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Crowdsourcing Multiple Choice Science Questions. In Proceedings of the 3rd Workshop on Noisy User-generated Text . 94–106
Johannes Welbl, Nelson F Liu, and Matt Gardner. 2017 · 2017
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
DeepEval – The LLM Evaluation Framework
2023 · 2023
Earlier work this paper cites.
LawBench: Benchmarking Legal Knowledge of Large Language Models
Kai Chen, D. Zhu, Jidong Ge, Zhiwei Fei, Zhuo Han, Xiaoyu Shen, Zongwen Shen, Fengzhe Zhou, and Songyang Zhang. 2023 · 2023
Earlier work this paper cites.
LLMeBench: A Flexible Framework for Accelerating LLMs Benchmarking. In Conference of the European Chapter of the Association for Computational Linguistics
Fahim Dalvi, Maram Hasanain, Sabri Boughorbel, Basel Mousi, Samir Abdaljalil, Nizi Nazar, Ahmed Abdelali, Shammur A. Chowdhury, Hamdy Mubarak, Ahmed M. Ali, Majd Hawasly, Nadir Durrani, and Firoj Alam. 2023 · 2023
Earlier work this paper cites.
RAGAS: Automated Evaluation of Retrieval Augmented Generation
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2023 · 2023
Earlier work this paper cites.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, et al · 2023
Earlier work this paper cites.
SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research
Yang Han, Baocai Chen, Da Ma, Lu Chen, Liangtai Sun, Zihan Zhao, Kai Yu, and Zhe-Wei Shen. 2023 · 2023
Earlier work this paper cites.
AQ Jiang, A Sablayrolles, A Mensch, C Bamford, DS Chaplot, D de las Casas, F Bressand, G Lengyel, G Lample, L Saulnier, et al · 2023
Earlier work this paper cites.
Qasa: advanced question answering on scientific articles. In International Conference on Machine Learning . PMLR, 19036–19052
Yoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang, Jaehyeon Kim, Hong-in Lee, and Moontae Lee. 2023 · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology . 1–22
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023 · 2023
Earlier work this paper cites.
Scibench: Evaluating college-level scientific problem-solving abilities of large language models
Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang. 2023 · 2023
Earlier work this paper cites.
C-Pack: Packaged Resources To Advance General Chinese Embedding
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023 · 2023
Cited alongside, same era.
Article about Claude 3 Models
Anthropic Team and Collaborators. 2024 · 2024
Cited alongside, same era.
RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture
Angels Balaguer, Vinamra Benara, Renato Luiz de Freitas Cunha, Roberto de M. Estevão Filho, Todd Hendry, Daniel Holstein, Jennifer Marsman, Nick Mecklenburg, Sara Malvar, Leonardo O. Nunes, Rafael Padilha, Morris Sharp, Bruno Silva, Swati Sharma, Vijay Aski, and Ranveer Chandra. 2024 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Mdagents: An adaptive collaboration of llms for medical decision-making
Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik S Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, and Hae W Park. 2024a · 2024
Closest in time.
Shuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng, Zhenhui Peng, Ming Yin, and Xiaojuan Ma. 2024 · 2024
Closest in time.
WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain. In unknown
Rounak Meyur, Hung Phan, S. Wagle, Jan Strube, M. Halappanavar, Sameera Horawalavithana, Anurag Acharya, and Sai Munikoti. 2024 · 2024
Closest in time.
OpenAI. 2024 · 2024
Closest in time.
LLM for Environmental Review
PolicyAI. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Nygård Ege, Henrik H Øvrebø, Vegar Stubberud, Martin Francis Berg, Christer Elverum, Martin Steinert, and Håvard Vestad. 2024 · 2024
Cited alongside, same era.
Evaluating Large Language Models with fmeval
Luca Franceschi, Muhammad Bilal Zafar, Pinal Tailor, Pola Schwöbel, Michael Diamond, Michele Donini, Keerthan Vasist, Tomer Shenhar, Pinar Yilmaz, and Aman Malhotra. 2024 · 2024
Cited alongside, same era.
The Language Model Evaluation Harness
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2024 · 2024
Cited alongside, same era.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team and Collaborators. 2024 · 2024
Cited alongside, same era.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai
Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al · 2024
Cited alongside, same era.
Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation
Gauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, and Laurent Callot. 2024 · 2024
Cited alongside, same era.
CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language Models
Tianyu Guo, Weijian Sun, Ying Nie, Wei He, Binfan Zheng, Qiang Li, Hao Liu, Binwei Yan, Dacheng Tao, Yunhe Wang, Weihao Wang, and Haoyu Wang. 2024 · 2024
Cited alongside, same era.
EnviroExam: Benchmarking Environmental Science Knowledge of Large Language Models
Yu Huang, Liang Guo, Wanqian Guo, Zhe Tao, Yang Lv, Zhihao Sun, and Dongfang Zhao. 2024 · 2024
Cited alongside, same era.
Optimal decision making through scenario simulations using large language models
Sumedh Rasal and EJ Hauer. 2024 · 2024
Closest in time.
Empirical evaluation of uncertainty quantification in retrieval-augmented language models for science. In Proceedings of the Workshop on Scientific Document Understanding (SDU) . Vancouver, Canada
Sridevi Wagle, Sai Munikoti, Anurag Acharya, Sara Smith, and Sameera Horawalavithana. 2024 · 2024
Closest in time.
Unlocking the Potential: Benchmarking Large Language Models in Water Engineering and Research
Yue Yang, Boyan Xu, How yong Ng, Xiongpeng Tang, Rui Tong, Zihao Li, Xueqing Shi, Liang Wen, Yu Li, Yuxin Yang, Qingxian Su, Zihao Wu, and Guanlan Wu. 2024 · 2024
Closest in time.
KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models. In Annual Meeting of the Association for Computational Linguistics
Wenjin Yao, Jindong Wang, Zhuohao Yu, Chang Gao, Yidong Wang, Shikun Zhang, Xing Xie, Wei Ye, and Yue Zhang. 2024 · 2024
Closest in time.
TransportationGames: Benchmarking Transportation Knowledge of (Multimodal) Large Language Models
Xue Zhang, Xiangyu Shi, Xinyue Lou, Rui Qi, Yufeng Chen, Jinan Xu, and Wenjuan Han. 2024 · 2024
Closest in time.
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
Hao Cui, Zahra Shamsi, Gowoon Cheon, Xuejian Ma, Shutong Li, Maria Tikhanovskaya, Peter Norgaard, Nayantara Mudur, Martyna Plomecka, Paul Raccuglia, et al · 2025
Closest in time.
Jing Guo, Nan Li, and Ming Xu. 2025 · 2025
Closest in time.
Exploring LLMs Applications in Law: A Literature Review on Current Legal NLP Approaches
Marco Siino, Mariana Falco, Daniele Croce, and Paolo Rosso. 2025 · 2025
Closest in time.
Physreason: A comprehensive benchmark towards physics-based reasoning
Xinyu Zhang, Yuxuan Dong, Yanrui Wu, Jiaxing Huang, Chengyou Jia, Basura Fernando, Mike Zheng Shou, Lingling Zhang, and Jun Liu. 2025 · 2025
Closest in time.