Fetching the paper…
Reading the bibliography…
AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks.
Construct validity in psychological tests
Lee J Cronbach and Paul E Meehl · 1955
Earlier work this paper cites.
The Concept of Ecological Validity: What Are Its Limitations and Is It Bad to Be Invalid?
David J Lewkowicz · 2001
Earlier work this paper cites.
Focus article: On the structure of educational assessments
Robert J Mislevy, Linda S Steinberg, and Russell G Almond · 2003
Earlier work this paper cites.
Good benchmarks are hard to find: Toward the benchmark for information retrieval applications in software engineering
Alex Dekhtyar and Jane Huffman Hayes · 2006
Earlier work this paper cites.
Issues in bioinformatics benchmarking: the case study of multiple sequence alignment
Mohamed Radhouene Aniba, Olivier Poch, and Julie D. Thompson · 2010
Earlier work this paper cites.
Protein sequence comparison and fold recognition: progress and good-practice benchmarking
Johannes Söding and Michael Remmert · 2011
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Peer review and the publication process
Parveen Azam Ali and Roger Watson · 2016
Earlier work this paper cites.
The FAIR Guiding Principles for scientific data management and stewardship
Mark D Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E Bourne, et al · 2016
Earlier work this paper cites.
Environmental Quality Benchmarks — The Good, The Bad, and The Ugly
Peter M Chapman · 2018
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Artificial intelligence faces reproducibility crisis
Matthew Hutson · 2018
Earlier work this paper cites.
Ethics and privacy in AI and big data: Implementing responsible research and innovation
Bernd Carsten Stahl and David Wright · 2018
Earlier work this paper cites.
Essential guidelines for computational method benchmarking
Lukas M Weber, Wouter Saelens, Robrecht Cannoodt, Charlotte Soneson, Alexander Hapfelmeier, Paul P Gardner, Anne-Laure Boulesteix, Yvan Saeys, and Mark D Robinson · 2019
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Benchmarking in Optimization: Best Practice and Open Issues
Thomas Bartz-Beielstein, Carola Doerr, Jakob Bossek, Sowmya Chandrasekaran, Tome Eftimov, Andreas Fischbach, Pascal Kerschke, Manuel López-Ibáñez, Katherine Mary Malan, Jason H. Moore, Boris Naujoks, Patryk Orzechowski, Vanessa Volz, Markus Wagner, and T. Weise · 2020
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 2020
Earlier work this paper cites.
RL unplugged: a suite of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gómez Colmenarejo, Konrad Zołna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, Gabriel Dulac-Arnold, Jerry Li, Mohammad Norouzi, Matt Hoffman, Nicolas Heess, and Nando de Freitas · 2020
Earlier work this paper cites.
WordCraft: An Environment for Benchmarking Commonsense Agents
Minqi Jiang, Jelena Luketina, Nantas Nardelli, Pasquale Minervini, Philip H.S. Torr, Shimon Whiteson, and Tim Rocktäschel · 2020
Earlier work this paper cites.
Is deep reinforcement learning ready for practical applications in healthcare? a sensitivity analysis of duel-ddqn for hemodynamic management in sepsis patients
MingYu Lu, Zachary Shahn, Daby Sow, Finale Doshi-Velez, and H Lehman Li-wei · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Earlier work this paper cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Precision medicine, AI, and the future of personalized health care
Kevin B Johnson, Wei-Qi Wei, Dilhan Weeraratne, Mark E Frisse, Karl Misulis, Kyu Rhee, Juan Zhao, and Jane L Snowdon · 2021
Earlier work this paper cites.
The real threat of deepfake pornography: A review of canadian policy
Vasileia Karasavva and Aalia Noorbhai · 2021
Earlier work this paper cites.
AI and the Everything in the Whole Wide World Benchmark
Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada · 2021
Earlier work this paper cites.
WinoGrande: an adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Earlier work this paper cites.
How to report and benchmark emerging field-effect transistors
Zhihui Cheng, Chin-Sheng Pang, Peiqi Wang, Son T Le, Yanqing Wu, Davood Shahrjerdi, Iuliana Radu, Max C Lemme, Lian-Mao Peng, Xiangfeng Duan, et al · 2022
Cited alongside, same era.
Good practices and recommendations for using and benchmarking computational metabolomics metabolite annotation tools
Niek F de Jonge, Kevin Mildau, David Meijer, Joris JR Louwen, Christoph Bueschl, Florian Huber, and Justin JJ van der Hooft · 2022
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
FinRL-Meta: Market Environments and Benchmarks for Data-Driven Financial Reinforcement Learning
Xiao-Yang Liu, Ziyi Xia, Jingyang Rui, Jiechao Gao, Hongyang Yang, Ming Zhu, Christina Dan Wang, Zhaoran Wang, and Jian Guo · 2022
Cited alongside, same era.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman · 2022
Cited alongside, same era.
MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification
Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni · 2023
Later among the works it cites.
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples, 2023
Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang · 2023
Later among the works it cites.
Don’t Make Your LLM an Evaluation Benchmark Cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han · 2023
Later among the works it cites.
https://www.euaiact.com/article/51
Art. 51 Classification of General-Purpose AI Models as General-Purpose AI Models with Systemic Risk - EU AI Act — euaiact.com · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data cards: Purposeful and transparent dataset documentation for responsible ai
Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson · 2022
Cited alongside, same era.
Precision farming in modern agriculture
E Fantin Irudaya Raj, M Appadurai, and K Athiappan · 2022
Cited alongside, same era.
Prompt sensitivity of language model for solving programming problems
Atsushi Shirafuji, Takumi Ito, Makoto Morishita, Yuki Nakamura, Yusuke Oda, Jun Suzuki, and Yutaka Watanobe · 2022
Cited alongside, same era.
PDEBench: An Extensive Benchmark for Scientific Machine Learning
Makoto Takamoto, Timothy Praditia, Raphael Leiteritz, Dan MacKinlay, Francesco Alesiani, Dirk Pflüger, and Mathias Niepert · 2022
Cited alongside, same era.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al · 2022
Cited alongside, same era.
SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles
Chejian Xu, Wenhao Ding, Weijie Lyu, Zuxin Liu, Shuai Wang, Yihan He, Hanjiang Hu, Ding Zhao, and Bo Li · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Closest in time.
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Norah Alzahrani, Hisham Abdullah Alyahya, et al · 2024
Closest in time.
Introducing the next generation of Claude
Anthropic · 2024
Closest in time.
Machine learning data practices through a data curation lens: An evaluation framework
Eshta Bhardwaj, Harshit Gujral, Siyi Wu, Ciara Zogheib, Tegan Maharaj, and Christoph Becker · 2024
Closest in time.
Lessons from the Trenches on Reproducible Evaluation of Language Models
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A. Wang, Genta Indra Winata, François Yvon, and Andy Zou · 2024
Closest in time.
A Survey on Evaluation of Large Language Models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie · 2024
Closest in time.
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, et al · 2024
Closest in time.
Red-Teaming for Generative AI: Silver Bullet or Security Theater?
Michael Feffer, Anusha Sinha, Wesley Hanwen Deng, Zachary C. Lipton, and Hoda Heidari · 2024
Closest in time.
Adding a workflow status badge
GitHub · 2024
Closest in time.
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation
Declan Grabb, Max Lamparth, and Nina Vasan · 2024
Closest in time.
Devising ML Metrics
Dan Hendrycks and Thomas Woodside · 2024
Closest in time.
Investigating Data Contamination for Pre-training Language Models
Minhao Jiang, Ken Ziyu Liu, Ming Zhong, Rylan Schaeffer, Siru Ouyang, Jiawei Han, and Sanmi Koyejo · 2024
Closest in time.
REFORMS: Consensus-based Recommendations for Machine-learning-based Science
Sayash Kapoor, Emily M. Cantrell, Kenny Peng, Thanh Hien Pham, Christopher A. Bail, Odd Erik Gundersen, Jake M. Hofman, Jessica Hullman, Michael A. Lones, Momin M. Malik, Priyanka Nanayakkara, Russell A. Poldrack, Inioluwa Deborah Raji, Michael Roberts, Matthew J. Salganik, Marta Serra-Garcia, Brandon M. Stewart, Gilles Vandewiele, and Arvind Narayanan · 2024
Closest in time.
Promises and pitfalls of artificial intelligence for legal applications
Sayash Kapoor, Peter Henderson, and Arvind Narayanan · 2024
Closest in time.
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
Max Lamparth, Anthony Corso, Jacob Ganz, Oriana Skylar Mastro, Jacquelyn Schneider, and Harold Trinkunas · 2024
Closest in time.
Rethinking Model Evaluation as Narrowing The Socio-Technical Gap
Q Vera Liao and Ziang Xiao · 2024
Closest in time.
AgentBench: Evaluating LLMs as Agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang · 2024
Closest in time.
Introducing agentbench v0.2
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang · 2024
Closest in time.
ECBD: Evidence-Centered Benchmark Design for NLP
Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Q Vera Liao, Alexandra Olteanu, and Ziang Xiao · 2024
Closest in time.
The AI Index 2024 Annual Report
Nestor Maslej, Loredana Fattorini, Raymond Perrault, Vanessa Parli, Anka Reuel, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Russell Wald, and Jack Clark · 2024
Closest in time.
Inadequacies of large language model benchmarks in the era of generative artificial intelligence
Timothy R McIntosh, Teo Susnjak, Tong Liu, Paul Watters, and Malka N Halgamuge · 2024
Closest in time.
Position: Technical Research and Talent is Needed for Effective AI Governance
Anka Reuel, Lisa Soder, Benjamin Bucknall, and Trond Arne Undheim · 2024
Closest in time.
Escalation risks from language models in military and diplomatic decision-making
Juan-Pablo Rivera, Gabriel Mukobi, Anka Reuel, Max Lamparth, Chandler Smith, and Jacquelyn Schneider · 2024
Closest in time.
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
Aryan Shrivastava, Jessica Hullman, and Max Lamparth · 2024
Closest in time.
AI Benchmarks: Why GenAI Scoreboards Need an Overhaul
Sumeet Wadhwani · 2024
Closest in time.
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen · 2024
Closest in time.