Fetching the paper…
Reading the bibliography…
Rigorous and reproducible evaluation is critical for assessing the state of the art and for guiding scientific advances in Artificial Intelligence.
The Rating of Chessplayers, Past and Present
A. E. Elo · 1978
Earlier work this paper cites.
Heads and tails: studies of web search with common and rare queries
D. Downey, S. Dumais, and E. Horvitz · 2007
Earlier work this paper cites.
Direct answers for search queries in the long tail
M. S. Bernstein, J. Teevan, S. Dumais, D. Liebling, and E. Horvitz · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books, 2015
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. J. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation, 2017
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Earlier work this paper cites.
On human intellect and machine failures: Troubleshooting integrative machine learning systems
B. Nushi, E. Kamar, E. Horvitz, and D. Kossmann · 2017
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba · 2017
Earlier work this paper cites.
Towards accountable ai: Hybrid human-machine analyses for characterizing system failure
B. Nushi, E. Kamar, and E. Horvitz · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Updates in human-ai teams: Understanding and addressing the performance/compatibility tradeoff
G. Bansal, B. Nushi, E. Kamar, D. S. Weld, W. S. Lasecki, and E. Horvitz · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Distributionally robust language modeling
Y. Oren, S. Sagawa, T. B. Hashimoto, and P. Liang · 2019
Earlier work this paper cites.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2019
Earlier work this paper cites.
Towards backward-compatible representation learning
Y. Shen, Y. Xiong, W. Xia, and S. Soatto · 2020
Earlier work this paper cites.
An empirical analysis of backward compatibility in machine learning systems
M. Srivastava, B. Nushi, E. Kamar, S. Shah, and E. Horvitz · 2020
Earlier work this paper cites.
Designing disaggregated evaluations of ai systems: Choices, considerations, and tradeoffs
S. Barocas, A. Guo, E. Kamar, J. Krones, M. R. Morris, J. W. Vaughan, W. D. Wadsworth, and H. Wallach · 2021
Earlier work this paper cites.
Distributionally robust group backwards compatibility
M. Bertran, N. Martinez, A. Oesterling, and G. Sapiro · 2021
Earlier work this paper cites.
To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making
Z. Buçinca, M. B. Malaya, and K. Z. Gajos · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation, 2021
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, K. McDonell, N. Muennighoff, J. Phang, L. Reynolds, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Understanding failures of deep networks via robust feature extraction
S. Singla, B. Nushi, S. Shah, E. Kamar, and E. Horvitz · 2021
Earlier work this paper cites.
Backward-compatible prediction updates: A probabilistic approach
F. Träuble, J. Von Kügelgen, M. Kleindessner, F. Locatello, B. Schölkopf, and P. Gehler · 2021
Earlier work this paper cites.
Why is winoground hard? investigating failures in visuolinguistic compositionality
A. Diwan, L. Berry, E. Choi, D. Harwath, and K. Mahowald · 2022
Earlier work this paper cites.
Domino: Discovering systematic errors with cross-modal embeddings
S. Eyuboglu, M. Varma, K. K. Saab, J. Delbrouck, C. Lee-Messer, J. Dunnmon, J. Zou, and C. Ré · 2022
Earlier work this paper cites.
Who goes first? influences of human-ai workflow on decision making in clinical imaging
R. Fogliato, S. Chappidi, M. Lungren, P. Fisher, D. Wilson, M. Fitzke, M. Parkinson, E. Horvitz, K. Inkpen, and B. Nushi · 2022
Earlier work this paper cites.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
T. Hartvigsen, S. Gabriel, H. Palangi, M. Sap, D. Ray, and E. Kamar · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. R. Joty, and E. Hoque · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Earlier work this paper cites.
Learning from few examples: A summary of approaches to few-shot learning
A. Parnami and M. Lee · 2022
Earlier work this paper cites.
Overreliance on ai literature review
S. Passi and M. Vorvoreanu · 2022
Earlier work this paper cites.
Pddl planning with pretrained large language models
T. Silver, V. Hariprasad, R. S. Shuttleworth, N. Kumar, T. Lozano-Pérez, and L. P. Kaelbling · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Earlier work this paper cites.
Reclip: A strong zero-shot baseline for referring expression comprehension, 2022
S. Subramanian, W. Merrill, T. Darrell, M. Gardner, S. Singh, and A. Rohrbach · 2022
Earlier work this paper cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross · 2022
Earlier work this paper cites.
Musique: Multihop questions via single-hop question composition
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal · 2022
Earlier work this paper cites.
Cpt: Colorful prompt tuning for pre-trained vision-language models, 2022
Y. Yao, A. Zhang, Z. Zhang, Z. Liu, T.-S. Chua, and M. Sun · 2022
Earlier work this paper cites.
Y. Zhang, K. Zhou, and Z. Liu · 2022
Earlier work this paper cites.
MEGA: multilingual evaluation of generative AI
K. Ahuja, H. Diddee, R. Hada, M. Ochieng, K. Ramesh, P. Jain, A. U. Nambi, T. Ganu, S. Segal, M. Ahmed, K. Bali, and S. Sitaram · 2023
Earlier work this paper cites.
Megaverse: Benchmarking large language models across languages, modalities, models and tasks
S. Ahuja, D. Aggarwal, V. Gumma, I. Watts, A. Sathe, M. Ochieng, R. Hada, P. Jain, M. Axmed, K. Bali, et al · 2023
Cited alongside, same era.
How is chatgpt’s behavior changing over time?
L. Chen, M. Zaharia, and J. Zou · 2023
Cited alongside, same era.
Egoplan-bench: Benchmarking egocentric embodied planning with multimodal large language models
Y. Chen, Y. Ge, Y. Ge, M. Ding, B. Li, R. Wang, R. Xu, Y. Shan, and X. Liu · 2023
Cited alongside, same era.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models
J. Cho, A. Zala, and M. Bansal · 2023
Cited alongside, same era.
Investigating data contamination in modern benchmarks for large language models
Understanding the impacts of language technologies’ performance disparities on african american language speakers
J. Cunningham, S. L. Blodgett, M. Madaio, H. D. Iii, C. Harrington, and H. Wallach · 2024
Closest in time.
Constat: Performance-based contamination detection in large language models
J. Dekoninck, M. N. Müller, and M. Vechev · 2024
Closest in time.
Multilingual jailbreak challenges in large language models
Y. Deng, W. Zhang, S. J. Pan, and L. Bing · 2024
Closest in time.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Closest in time.
Muscle: A model update strategy for compatible llm evolution
J. Echterhoff, F. Faghri, R. Vemulapalli, T.-Y. Hu, C.-L. Li, O. Tuzel, and H. Pouransari · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Deng, Y. Zhao, X. Tang, M. Gerstein, and A. Cohan · 2023
Cited alongside, same era.
Data contamination quiz: A tool to detect and estimate contamination in large language models
S. Golchin and M. Surdeanu · 2023
Cited alongside, same era.
Holistic evaluation of text-to-image models
T. Lee, M. Yasunaga, C. Meng, Y. Mai, J. S. Park, A. Gupta, Y. Zhang, D. Narayanan, H. Teufel, M. Bellagente, M. Kang, T. Park, J. Leskovec, J.-Y. Zhu, F.-F. Li, J. Wu, S. Ermon, and P. S. Liang · 2023
Cited alongside, same era.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. A. Cosgrove, C. D. Manning, C. Re, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. WANG, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. S. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. A. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda · 2023
Cited alongside, same era.
Visual spatial reasoning
F. Liu, G. Emerson, and N. Collier · 2023
Cited alongside, same era.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, K. Chen, and D. Lin · 2023
Cited alongside, same era.
Egoschema: A diagnostic benchmark for very long-form video language understanding
K. Mangalam, R. Akshulakov, and J. Malik · 2023
Cited alongside, same era.
A robust backward compatibility metric for model retraining
R. Matsuno and K. Sakuma · 2023
Cited alongside, same era.
Closest in time.
Open llm leaderboard v2
C. Fourrier, N. Habib, A. Lozovskaya, K. Szafer, and T. Wolf · 2024
Closest in time.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang, et al · 2024
Closest in time.
Human feedback is not gold standard
T. Hosking, P. Blunsom, and M. Bartolo · 2024
Closest in time.
As generative models improve, people adapt their prompts
E. Jahani, B. S. Manning, J. Zhang, H.-Y. TuYe, M. Alsobay, C. Nicolaides, S. Suri, and D. Holtz · 2024
Closest in time.
D \ \backslash ’ej \ \backslash a vu memorization in vision-language models
B. Jayaraman, C. Guo, and K. Chaudhuri · 2024
Closest in time.
Apeer: Automatic prompt engineering enhances large language model reranking
C. Jin, H. Peng, S. Zhao, Z. Wang, W. Xu, L. Han, J. Zhao, K. Zhong, S. Rajasekaran, and D. N. Metaxas · 2024
Closest in time.
Benchmarking cognitive biases in large language models as evaluators
R. Koo, M. Lee, V. Raheja, J. I. Park, Z. M. Kim, and D. Kang · 2024
Closest in time.
Same task, more tokens: the impact of input length on the reasoning performance of large language models, 2024
M. Levy, A. Jacoby, and Y. Goldberg · 2024
Closest in time.
Culturellm: Incorporating cultural differences into large language models
C. Li, M. Chen, J. Wang, S. Sitaram, and X. Xie · 2024
Closest in time.
Mvbench: A comprehensive multi-modal video understanding benchmark
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo, et al · 2024
Closest in time.
Long-context llms struggle with long in-context learning, 2024
T. Li, G. Zhang, Q. D. Do, X. Yue, and W. Chen · 2024
Closest in time.
Autobencher: Creating salient, novel, difficult datasets for language models
X. L. Li, E. Z. Liu, P. Liang, and T. Hashimoto · 2024
Closest in time.
Ux matters: The critical role of ux in responsible ai
Q. V. Liao, M. Vorvoreanu, H. Subramonyam, and L. Wilcox · 2024
Closest in time.
Ecbd: Evidence-centered benchmark design for nlp
Y. L. Liu, S. L. Blodgett, J. C. K. Cheung, Q. V. Liao, A. Olteanu, and Z. Xiao · 2024
Closest in time.
Stable bias: Evaluating societal representations in diffusion models
S. Luccioni, C. Akiki, M. Mitchell, and Y. Jernite · 2024
Closest in time.
“one-size-fits-all”? examining expectations around what constitute “fair” or “good” nlg system behaviors
L. Lucy, S. L. Blodgett, M. Shokouhi, H. Wallach, and A. Olteanu · 2024
Closest in time.
(why) is my prompt getting worse? rethinking regression testing for evolving llm apis
W. Ma, C. Yang, and C. Kästner · 2024
Closest in time.
Worldbench: Quantifying geographic disparities in llm factual recall
M. Moayeri, E. Tabassi, and S. Feizi · 2024
Closest in time.
Evaluating cognitive maps and planning in large language models with cogeval
I. Momennejad, H. Hasanbeig, F. Vieira Frujeri, H. Sharma, N. Jojic, H. Palangi, R. Ness, and J. Larson · 2024
Closest in time.
OpenAI · 2024
Closest in time.
Perception test: A diagnostic benchmark for multimodal video models
V. Patraucean, L. Smaira, A. Gupta, A. Recasens, L. Markeeva, D. Banarse, S. Koppula, M. Malinowski, Y. Yang, C. Doersch, et al · 2024
Closest in time.
Benchmark agreement testing done right: A guide for llm benchmark evaluation
Y. Perlitz, A. Gera, O. Arviv, A. Yehudai, E. Bandel, E. Shnarch, M. Shmueli-Scheuer, and L. Choshen · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Image understanding benchmark
M. Research · 2024
Closest in time.
Prompts as programs: A structure-aware approach to efficient compile-time prompt optimization
T. Schnabel and J. Neville · 2024
Closest in time.
Robovqa: Multimodal long-horizon reasoning for robotics
P. Sermanet, T. Ding, J. Zhao, F. Xia, D. Dwibedi, K. Gopalakrishnan, C. Chan, G. Dulac-Arnold, S. Maddineni, N. J. Joshi, P. Florence, W. Han, R. Baruch, Y. Lu, S. Mirchandani, P. Xu, P. Sanketi, K. Hausman, I. Shafran, B. Ichter, and Y. Cao · 2024
Closest in time.
Introducing v0. 5 of the ai safety benchmark from mlcommons
B. Vidgen, A. Agrawal, A. M. Ahmed, V. Akinwande, N. Al-Nuaimi, N. Alfaraj, E. Alhajjar, L. Aroyo, T. Bavalatti, B. Blili-Hamelin, et al · 2024
Closest in time.
Is a picture worth a thousand words? delving into spatial reasoning for vision language models, 2024
J. Wang, Y. Ming, Z. Shi, V. Vineet, X. Wang, and N. Joshi · 2024
Closest in time.
Large language models are not fair evaluators
P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, L. Kong, Q. Liu, T. Liu, and Z. Sui · 2024
Closest in time.
Benchmark self-evolving: A multi-agent framework for dynamic llm evaluation
S. Wang, Z. Long, Z. Fan, Z. Wei, and X. Huang · 2024
Closest in time.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
T. Xie, X. Qi, Y. Zeng, Y. Huang, U. M. Sehwag, K. Huang, L. He, B. Wei, D. Li, Y. Sheng, et al · 2024
Closest in time.
X. Yuan, J. Li, D. Wang, Y. Chen, X. Mao, L. Huang, H. Xue, W. Wang, K. Ren, and J. Wang · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2024
Closest in time.
Attention satisfies: A constraint-satisfaction lens on factual errors of language models
M. Yüksekgönül, V. Chandrasekaran, E. Jones, S. Gunasekar, R. Naik, H. Palangi, E. Kamar, and B. Nushi · 2024
Closest in time.
Cybench: A framework for evaluating cybersecurity capabilities and risk of language models
A. K. Zhang, N. Perry, R. Dulepet, E. Jones, J. W. Lin, J. Ji, C. Menders, G. Hussein, S. Liu, D. Jasper, et al · 2024
Closest in time.
Inherent trade-offs between diversity and stability in multi-task benchmark
G. Zhang and M. Hardt · 2024
Closest in time.
A careful examination of large language model performance on grade school arithmetic
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, D. Slack, Q. Lyu, et al · 2024
Closest in time.
Natural plan: Benchmarking llms on natural language planning
H. S. Zheng, S. Mishra, H. Zhang, X. Chen, M. Chen, A. Nova, L. Hou, H.-T. Cheng, Q. V. Le, E. H. Chi, et al · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.
Dyval: Dynamic evaluation of large language models for reasoning tasks
K. Zhu, J. Chen, J. Wang, N. Z. Gong, D. Yang, and X. Xie · 2024
Closest in time.