Fetching the paper…
Reading the bibliography…
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models, 2022
Pan, A., Bhatia, K., and Steinhardt, J · 2022
Earlier work this paper cites.
Scienceworld: Is Your Agent Smarter than a 5th Grader?
Wang, R., Jansen, P., Côté, M.-A., and Ammanabrolu, P · 2022
Earlier work this paper cites.
Can Large Language Models Be an Alternative to Human Evaluations?, 2023
Chiang, C.-H. and yi Lee, H · 2023
Earlier work this paper cites.
GPTScore: Evaluate as You Desire
Fu, J., Ng, S.-K., Jiang, Z., and Liu, P · 2023
Earlier work this paper cites.
Preparedness Framework, December 2023
OpenAI · 2023
Earlier work this paper cites.
ARB: Advanced Reasoning Benchmark for Large Language Models, 2023
Sawada, T., Paleka, D., Havrilla, A., Tadepalli, P., Vidas, P., Kranias, A., Nay, J. J., Gupta, K., and Komatsuzaki, A · 2023
Earlier work this paper cites.
ReAct: Synergizing Reasoning and Acting in Language Models, March 2023
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2023
Earlier work this paper cites.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Earlier work this paper cites.
Stochastic Interpolants with Data-Dependent Couplings
Albergo, M. S., Goldstein, M., Boffi, N. M., Ranganath, R., and Vanden-Eijnden, E · 2024
Earlier work this paper cites.
MLE-bench: Evaluating machine learning agents on machine learning engineering
Chan, J. S., Chowdhury, N., Jaffe, O., Aung, J., Sherburn, D., Mays, E., Starace, G., Liu, K., Maksin, L., Patwardhan, T., et al · 2024
Earlier work this paper cites.
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
Chen, D., Chen, R., Zhang, S., Liu, Y., Wang, Y., Zhou, H., Zhang, Q., Wan, Y., Zhou, P., and Sun, L · 2024
Earlier work this paper cites.
Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
Frans, K., Park, S., Abbeel, P., and Levine, S · 2024
Earlier work this paper cites.
All-in-one simulation-based inference
Gloeckler, M., Deistler, M., Weilbach, C. D., Wood, F., and Macke, J. H · 2024
Earlier work this paper cites.
Frontier Safety Framework, May 2024
Google DeepMind · 2024
Cited alongside, same era.
Introducing BigLaw Bench, August 2024
Harvey Team · 2024
Cited alongside, same era.
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Huang, Q., Vora, J., Liang, P., and Leskovec, J · 2024
Cited alongside, same era.
Jansen, P., Côté, M.-A., Khot, T., Bransom, E., Mishra, B. D., Majumder, B. P., Tafjord, O., and Clark, P · 2024
Cited alongside, same era.
What Will My Model Forget? Forecasting Forgotten Examples in Language Model Refinement
Jin, X. and Ren, X · 2024
Cited alongside, same era.
DSBench: How Far Are Data Science Agents to Becoming Data Science Experts?
LCA-on-the-Line: Benchmarking Out of Distribution Generalization with Class Taxonomies
Shi, J., Gare, G. R., Tian, J., Chai, S., Lin, Z., Vasudevan, A. B., Feng, D., Ferroni, F., and Kong, S · 2024
Later among the works it cites.
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
Si, C., Yang, D., and Hashimoto, T · 2024
Later among the works it cites.
Siegel, Z. S., Kapoor, S., Nagdir, N., Stroebl, B., and Narayanan, A · 2024
Later among the works it cites.
SAPG: Split and Aggregate Policy Gradients
Singla, J., Agarwal, A., and Pathak, D · 2024
Later among the works it cites.
AI Sandbagging: Language Models Can Strategically Underperform on Evaluations
van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., and Ward, F. R · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jing, L., Huang, Z., Wang, X., Yao, W., Yu, W., Ma, K., Zhang, H., Du, X., and Yu, D · 2024
Cited alongside, same era.
RewardBench: Evaluating Reward Models for Language Modeling, 2024
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N. A., and Hajishirzi, H · 2024
Cited alongside, same era.
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J. K., and Mihalcea, R · 2024
Cited alongside, same era.
Test-Time Model Adaptation with Only Forward Passes
Niu, S., Miao, C., Chen, G., Wu, P., and Zhao, P · 2024
Cited alongside, same era.
Challenges in Training PINNs: A Loss Landscape Perspective
Rathore, P., Lei, W., Frangella, Z., Lu, L., and Udell, M · 2024
Cited alongside, same era.
Stay on Topic with Classifier-Free Guidance
Sanchez, G., Spangher, A., Fan, H., Levi, E., and Biderman, S · 2024
Cited alongside, same era.
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
Schlarmann, C., Singh, N. D., Croce, F., and Hein, M · 2024
Cited alongside, same era.
Later among the works it cites.
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Wijk, H., Lin, T., Becker, J., Jawhar, S., Parikh, N., Broadley, T., Chan, L., Chen, M., Clymer, J., Dhyani, J., et al · 2024
Later among the works it cites.
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
Wolczyk, M., Cupial, B., Ostaszewski, M., Bortkiewicz, M., Zajac, M., Pascanu, R., Kucinski, L., and Milos, P · 2024
Later among the works it cites.
Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance Constraints
Xia, X., Liu, J., Zhang, S., Wu, Q., Wei, H., and Liu, T · 2024
Later among the works it cites.
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
Zhao, B., Hajishirzi, H., and Cao, Q · 2024
Later among the works it cites.
Agent-as-a-Judge: Evaluate Agents with Agents
Zhuge, M., Zhao, C., Ashley, D., Wang, W., Khizbullin, D., Xiong, Y., Liu, Z., Chang, E., Krishnamoorthi, R., Tian, Y., et al · 2024
Later among the works it cites.
Responsible Scaling Policy
Anthropic · 2025
Closest in time.
Specification Gaming: The Flip Side of AI Ingenuity, 2024
DeepMind · 2025
Closest in time.
Inspect, 2025
UK AI Safety Institute · 2025
Closest in time.