Fetching the paper…
Reading the bibliography…
AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research.
A Step Toward Quantifying Independently Reproducible Machine Learning Research, September 2019
Edward Raff · 1909
Earlier work this paper cites.
Prediction Interval: What to Expect When You’re Expecting … A Replication
Jeffrey R. Spence and David J. Stanley · 1932
Earlier work this paper cites.
Push button replication: Is impact evaluation evidence for international development verifiable?
Benjamin D. K. Wood, Rui Müller, and Annette N. Brown · 1932
Earlier work this paper cites.
WaveLab and Reproducible Research
Jonathan B. Buckheit and David L. Donoho · 1995
Earlier work this paper cites.
Lessons from the JMCB Archive
B. D. McCullough, Kerry Anne McGeary, and Teresa D. Harrison · 2006
Earlier work this paper cites.
Repeatability of published microarray gene expression analyses
John P. A. Ioannidis, David B. Allison, Catherine A. Ball, Issa Coulibaly, Xiangqin Cui, Aedín C. Culhane, Mario Falchi, Cesare Furlanello, Laurence Game, Giuseppe Jurman, Jon Mangion, Tapan Mehta, Michael Nitzberg, Grier P. Page, Enrico Petretto, and Vera van Noort · 2009
Earlier work this paper cites.
Recommendations for utilizing and reporting population genetic analyses: the reproducibility of genetic clustering using the program structure
Kimberly J. Gilbert, Rose L. Andrew, Dan G. Bock, Michelle T. Franklin, Nolan C. Kane, Jean-Sébastien Moore, Brook T. Moyers, Sébastien Renaut, Diana J. Rennison, Thor Veen, and Timothy H. Vines · 2012
Earlier work this paper cites.
Assessing the reproducibility of discriminant function analyses
Rose L. Andrew, Arianne Y.K. Albert, Sebastien Renaut, Diana J. Rennison, Dan G. Bock, and Tim Vines · 2015
Earlier work this paper cites.
Repeatability in computer systems research
Christian Collberg and Todd A. Proebsting · 2016
Earlier work this paper cites.
Computational Reproducibility via Containers in Psychology
April Clyburne-Sherin, Xu Fei, and Seth Ariel Green · 2018
Earlier work this paper cites.
How to make replication the norm
Paul Gertler, Sebastian Galiani, and Mauricio Romero · 2018
Earlier work this paper cites.
Data availability, reusability, and analytic reproducibility: evaluating the impact of a mandatory open data policy at the journal Cognition
Tom E. Hardwicke, Maya B. Mathur, Kyle MacDonald, Gustav Nilsonne, George C. Banks, Mallory C. Kidwell, Alicia Hofelich Mohr, Elizabeth Clayton, Erica J. Yoon, Michael Henry Tessler, Richie L. Lenne, Sara Altman, Bria Long, and Michael C. Frank · 2018
Earlier work this paper cites.
Computational reproducibility in geoscientific papers: Insights from a series of studies with geoscientists and a reproduction study
Markus Konkol, Christian Kray, and Max Pfeiffer · 2018
Earlier work this paper cites.
Data sharing and reanalysis of randomized controlled trials in leading biomedical journals with a full data sharing policy: survey of studies published in The BMJ and PLOS Medicine
Florian Naudet, Charlotte Sakarovitch, Perrine Janiaud, Ioana Cristea, Daniele Fanelli, David Moher, and John P. A. Ioannidis · 2018
Earlier work this paper cites.
Data Access, Transparency, and Replication: New Insights from the Political Behavior Literature
Daniel Stockemer, Sebastian Koehler, and Tobias Lentz · 2018
Earlier work this paper cites.
Successes and Struggles with Computational Reproducibility: Lessons from the Fragile Families Challenge
David M. Liu and Matthew J. Salganik · 2019
Earlier work this paper cites.
Reproducibility and Replicability in Science
National Academies of Sciences Engineering and Medicine · 2019
Earlier work this paper cites.
Analysis of Open Data and Computational Reproducibility in Registered Reports in Psychology
Pepijn Obels, Daniël Lakens, Nicholas A. Coles, Jaroslav Gottfried, and Seth A. Green · 2020
Earlier work this paper cites.
A Systematic Review of Reproducibility Research in Natural Language Processing
Anya Belz, Shubham Agarwal, Anastasia Shimorina, and Ehud Reiter · 2021
Cited alongside, same era.
Evaluating Large Language Models Trained on Code, July 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba · 2021
Cited alongside, same era.
Improving Reproducibility in Machine Learning Research(A Report from the NeurIPS 2019 Reproducibility Program)
Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Lariviere, Alina Beygelzimer, Florence d’Alche Buc, Emily Fox, and Hugo Larochelle · 2021
Cited alongside, same era.
AI and the Everything in the Whole Wide World Benchmark, 2021
Yifeng He, Ethan Wang, Yuyang Rong, Zifei Cheng, and Hao Chen · 2024
Closest in time.
AI Agents That Matter, July 2024
Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan · 2024
Closest in time.
Everyone’s talking about Sakana’s AI scientist. But no-one’s answering the big question: is its output good?, August 2024
Jimmy Koppel · 2024
Closest in time.
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, August 2024
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha · 2024
Closest in time.
OpenEQA: Embodied Question Answering in the Era of Foundation Models
Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma, Vincent Berges, Shiqi Zhang, Pulkit Agrawal, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton, Alexander Sax, and Aravind Rajeswaran · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna · 2021
Cited alongside, same era.
MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation, 2022
Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda · 2022
Cited alongside, same era.
Competition-Level Code Generation with AlphaCode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals · 2022
Cited alongside, same era.
[Re] Exploring the Representation of Word Meanings in Context
Matteo Brivio and Çağrı Çöltekin · 2023
Cited alongside, same era.
Benchmarking Large Language Models As AI Research Agents, October 2023
Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec · 2023
Cited alongside, same era.
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, October 2023
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
Evaluating LLMs is a minefield, October 2023
Sayash Kapoor and Arvind Narayanan · 2023
Cited alongside, same era.
[Re] A Reproduction of Automatic Multi-Label Prompting: Simple and Interpretable Few-Shot Classification
Victor Livernoche and Vidya Sujaya · 2023
Cited alongside, same era.
AutoGPT, September 2024
Significant Gravitas · 2023
Cited alongside, same era.
Closest in time.
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models, July 2024
Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Bhavana Dalvi Mishra, Abhijeetsingh Meena, Aryan Prakhar, Tirth Vora, Tushar Khot, Ashish Sabharwal, and Peter Clark · 2024
Closest in time.
Vivaria, 2024
METR · 2024
Closest in time.
CiteME: Can Language Models Accurately Cite Scientific Claims?, July 2024
Ori Press, Andreas Hochlehnert, Ameya Prabhu, Vishaal Udandarao, Ofir Press, and Matthias Bethge · 2024
Closest in time.
Computational Reproducibility in Finance: Evidence from 1,000 Tests, February 2024
Christophe Pérignon, Olivier Akmansoy, Christophe Hurlin, Anna Dreber, Felix Holzmeister, Juergen Huber, Magnus Johannesson, Michael Kirchler, Albert J. Menkveld, Michael Razen, and Utz Weitzel · 2024
Closest in time.
SciCode: A Research Coding Benchmark Curated by Scientists, July 2024
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, Shengyan Liu, Di Luo, Yutao Ma, Hao Tong, Kha Trinh, Chenyu Tian, Zihan Wang, Bohao Wu, Yanyu Xiong, Shengzhu Yin, Minhui Zhu, Kilian Lieret, Yanxin Lu, Genglin Liu, Yufeng Du, Tianhua Tao, Ofir Press, Jamie Callan, Eliu Huerta, and Hao Peng · 2024
Closest in time.
ChartBench: A Benchmark for Complex Visual Reasoning in Charts, June 2024
Zhengzhuo Xu, Sinan Du, Yiyan Qi, Chengjin Xu, Chun Yuan, and Jian Guo · 2024
Closest in time.
SWE-AGENT: Agent-Computer Interfaces Enable Automated Software Engineering
John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press · 2024
Closest in time.
$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, June 2024
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan · 2024
Closest in time.
The Shift from Models to Compound AI Systems, February 2024
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi · 2024
Closest in time.
PyBench: Evaluating LLM Agent on various real-world coding tasks, August 2024
Yaolun Zhang, Yinxu Pan, Yudong Wang, and Jie Cai · 2024
Closest in time.
Computational reproducibility of Jupyter notebooks from biomedical publications
Sheeba Samuel and Daniel Mietchen · 2047
Closest in time.
A large-scale study on research code quality and execution
Ana Trisovic, Matthew K. Lau, Thomas Pasquier, and Mercè Crosas · 2052
Closest in time.
Analytic reproducibility in articles receiving open data badges at the journal Psychological Science
Tom E. Hardwicke, Manuel Bohn, Kyle MacDonald, Emily Hembacher, Michèle B. Nuijten, Benjamin N. Peloquin, Benjamin E. deMayo, Bria Long, Erica J. Yoon, and Michael C. Frank · 2054
Closest in time.