Fetching the paper…
Reading the bibliography…
Scientific experimentation, a cornerstone of human progress, demands rigor in reliability, methodical control, and interpretability to yield meaningful results.
Using context to build rigor: Application to two hermeneutic phenomenological studies
Armour, M., Rivaux, S. L., and Bell, H · 2009
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021a
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Earlier work this paper cites.
Getting rigorous with scientific rigor
Hofseth, L. J · 2018
Earlier work this paper cites.
What is research rigor? lessons for a transdiscipline
Gill, T. and Gill, T · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Plant growth: the what, the how, and the why
Hilty, J., Muller, B., Pantin, F., and Leuzinger, S · 2021
Earlier work this paper cites.
Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., Do, Q. V., Xu, Y., and Fung, P · 2023
Earlier work this paper cites.
Ai for science: An emerging agenda, 2023
Berens, P., Cranmer, K., Lawrence, N. D., von Luxburg, U., and Montgomery, J · 2023
Earlier work this paper cites.
Mathematical capabilities of chatgpt
Frieder, S., Pinchetti, L., , Griffiths, R.-R., Salvatori, T., Lukasiewicz, T., Petersen, P., and Berner, J · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Earlier work this paper cites.
Automated scientific discovery: From equation discovery to autonomous discovery systems, 2023
Kramer, S., Cerrato, M., Džeroski, S., and King, R · 2023
Earlier work this paper cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
Accelerating science with human-aware artificial intelligence
Sourati, J. and Evans, J. A · 2023
Earlier work this paper cites.
Automl-gpt: Automatic machine learning with gpt, 2023
Zhang, S., Gong, C., Wu, L., Liu, X., and Zhou, M · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Cited alongside, same era.
Litllm: A toolkit for scientific literature review
Agarwal, S., Laradji, I. H., Charlin, L., and Pal, C · 2024
Cited alongside, same era.
Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku
Anthropic · 2024
Cited alongside, same era.
Conversation patterns
AutoGen · 2024
Cited alongside, same era.
Knowledge graph extraction from total synthesis documents
The hallmark effect: Supporting provenance and transparent use of large language models in writing with interactive visualization
Hoque, M. N., Mashiat, T., Ghai, B., Shelton, C. D., Chevalier, F., Kraus, K., and Elmqvist, N · 2024
Later among the works it cites.
Mlagentbench: Evaluating language agents on machine learning experimentation, 2024
Huang, Q., Vora, J., Liang, P., and Leskovec, J · 2024
Later among the works it cites.
The impact of reasoning step length on large language models
Jin, M., Yu, Q., Shu, D., Zhao, H., Hua, W., Meng, Y., Zhang, Y., and Du, M · 2024
Later among the works it cites.
The ai scientist: Towards fully automated open-ended scientific discovery
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D · 2024
Later among the works it cites.
Large language models as biomedical hypothesis generators: A comprehensive evaluation, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bran, A. M., Jončev, Z., and Schwaller, P · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Chen, Z., Chen, S., Ning, Y., Zhang, Q., Wang, B., Yu, B., Li, Y., Liao, Z., Wei, C., Lu, Z., et al · 2024
Cited alongside, same era.
Language models as science tutors, 2024
Chevalier, A., Geng, J., Wettig, A., Chen, H., Mizera, S., Annala, T., Aragon, M. J., Fanlo, A. R., Frieder, S., Machado, S., Prabhakar, A., Thieu, E., Wang, J. T., Wang, Z., Wu, X., Xia, M., Xia, W., Yu, J., Zhu, J.-J., Ren, Z. J., Arora, S., and Chen, D · 2024
Cited alongside, same era.
The faiss library
Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazaré, P.-E., Lomeli, M., Hosseini, L., and Jégou, H · 2024
Cited alongside, same era.
Magentic-one: A generalist multi-agent system for solving complex tasks
Fourney, A., Bansal, G., Mozannar, H., Tan, C., Salinas, E., Niedtner, F., Proebsting, G., Bassman, G., Gerrits, J., Alber, J., et al · 2024
Cited alongside, same era.
Sciagents: Automating scientific discovery through multi-agent intelligent graph reasoning, 2024
Ghafarollahi, A. and Buehler, M. J · 2024
Cited alongside, same era.
Qi, B., Zhang, K., Tian, K., Li, H., Chen, Z.-R., Zeng, S., Hua, E., Jinfang, H., and Zhou, B · 2024
Later among the works it cites.
Ai-driven review systems: Evaluating llms in scalable and bias-aware academic reviews, 2024
Tyser, K., Segev, B., Longhitano, G., Zhang, X.-Y., Meeks, Z., Lee, J., Garg, U., Belsten, N., Shporer, A., Udell, M., Te’eni, D., and Drori, I · 2024
Later among the works it cites.
Knowledge conflicts for llms: A survey
Xu, R., Qi, Z., Guo, Z., Wang, C., Wang, H., Zhang, Y., and Xu, W · 2024
Later among the works it cites.
Swe-agent: Agent-computer interfaces enable automated software engineering
Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., and Press, O · 2024
Later among the works it cites.
Hypothesis generation with large language models
Zhou, Y., Liu, H., Srivastava, T., Mei, H., and Tan, C · 2024
Later among the works it cites.
Agent laboratory: Using llm agents as research assistants
Schmidgall, S., Su, Y., Wang, Z., Sun, X., Wu, J., Yu, X., Liu, J., Liu, Z., and Barsoum, E · 2025
Closest in time.
Dolphin: Closed-loop open-ended auto-research through thinking, practice, and feedback, 2025
Yuan, J., Yan, X., Shi, B., Chen, T., Ouyang, W., Zhang, B., Bai, L., Qiao, Y., and Zhou, B · 2025
Closest in time.
Nobel turing challenge: creating the engine for scientific discovery
Kitano, H · 2056
Closest in time.