Fetching the paper…
Reading the bibliography…
Numerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation.
Rouge: A Package for Automatic Evaluation of Summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Pdffigures 2.0: Mining figures from research papers
Clark, C. and Divvala, S · 2016
Earlier work this paper cites.
River classification as a geographic tool in the age of big data and global change
Praskievicz, S · 2018
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2020
Earlier work this paper cites.
Artificial intelligence in medicine for chronic disease classification using machine learning
Rakhimov, M., Akhmadjonov, R., and Javliev, S · 2022
Earlier work this paper cites.
Falcon-40B: an open large language model with state-of-the-art performance, 2023
Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, E., Heslow, D., Launay, J., Malartic, Q., Noune, B., Pannier, B., and Penedo, G · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Team, G., et al · 2023
Earlier work this paper cites.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Earlier work this paper cites.
Traineragent: Customizable and efficient model training through llm-powered multi-agent system
Li, H., Jiang, H., Zhang, T., Yu, Z., Yin, A., Cheng, H., Fu, S., Zhang, Y., and He, W · 2023
Earlier work this paper cites.
Vila: On pre-training for visual language models, 2023
Lin, J., Yin, H., Ping, W., Lu, Y., Molchanov, P., Tao, A., Mao, H., Kautz, J., Shoeybi, M., and Han, S · 2023
Earlier work this paper cites.
Papermage: A unified toolkit for processing, representing, and manipulating visually-rich scientific documents
Lo, K., Shen, Z., Newman, B., Chang, J. Z., Authur, R., Bransom, E., Candra, S., Chandrasekhar, Y., Huff, R., Kuehl, B., et al · 2023
Earlier work this paper cites.
Identifying social norm violation in movie plots: from borat to american pie
Neuman, Y., Cohen, Y., and Yin, W · 2023
Earlier work this paper cites.
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Nlpbench: Evaluating large language models on solving nlp problems
Song, L., Zhang, J., Cheng, L., Zhou, P., Zhou, T., and Li, I · 2023
Cited alongside, same era.
Ml-bench: Evaluating large language models and agents for machine learning tasks on repository-level code
Tang, X., Liu, Y., Cai, Z., Shao, Y., Lu, J., Zhang, Y., Deng, Z., Hu, H., An, K., Huang, R., et al · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2023
Cited alongside, same era.
Llama 3 model card, 2024
AI@Meta · 2024
Cited alongside, same era.
Can large language models unlock novel scientific research ideas?
Kumar, S., Ghosal, T., Goyal, V., and Ekbal, A · 2024
Closest in time.
Biomistral: A collection of open-source pretrained large language models for medical domains
Labrak, Y., Bazoge, A., Morin, E., Gourraud, P.-A., Rouvier, M., and Dufour, R · 2024
Closest in time.
Mlr-copilot: Autonomous machine learning research based on large language models agents
Li, R., Patel, T., Wang, Q., and Du, X · 2024
Closest in time.
Can large language models provide useful feedback on research papers? a large-scale empirical analysis
Liang, W., Zhang, Y., Cao, H., Wang, B., Ding, D. Y., Yang, X., Vodrahalli, K., He, S., Smith, D. S., Yin, Y., et al · 2024
Closest in time.
Lost in the middle: How language models use long contexts
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automated focused feedback generation for scientific writing assistance
Chamoun, E., Schlichktrull, M., and Vlachos, A · 2024
Cited alongside, same era.
Mle-bench: Evaluating machine learning agents on machine learning engineering
Chan, J. S., Chowdhury, N., Jaffe, O., Aung, J., Sherburn, D., Mays, E., Starace, G., Liu, K., Maksin, L., Patwardhan, T., et al · 2024
Cited alongside, same era.
Llms assist NLP researchers: Critique paper (meta-)reviewing
Du, J., Wang, Y., Zhao, W., Deng, Z., Liu, S., Lou, R., Zou, H. P., Venkit, P. N., Zhang, N., Srinath, M., Zhang, H. R., Gupta, V., Li, Y., Li, T., Wang, F., Liu, Q., Liu, T., Gao, P., Xia, C., Xing, C., Cheng, J., Wang, Z., Su, Y., Shah, R. S., Guo, R., Gu, J., Li, H., Wei, K., Wang, Z., Cheng, L., Ranathunga, S., Fang, M., Fu, J., Liu, F., Huang, R., Blanco, E., Cao, Y., Zhang, R., Yu, P. S., and Yin, W · 2024
Cited alongside, same era.
Reviewer2: Optimizing review generation through prompt generation
Gao, Z., Brantley, K., and Joachims, T · 2024
Cited alongside, same era.
Olmo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., Arora, S., Atkinson, D., Authur, R., Chandu, K., Cohan, A., Dumas, J., Elazar, Y., Gu, Y., Hessel, J., Khot, T., Merrill, W., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Pyatkin, V., Ravichander, A., Schwenk, D., Shah, S., Smith, W., Subramani, N., Wortsman, M., Dasigi, P., Lambert, N., Richardson, K., Dodge, J., Lo, K., Soldaini, L., Smith, N. A., and Hajishirzi, H · 2024
Cited alongside, same era.
Adaptive and explainable margin trading via large language models on portfolio management
Gu, J., Ye, J., Yin, W., and Wang, G · 2024
Cited alongside, same era.
Mlagentbench: Evaluating language agents on machine learning experimentation
Huang, Q., Vora, J., Liang, P., and Leskovec, J · 2024
Cited alongside, same era.
Closest in time.
The AI Scientist: Towards fully automated open-ended scientific discovery
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D · 2024
Closest in time.
Introducing llama 3.1: Our most capable models to date
MetaAI · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Closest in time.
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers
Si, C., Yang, D., and Hashimoto, T · 2024
Closest in time.
Leave no document behind: Benchmarking long-context llms with extended multi-doc qa
Wang, M., Chen, L., Fu, C., Liao, S., Zhang, X., Wu, B., Yu, H., Xu, N., Zhang, L., Luo, R., et al · 2024
Closest in time.
Cycleresearcher: Improving automated research via automated review
Weng, Y., Zhu, M., Bao, G., Zhang, H., Wang, J., Zhang, Y., and Yang, L · 2024
Closest in time.
Openai o3-mini
OpenAI · 2025
Closest in time.
Deepreview: Improving llm-based paper review with human-like deep thinking process
Zhu, M., Weng, Y., Yang, L., and Zhang, Y · 2025
Closest in time.