Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly utilized in scientific research assessment, particularly in automated paper review.
Zero: Memory optimizations toward training trillion parameter models, 2020
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 1910
Earlier work this paper cites.
Scientific discovery: Computational explorations of the creative processes
P. Langley · 1987
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Reviewing peer review, 2008
B. Alberts, B. Hanson, and K. L. Kelner · 2008
Earlier work this paper cites.
A dataset of peer reviews (peerread): Collection, insights and nlp applications
D. Kang, W. Ammar, B. Dalvi, M. Van Zuylen, S. Kohlmeier, E. Hovy, and R. Schwartz · 2018
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
ReviewRobot: Explainable paper review generation based on knowledge synthesis
Q. Wang, Q. Zeng, L. Huang, K. Knight, H. Ji, and N. F. Rajani · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al · 2021
Earlier work this paper cites.
Can we automate scientific reviewing?, 2021
W. Yuan, P. Liu, and G. Neubig · 2021
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2022
Earlier work this paper cites.
Citebench: A benchmark for scientific citation text generation
M. Funkquist, I. Kuznetsov, Y. Hou, and I. Gurevych · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. H. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Earlier work this paper cites.
Emergent autonomous scientific research capabilities of large language models
B. Daniil, A., M. Robert, and G. Gabe · 2023
Earlier work this paper cites.
Towards mitigating LLM hallucination via self reflection
Z. Ji, T. Yu, Y. Xu, N. Lee, E. Ishii, and P. Fung · 2023
Earlier work this paper cites.
Summarizing multiple documents with conversational structure for meta-review generation
M. Li, E. Hovy, and J. H. Lau · 2023
Earlier work this paper cites.
A critical examination of the ethics of ai-mediated peer review, 2023
L. A. Schintler, C. L. McNeely, and J. Witte · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Earlier work this paper cites.
Large language models are better reasoners with self-verification
Y. Weng, M. Zhu, F. Xia, B. Li, S. He, S. Liu, B. Sun, K. Liu, and J. Zhao · 2023
Cited alongside, same era.
Large language models for automated open-domain scientific hypotheses discovery
Y. Zonglin, D. Xinya, L. Junxian, Z. Jie, P. Soujanya, and C. Erik · 2023
Cited alongside, same era.
Care: Collaborative ai-assisted reading environment
D. Zyska, N. Dycke, J. Buchmann, I. Kuznetsov, and I. Gurevych · 2023
Cited alongside, same era.
M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kauffmann, J. R. Lee, Y. T. Lee, Y. Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y. Wu, D. Yu, C. Zhang, and Y. Zhang · 2024
Cited alongside, same era.
Openscholar: Synthesizing scientific literature with retrieval-augmented lms, 2024
Prompting llms to compose meta-review drafts from peer-review narratives of scholarly manuscripts
S. K. K. Santu, S. K. Sinha, N. Bansal, A. Knipper, S. Sarkar, J. Salvador, Y. Mahajan, S. Guttikonda, M. Akter, M. Freestone, et al · 2024
Later among the works it cites.
D. Scherbakov, N. Hubig, V. Jansari, A. Bakumenko, and L. A. Lenert · 2024
Later among the works it cites.
H. Su, R. Chen, S. Tang, X. Zheng, J. Li, Z. Yin, W. Ouyang, and N. Dong · 2024
Later among the works it cites.
Ai-driven review systems: evaluating llms in scalable and bias-aware academic reviews
K. Tyser, B. Segev, G. Longhitano, X.-Y. Zhang, Z. Meeks, J. Lee, U. Garg, N. Belsten, A. Shporer, M. Udell, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Asai, J. He, R. Shao, W. Shi, A. Singh, J. C. Chang, K. Lo, L. Soldaini, S. Feldman, M. D’arcy, D. Wadden, M. Latzke, M. Tian, P. Ji, S. Liu, H. Tong, B. Wu, Y. Xiong, L. Zettlemoyer, G. Neubig, D. Weld, D. Downey, W. tau Yih, P. W. Koh, and H. Hajishirzi · 2024
Cited alongside, same era.
Iclr 2025: Assisting reviewers
I. Blog · 2024
Cited alongside, same era.
The ai scientist: Towards fully automated open-ended scientific discovery
L. Chris, L. Cong, L. Robert, Tjarko, F. Jakob, C. Jeff, and H. David · 2024
Cited alongside, same era.
Marg: Multi-agent review generation for scientific papers
M. D’Arcy, T. Hope, L. Birnbaum, and D. Downey · 2024
Cited alongside, same era.
LongroPE: Extending LLM context window beyond 2 million tokens
Y. Ding, L. L. Zhang, C. Zhang, Y. Xu, N. Shang, J. Xu, F. Yang, and M. Yang · 2024
Cited alongside, same era.
Human-in-the-loop ai reviewing: Feasibility, opportunities, and risks
I. Drori and D. Te’eni · 2024
Cited alongside, same era.
LLMs assist NLP researchers: Critique paper (meta-)reviewing
J. Du, Y. Wang, W. Zhao, Z. Deng, S. Liu, R. Lou, H. P. Zou, P. Narayanan Venkit, N. Zhang, M. Srinath, H. R. Zhang, V. Gupta, Y. Li, T. Li, F. Wang, Q. Liu, T. Liu, P. Gao, C. Xia, C. Xing, C. Jiayang, Z. Wang, Y. Su, R. S. Shah, R. Guo, J. Gu, H. Li, K. Wei, Z. Wang, L. Cheng, S. Ranathunga, M. Fang, J. Fu, F. Liu, R. Huang, E. Blanco, Y. Cao, R. Zhang, P. S. Yu, and W. Yin · 2024
Cited alongside, same era.
Reviewer2: Optimizing review generation through prompt generation
Z. Gao, K. Brantley, and T. Joachims · 2024
Cited alongside, same era.
Large language models for automated open-domain scientific hypotheses discovery
Z. Yang, X. Du, J. Li, J. Zheng, S. Poria, and E. Cambria · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Later among the works it cites.
Are we there yet? revealing the risks of utilizing large language models in scholarly peer review
R. Ye, X. Pang, J. Chai, J. Chen, Z. Yin, Z. Xiang, X. Dong, J. Shao, and S. Chen · 2024
Later among the works it cites.
Automated peer reviewing in paper sea: Standardization, evaluation, and analysis
J. Yu, Z. Ding, J. Tan, K. Luo, Z. Weng, C. Gong, L. Zeng, R. Cui, C. Han, Q. Sun, et al · 2024
Later among the works it cites.
Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation
Q. Zeng, M. Sidhu, H. P. Chan, L. Wang, and H. Ji · 2024
Later among the works it cites.
Is LLM a reliable reviewer? a comprehensive evaluation of LLM on automatic paper reviewing tasks
R. Zhou, L. Chen, and K. Yu · 2024
Later among the works it cites.
Is llm a reliable reviewer? a comprehensive evaluation of llm on automatic paper reviewing tasks
R. Zhou, L. Chen, and K. Yu · 2024
Later among the works it cites.
Aider is ai pair programming in your terminal
A. AI · 2025
Closest in time.
rstar-math: Small llms can master math reasoning with self-evolved deep thinking, 2025
X. Guan, L. L. Zhang, Y. Liu, N. Shang, Y. Sun, Y. Zhu, F. Yang, and M. Yang · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2025
Closest in time.
Potential and perils of large language models as judges of unstructured textual data
B. Rewina, P. Natalie, B. Sreyoshi, K. Satya, G. Alex, C. Elizabeth, I. Ikkei, T. David, C. Aman, and N. Naumaan · 2025
Closest in time.
Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers
C. Si, D. Yang, and T. Hashimoto · 2025
Closest in time.
Learning to plan & reason for evaluation with thinking-llm-as-a-judge
S. Swarnadeep, L. Xian, G. Marjan, W. Jason, and W. Tianlu · 2025
Closest in time.
Cycleresearcher: Improving automated research via automated review
Y. Weng, M. Zhu, G. Bao, H. Zhang, J. Wang, Y. Zhang, and L. Yang · 2025
Closest in time.
Towards system 2 reasoning in llms: Learning how to think with meta chain-of-thought, 2025
V. Xiang, C. Snell, K. Gandhi, A. Albalak, A. Singh, C. Blagden, D. Phung, R. Rafailov, N. Lile, D. Mahan, L. Castricato, J.-P. Franken, N. Haber, and C. Finn · 2025
Closest in time.
Large language models for automated scholarly paper review: A survey
Z. Zhuang, J. Chen, H. Xu, Y. Jiang, and J. Lin · 2025
Closest in time.