Fetching the paper…
Reading the bibliography…
Given the rapid ascent of large language models (LLMs), we study the question: (How) can large language models help in reviewing of scientific papers or proposals? We first conduct some pilot studies where we find that (i) GPT-4 outperforms other LLMs (Bard, Vicuna, Koala, Alpaca, LLaMa, Dolly, OpenAssistant, StableLM), and (ii) prompting with a specific question (e.g., to identify errors) outperforms prompting to simply write a review.
“RoBERTa: A Robustly Optimized BERT Pretraining Approach”, 2019
Yinhan Liu · 1907
Earlier work this paper cites.
“DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond and Thomas Wolf · 1910
Earlier work this paper cites.
“Rank analysis of incomplete block designs: I. The method of paired comparisons”
Ralph Bradley and Milton Terry · 1952
Earlier work this paper cites.
“Individual Choice Behavior”
R Luce · 1959
Earlier work this paper cites.
“Who reviews the reviewers? Feasibility of using a fictitious manuscript to evaluate peer reviewer performance”
William Baxt, Joseph Waeckerle, Jesse Berlin and Michael Callaham · 1998
Earlier work this paper cites.
“Effect on the quality of peer review of blinding reviewers and asking them to sign their reports: a randomized controlled trial”
Fiona Godlee, Catharine Gale and Christopher Martyn · 1998
Earlier work this paper cites.
“The myth of the double-blind review? Author identification using only citations”
Shawndra Hill and Foster J · 2003
Earlier work this paper cites.
“Effects of training on quality of peer review: randomised controlled trial”
Sara Schroter et al · 2004
Earlier work this paper cites.
“Is peer review broken? Submissions are up, reviewers are overtaxed, and authors are lodging complaint after complaint about the process at top-tier journals. What’s wrong with peer review?”
Alison McCook · 2006
Earlier work this paper cites.
“What errors do peer reviewers detect, and does training improve their ability to detect them?”
Sara Schroter et al · 2008
Earlier work this paper cites.
“The Toronto Paper Matching System: An automated paper-reviewer assignment system”
Laurent Charlin and Richard Zemel · 2013
Earlier work this paper cites.
“A Bayesian model for calibrating conference review scores”, Manuscript, 2013
Hong Ge, Max Welling and Zoubin Ghahramani · 2013
Earlier work this paper cites.
“Peer Review: How We Found 15 Million Hours of Lost Time”, 2013
The AJE Team · 2013
Earlier work this paper cites.
“Commensuration bias in peer review”
Carole Lee · 2015
Earlier work this paper cites.
“An Introduction to StatReviewer”, EMUG, 2016
Tim Houle · 2016
Earlier work this paper cites.
“Reviewer bias in single-versus double-blind peer review”
Andrew Tomkins, Min Zhang and William Heavlin · 2017
Earlier work this paper cites.
Jia-Bin Huang · 2018
Earlier work this paper cites.
“A dataset of peer reviews (peerread): Collection, insights and NLP applications”
Dongyeop Kang · 2018
Earlier work this paper cites.
“The myth of double-blind review revisited: ACL vs. EMNLP”
Cornelia Caragea, Ana Uban and Liviu Dinu · 2019
Earlier work this paper cites.
“Academic plagiarism detection: a systematic literature review”
Tomáš Foltỳnek, Norman Meuschke and Bela Gipp · 2019
Earlier work this paper cites.
“Argument mining for understanding peer reviews”
Xinyu Hua, Mitko Nikolov, Nikhil Badugu and Lu Wang · 2019
Earlier work this paper cites.
“Paper Matching with Local Fairness Constraints”
Ari Kobren, Barna Saha and Andrew McCallum · 2019
Earlier work this paper cites.
“Simple and Effective Paraphrastic Similarity from Parallel Translations”
John Wieting, Kevin Gimpel, Graham Neubig and Taylor Berg-Kirkpatrick · 2019
Earlier work this paper cites.
“Meta-Research: Large-scale language analysis of peer review reports”
Ivan Buljan et al · 2020
Earlier work this paper cites.
“Aspect-based Sentiment Analysis of Scientific Reviews”
Souvic Chakraborty, Pawan Goyal and Animesh Mukherjee · 2020
Earlier work this paper cites.
“SPECTER: Document-level Representation Learning using Citation-informed Transformers”
Arman Cohan et al · 2020
Earlier work this paper cites.
“Argument Mining Driven Analysis of Peer-Reviews”
Michael Fromm · 2020
Earlier work this paper cites.
“Mitigating manipulation in peer review via randomized reviewer assignments”
Steven Jecmen et al · 2020
Earlier work this paper cites.
“Distributed peer review enhanced with natural language processing and machine learning”
Wolfgang Kerzendorf et al · 2020
Earlier work this paper cites.
“Citations Beyond Self Citations: Identifying Authors, Affiliations, and Nationalities in Scientific Papers”
Yoshitomo Matsubara and Sameer Singh · 2020
Earlier work this paper cites.
“Potential Organized Fraud in ACM/IEEE Computer Architecture Conferences”, 2020
T.. Vijaykumar · 2020
Earlier work this paper cites.
“ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis”
Qingyun Wang et al · 2020
Cited alongside, same era.
“A Dataset for Discourse Structure in Peer Review Discussions”
Neha Kennard · 2021
Cited alongside, same era.
“Collusion rings threaten the integrity of computer science research”
Michael Littman · 2021
Cited alongside, same era.
“Uncovering Latent Biases in Text: Method and Application to Peer Review”
Emaad Manzoor and Nihar Shah · 2021
Cited alongside, same era.
“Loss Functions, Axioms, and Peer Review”
Ritesh Noothigattu, Nihar Shah and Ariel Procaccia · 2021
Cited alongside, same era.
“Generating Datasets with Pretrained Language Models”
“Hello Dolly: Democratizing the magic of ChatGPT with open models”, Blog post, 2023
Mike Conover · 2023
Closest in time.
“Improving Factuality and Reasoning in Language Models through Multiagent Debate”, 2023
Yilun Du et al · 2023
Closest in time.
“Koala: A Dialogue Model for Academic Research”, Blog post, 2023
Xinyang Geng · 2023
Closest in time.
“Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection”, 2023
Kai Greshake et al · 2023
Closest in time.
“Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other Large Language Models in scholarly peer review”
Mohammad Hosseini and Serge Horbach · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timo Schick and Hinrich Schütze · 2021
Cited alongside, same era.
“Catch Me if I Can: Detecting Strategic Behaviour in Peer Assessment”
Ivan Stelmakh, Nihar Shah and Aarti Singh · 2021
Cited alongside, same era.
“PeerReview4All: Fair and Accurate Reviewer Assignment in Peer Review”
Ivan Stelmakh, Nihar Shah and Aarti Singh · 2021
Cited alongside, same era.
“Want To Reduce Labeling Cost? GPT-3 Can Help”
Shuohang Wang et al · 2021
Cited alongside, same era.
“Can We Automate Scientific Reviewing?”
Weizhe Yuan, Pengfei Liu and Graham Neubig · 2021
Cited alongside, same era.
“Peer review analyze: A novel benchmark resource for computational analysis of peer reviews”
Tirthankar Ghosal, Sandeep Kumar, Prabhat Bharti and Asif Ekbal · 2022
Cited alongside, same era.
“Nobel and novice: Author prominence affects peer review”
Jürgen Huber et al · 2022
Cited alongside, same era.
“Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case Study”
Perttu Hämäläinen, Mikke Tavast and Anton Kunnari · 2023
Closest in time.
“A Dataset on Malicious Paper Bidding in Peer Review”
Steven Jecmen et al · 2023
Closest in time.
“Causal reasoning and large language models: Opening a new frontier for causality”
Emre Kıcıman, Robert Ness, Amit Sharma and Chenhao Tan · 2023
Closest in time.
“Large Language Models are Zero-Shot Reasoners”, 2023
Takeshi Kojima et al · 2023
Closest in time.
“Open Assistant”, 2023
LAION AI · 2023
Closest in time.
“Large Language Models are Few-Shot Health Learners”, 2023
Xin Liu · 2023
Closest in time.
“An overview of Bard: an early experiment with generative AI”, 2023
James Manyika · 2023
Closest in time.
“Using Transformer Language Models to Validate Peer-Assigned Essay Scores in Massive Open Online Courses (MOOCs)”
Wesley Morris, Scott Crossley, Langdon Holmes and Anne Trumbore · 2023
Closest in time.
“Large Language Models as Corporate Lobbyists”
John. Nay · 2023
Closest in time.
“ChatGPT – Release Notes”, Blog post, 2023
OpenAI · 2023
Closest in time.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Closest in time.
“Generative Agents: Interactive Simulacra of Human Behavior”
Joon Park et al · 2023
Closest in time.
“Stability AI Launches the First of its StableLM Suite of Language Models”, Blog post, 2023
Stability AI · 2023
Closest in time.
“A Gold Standard Dataset for the Reviewer Assignment Problem”
Ivan Stelmakh, John Wieting, Graham Neubig and Nihar Shah · 2023
Closest in time.
“Alpaca: A Strong, Replicable Instruction-Following Model”, Blog post, 2023
Rohan Taori · 2023
Closest in time.
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality”, 2023
The Vicuna Team · 2023
Closest in time.
“LLaMA: Open and Efficient Foundation Language Models”, 2023
Hugo Touvron · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez and Samuel Bowman · 2023
Closest in time.
“Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks”
Tomer Ullman · 2023
Closest in time.
“Voyager: An Open-Ended Embodied Agent with Large Language Models”, 2023
Guanzhi Wang · 2023
Closest in time.
“How Language Model Hallucinations Can Snowball”, 2023
Muru Zhang et al · 2023
Closest in time.
“Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models”
Shuai Zhao et al · 2023
Closest in time.
“Can Large Language Models Transform Computational Social Science?”, 2023
Caleb Ziems et al · 2023
Closest in time.
“Navigating the Grey Area: Expressions of Overconfidence and Uncertainty in Language Models”, 2023
Kaitlyn Zhou, Dan Jurafsky and Tatsunori Hashimoto · 2023
Closest in time.
“Estimation from pairwise comparisons: sharp minimax bounds with topology dependence”
Nihar Shah et al · 2095
Closest in time.