Fetching the paper…
Reading the bibliography…
Recent advances in large language models (LLMs) have enabled zero-shot automated essay scoring (AES), providing a promising way to reduce the cost and effort of essay scoring in comparison with manual grading.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry. 1952 · 1952
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
The Rating of Chessplayers, Past and Present
Arpad E. Elo. 1978 · 1978
Earlier work this paper cites.
Learning to rank using gradient descent
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. 2005 · 2005
Earlier work this paper cites.
Robert Ridley, Liang He, Xinyu Dai, Shujian Huang, and Jiajun Chen. 2020 · 2008
Earlier work this paper cites.
A new dataset and method for automatically grading ESOL texts
Helen Yannakoudakis, Ted Briscoe, and Ben Medlock. 2011 · 2011
Earlier work this paper cites.
TOEFL11: A corpus of non-native english
Daniel Blanchard, Joel R. Tetreault, Derrick Higgins, A. Cahill, and Martin Chodorow. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Automatic text scoring using neural networks
Dimitrios Alikaniotis, Helen Yannakoudakis, and Marek Rei. 2016 · 2016
Earlier work this paper cites.
A neural approach to automated essay scoring
Kaveh Taghipour and Hwee Tou Ng. 2016 · 2016
Earlier work this paper cites.
Attention-based recurrent convolutional neural network for automatic essay scoring
Fei Dong, Yue Zhang, and Jie Yang. 2017 · 2017
Earlier work this paper cites.
Deep learning using rectified linear units (ReLU)
Abien Fred Agarap. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Enhancing automated essay scoring performance via fine-tuning pre-trained language models with combination of regression and ranking
Ruosong Yang, Jiannong Cao, Zhiyuan Wen, Youzheng Wu, and Xiaodong He. 2020 · 2020
Cited alongside, same era.
Automated cross-prompt scoring of essay traits
Robert Ridley, Liang He, Xin yu Dai, Shujian Huang, and Jiajun Chen. 2021 · 2021
Cited alongside, same era.
A review of deep-neural automated essay scoring models
Masaki Uto. 2021 · 2021
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Analytic automated essay scoring based on deep neural networks integrating multidimensional item response theory
Takumi Shibata and Masaki Uto. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
PLAES: Prompt-generalized and level-aware learning framework for cross-prompt automated essay scoring
Yuan Chen and Xia Li. 2024 · 2024
Later among the works it cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024 · 2024
Later among the works it cites.
Unleashing large language models’ proficiency in zero-shot essay scoring
Sanwoo Lee, Yida Cai, Desong Meng, Ziyang Wang, and Yunfang Wu. 2024 · 2024
Later among the works it cites.
Aligning with human judgement: The role of pairwise preference in large language model evaluators
Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulić, Anna Korhonen, and Nigel Collier. 2024 · 2024
Later among the works it cites.
Efficient LLM comparative assessment: A product of experts framework for pairwise comparisons
Adian Liusie, Vatsal Raina, Yassir Fathullah, and Mark Gales. 2024b · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Automated essay scoring via pairwise contrastive regression
Jiayi Xie, Kaiwei Cai, Li Kong, Junsheng Zhou, and Weiguang Qu. 2022 · 2022
Cited alongside, same era.
PMAES: Prompt-mapping contrastive learning for cross-prompt automated essay scoring
Yuan Chen and Xia Li. 2023 · 2023
Cited alongside, same era.
Prompt- and trait relation-aware cross-prompt essay trait scoring
Heejin Do, Yunsu Kim, and Gary Geunbae Lee. 2023 · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Exploring the potential of using an ai language model for automated essay scoring
Atsushi Mizumoto and Masaki Eguchi. 2023 · 2023
Cited alongside, same era.
Rating short L2 essays on the CEFR scale with GPT-4
Kevin P. Yancey, Geoffrey Laflair, Anthony Verardi, and Jill Burstein. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Can large language models automatically score proficiency of written essays?
Watheq Ahmad Mansour, Salam Albatarni, Sohaila Eltanbouly, and Tamer Elsayed. 2024 · 2024
Later among the works it cites.
OpenAI. 2024 · 2024
Later among the works it cites.
PairEval: Open-domain dialogue evaluation with pairwise comparison
ChaeHun Park, Minseok Choi, Dohyun Lee, and Jaegul Choo. 2024 · 2024
Later among the works it cites.
Large language models are effective text rankers with pairwise ranking prompting
Zhen Qin, Rolf Jagerman, Kai Hui, and 1 others. 2024 · 2024
Later among the works it cites.
Beyond agreement: Diagnosing the rationale alignment of automated essay scoring methods based on linguistically-informed counterfactuals
Yupei Wang, Renfen Hu, and Zhe Zhao. 2024 · 2024
Later among the works it cites.
From generation to judgment: Opportunities and challenges of LLM-as-a-judge
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2025 · 2025
Closest in time.
KAES: Multi-aspect shared knowledge finding and aligning for cross-prompt automated scoring of essay traits
Xia Li and Wenjing Pan. 2025 · 2025
Closest in time.
T-MES: Trait-aware mix-of-experts representation learning for multi-trait essay scoring
Jiong Wang and Jie Liu. 2025 · 2025
Closest in time.