Fetching the paper…
Reading the bibliography…
While large language models (LLMs) have been used for automated grading, they have not yet achieved the same level of performance as humans, especially when it comes to grading complex questions.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Text-to-text semantic similarity for automatic short answer grading
Michael Mohler and Rada Mihalcea · 2009
Earlier work this paper cites.
Learning to grade short answer questions using semantic similarity measures and dependency graph alignments
Michael Mohler, Razvan Bunescu, and Rada Mihalcea · 2011
Earlier work this paper cites.
Stratified sampling meets machine learning
Edo Liberty, Kevin Lang, and Konstantin Shmakov · 2016
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling, 2022
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Earlier work this paper cites.
Gradeaid: a framework for automatic short answers grading in educational contexts—design, implementation and evaluation
Emiliano Del Gobbo, Alfonso Guarino, Barbara Cafarelli, and Luca Grilli · 2023
Earlier work this paper cites.
Automation of short answer grading techniques: Comparative study using deep learning techniques
Arunima Divya, Vivek Haridas, and Jayasree Narayanan · 2023
Earlier work this paper cites.
Hosein Hasanbeig, Hiteshi Sharma, Leo Betthauser, Felipe Vieira Frujeri, and Ida Momennejad · 2023
Earlier work this paper cites.
Owen Henkel, Libby Hills, Bill Roberts, and Joshua McGrane · 2023
Earlier work this paper cites.
Metagpt: Meta programming for a multi-agent collaborative framework, 2023
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber · 2023
Cited alongside, same era.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, 2023
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu · 2023
Cited alongside, same era.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Gpt-4 as an automatic grader: The accuracy of grades set by gpt-4 on introductory programming assignments, 2023
Filippa Nilsson and Jonatan Tuvstedt · 2023
Cited alongside, same era.
Large language models for education: Grading open-ended questions using chatgpt
How reliable are automatic evaluation methods for instruction-tuned llms?
Ehsan Doostmohammadi, Oskar Holmström, and Marco Kuhlmann · 2024
Closest in time.
Beyond traditional assessment: Exploring the impact of large language models on grading practices
O Fagbohun, NP Iduwe, M Abdullahi, A Ifaturoti, and OM Nwanna · 2024
Closest in time.
Rankprompt: Step-by-step comparisons make language models better reasoners
Chi Hu, Yuan Ge, Xiangnan Ma, Hang Cao, Qiang Li, Yonghua Yang, Tong Xiao, and Jingbo Zhu · 2024
Closest in time.
Evaluating students’ open-ended written responses with llms: Using the rag framework for gpt-3.5, gpt-4, claude-3, and mistral-large, 2024
Jussi S. Jauhiainen and Agustín Garagorry Guerra · 2024
Closest in time.
Fine-tuning chatgpt for automatic scoring
Ehsan Latif and Xiaoming Zhai · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gustavo Pinto, Isadora Cardoso-Pereira, Danilo Monteiro, Danilo Lucena, Alberto Souza, and Kiev Gama · 2023
Cited alongside, same era.
Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2023
Cited alongside, same era.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face, 2023
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang · 2023
Cited alongside, same era.
Short answer grading using one-shot prompting and text similarity scoring model
Su-Youn Yoon · 2023
Cited alongside, same era.
Detectors for safe and reliable llms: Implementations, uses, and limitations
Swapnaja Achintalwar, Adriana Alvarado Garcia, Ateret Anaby-Tavor, Ioana Baldini, Sara E Berger, Bishwaranjan Bhattacharjee, Djallel Bouneffouf, Subhajit Chaudhury, Pin-Yu Chen, Lamogha Chiazor, et al · 2024
Cited alongside, same era.
Automatic short answer grading for finnish with chatgpt
Li-Hsin Chang and Filip Ginter · 2024
Cited alongside, same era.
Humans or llms as the judge? a study on judgement biases
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang · 2024
Cited alongside, same era.
A chain-of-thought prompting approach with llms for evaluating students’ formative assessment responses in science
Clayton Cohn, Nicole Hutchins, Tuan Le, and Gautam Biswas · 2024
Cited alongside, same era.
Aligning with human judgement: The role of pairwise preference in large language model evaluators
Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulic, Anna Korhonen, and Nigel Collier · 2024
Closest in time.
Llm comparative assessment: Zero-shot nlg evaluation through pairwise comparisons using large language models
Adian Liusie, Potsawee Manakul, and Mark Gales · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2024
Closest in time.
Openai api, 2023
OpenAI · 2024
Closest in time.
Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment
Vyas Raina, Adian Liusie, and Mark Gales · 2024
Closest in time.
Automated essay scoring and revising based on open-source large language models
Yishen Song, Qianta Zhu, Huaibo Wang, and Qinhua Zheng · 2024
Closest in time.
Enhancing large language models against inductive instructions with dual-critique prompting, 2024
Rui Wang, Hongru Wang, Fei Mi, Yi Chen, Boyang Xue, Kam-Fai Wong, and Ruifeng Xu · 2024
Closest in time.