Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated strong potential in performing automatic scoring for constructed response assessments.
Graesser, A.C., Chipman, P., Haynes, B.C., Olney, A.: Autotutor: An intelligent tutoring system with mixed-initiative dialogue. IEEE Transactions on Education 48
2005
Earlier work this paper cites.
Parikh, D., Lu, Y., Xin, Y., Wu, D., Pelz, J., Lu, G.: Where am i looking: Localizing gaze in reconstructed 3d space. In: 2019 IEEE Global Conference on Signal and Information Processing (GlobalSIP), pp. 1–5 (2019). IEEE
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Zhai, X., Haudek, L. K.C.and Shi, Nehm, R., M., U.-L.: From substitution to redefinition: A framework of machine learning-based science assessment. Journal of Research in Science Teaching (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Zhai, X.: Advancing automatic guidance in virtual science inquiry: from ease of use to personalization. Educational Technology Research and Development (2021)
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., Iwasawa, Y.: Large language models are zero-shot reasoners. Advances in neural information processing systems 35
2022
Earlier work this paper cites.
Haque, S., Eberhart, Z., Bansal, A., McMillan, C.: Semantic similarity metrics for evaluating source code summarization. In: Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension, pp. 36–47 (2022)
2022
Earlier work this paper cites.
Wu, D., Wang, M., Li, X., et al.: Automatic scoring for translations based on language models. Computational Intelligence and Neuroscience 2022
2022
Earlier work this paper cites.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al
2022
Earlier work this paper cites.
Zhai, X., He, P., Krajcik, J.: Applying machine learning to automatically assess scientific models. Journal of Research in Science Teaching 59
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Zhai, X.: Chatgpt for next generation science learning. XRDS: Crossroads, The ACM Magazine for Students 29
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Du, M., He, F., Zou, N., Tao, D., Hu, X.: Shortcut learning of large language models in natural language understanding (2023)
2023
Cited alongside, same era.
Tang, R., Kong, D., Huang, L., Xue, H.: Large language models can be lazy learners: Analyze shortcuts in in-context learning. In: Findings of the Association for Computational Linguistics: ACL 2023, pp. 4645–4657 (2023)
2023
Cited alongside, same era.
Google Deepmind: Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context. (2024). Google Deepmind
2024
Closest in time.
2024
Closest in time.
Harris, C.J., Krajcik, J.S., Pellegrino, J.W.: Creating and Using Instructionally Supportive Assessments in NGSS Classrooms. NSTA Press, National Science Teaching Association, ??? (2024)
2024
Closest in time.
Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G., Li, X., Jin, Y., Gašević, D.: Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology 55
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Organisciak, P., Acar, S., Dumas, D., Berthiaume, K.: Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models. Thinking Skills and Creativity 49
2023
Cited alongside, same era.
Bewersdorff, A., Seßler, K., Baur, A., Kasneci, E., Nerdel, C.: Assessing student errors in experimentation using artificial intelligence and large language models: A comparative study with human raters. Computers and Education: Artificial Intelligence 5
2023
Cited alongside, same era.
OpenAI: Gpt-4 technical report. arXiv:2303.08774 (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Wilson, C.D., Haudek, K.C., Osborne, J.F., Buck Bracey, Z.E., Cheuk, T., Donovan, B.M., Stuhlsatz, M.A., Santiago, M.M., Zhai, X.: Using automated analysis to assess middle school students’ competence with scientific argumentation. Journal of Research in Science Teaching 61
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Latif, E., Zhai, X.: Fine-tuning chatgpt for automatic scoring. Computers and Education: Artificial Intelligence, 100210 (2024)
2024
Closest in time.
Zhai, X.: Ai and machine learning for next generation sci-ence assessments. Machine Learning, Natural Language Processing, and Psychometrics, 201 (2024)
2024
Closest in time.
Cohn, C., Hutchins, N., Le, T., Biswas, G.: A chain-of-thought prompting approach with llms for evaluating students’ formative assessment responses in science. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 23182–23190 (2024)
2024
Closest in time.
2024
Closest in time.
Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H., Wang, S., Yin, D., Du, M.: Explainability for large language models: A survey. ACM Transactions on Intelligent Systems and Technology 15
2024
Closest in time.
2024
Closest in time.
Wu, X., Yao, W., Chen, J., Pan, X., Wang, X., Liu, N., Yu, D.: From language modeling to instruction following: Understanding the behavior shift in llms after instruction tuning. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 2341–2369 (2024)
2024
Closest in time.
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al.: Judging llm-as-a-judge with mt-bench and chatbot arena. NIPS 36
2024
Closest in time.