Fetching the paper…
Reading the bibliography…
In this paper, we investigate whether current state-of-the-art large language models (LLMs) are effective as AI tutors and whether they demonstrate pedagogical abilities necessary for good AI tutoring in educational dialogues.
The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring
Benjamin S Bloom. 1984 · 1984
Earlier work this paper cites.
Development and use of the ARCS model of instructional design
John M Keller. 1987 · 1987
Earlier work this paper cites.
Multimedia learning
Richard E Mayer. 2002 · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
The effects of choice on intrinsic motivation and related outcomes: a meta-analysis of research findings
Erika A Patall, Harris Cooper, and Jorgianne Civey Robinson. 2008 · 2008
Earlier work this paper cites.
Dialogue response ranking training with large-scale human feedback data
Xiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett, and Bill Dolan. 2020 · 2009
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh. 2012 · 2012
Earlier work this paper cites.
Pearson’s correlation coefficient
Philip Sedgwick. 2012 · 2012
Earlier work this paper cites.
The ICAP framework: Linking cognitive engagement to active learning outcomes
Michelene TH Chi and Ruth Wylie. 2014 · 2014
Earlier work this paper cites.
Developing a coding scheme for analysing classroom dialogue across educational contexts
Sara Hennessy, Sylvia Rojas-Drummond, Rupert Higham, Ana María Márquez, Fiona Maine, Rosa María Ríos, Rocío García-Carrión, Omar Torreblanca, and María José Barrera. 2016 · 2016
Earlier work this paper cites.
Reimagining the role of technology in higher education: A supplement to the national education technology plan
John King and Joseph South. 2017 · 2017
Earlier work this paper cites.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
How we learn: The new science of education and the brain
Stanislas Dehaene. 2020 · 2020
Earlier work this paper cites.
The Metacognitive Student: How to Teach Academic, Social, and Emotional Intelligence in Every Content Area
Richard K Cohen, Deanne Kildare Opatosky, James Savage, Susan Olsen Stevens, and Edward P Darrah. 2021 · 2021
Earlier work this paper cites.
Measuring conversational uptake: A case study on student-teacher interactions
Dorottya Demszky, Jing Liu, Zid Mancenido, Julie Cohen, Heather Hill, Dan Jurafsky, and Tatsunori Hashimoto. 2021 · 2021
Cited alongside, same era.
Uncommon sense teaching: Practical insights in brain science to help students learn
Barbara Oakley and Terrence J Sejnowski. 2021 · 2021
Cited alongside, same era.
Are We There Yet? - A Systematic Literature Review Chatbots in Education
Sebastian Wollny, Jan Schneider, Daniele Di Mitri, Joshua Weidlich, Marc Rittberger, and Hendrik Drachsler. 2021 · 2021
Cited alongside, same era.
The AI teacher test: Measuring the pedagogical ability of blender and GPT-3 in educational dialogues
Anaïs Tack and Chris Piech. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Can language models employ the socratic method? Experiments with code debugging
Erfan Al-Hossami, Razvan Bunescu, Justin Smith, and Ryan Teehan. 2024 · 2024
Closest in time.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic. 2024 · 2024
Closest in time.
Let me teach you: Pedagogical foundations of feedback for language models
Beatriz Borges, Niket Tandon, Tanja Käser, and Antoine Bosselut. 2024 · 2024
Closest in time.
A Survey on Evaluation of Large Language Models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024 · 2024
Closest in time.
Language models as science tutors
Alexis Chevalier, Jiayi Geng, Alexander Wettig, Howard Chen, Sebastian Mizera, Toni Annala, Max Jameson Aragon, Arturo Rodríguez Fanlo, Simon Frieder, Simon Machado, et al. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
RETUYT-InCo at BEA 2023 shared task: Tuning open-source LLMs for generating teacher responses
Alexis Baladón, Ignacio Sastre, Luis Chiruzzo, and Aiala Rosá. 2023 · 2023
Cited alongside, same era.
Large language models in education: Vision and opportunities
Wensheng Gan, Zhenlian Qi, Jiayang Wu, and Jerry Chun-Wei Lin. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Cited alongside, same era.
Changyoon Lee, Junho Myung, Jieun Han, Jiho Jin, and Alice Oh. 2023 · 2023
Cited alongside, same era.
G-Eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
Jakub Macina, Nico Daheim, Sankalan Pal Chowdhury, Tanmay Sinha, Manu Kapur, Iryna Gurevych, and Mrinmaya Sachan. 2023 · 2023
Cited alongside, same era.
The BEA 2023 shared task on generating AI teacher responses in educational dialogues
Anaïs Tack, Ekaterina Kochmar, Zheng Yuan, Serge Bibauw, and Chris Piech. 2023 · 2023
Cited alongside, same era.
Nico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych, and Mrinmaya Sachan. 2024 · 2024
Closest in time.
Generative AI for Education (GAIED): Advances, Opportunities, and Challenges
Paul Denny, Sumit Gulwani, Neil T Heffernan, Tanja Käser, Steven Moore, Anna N Rafferty, and Adish Singla. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Towards responsible development of generative AI for education: An evaluation-driven approach
Irina Jurenka, Markus Kunesch, Kevin R McKee, Daniel Gillick, Shaojian Zhu, Sara Wiltberger, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, et al. 2024 · 2024
Closest in time.
Prometheus 2: An open source language model specialized in evaluating other language models
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024 · 2024
Closest in time.
Teaching CS50 with AI: leveraging generative artificial intelligence in computer science education
Rongxin Liu, Carter Zenke, Charlie Liu, Andrew Holmes, Patrick Thornton, and David J Malan. 2024 · 2024
Closest in time.
Instructors as innovators: A future-focused approach to new AI learning opportunities, with prompts
Ethan Mollick and Lilach Mollick. 2024 · 2024
Closest in time.
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
Sankalan Pal Chowdhury, Vilém Zouhar, and Mrinmaya Sachan. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024 · 2024
Closest in time.
Bridging the novice-expert gap via models of decision-making: A case study on remediating math mistakes
Rose Wang, Qingyang Zhang, Carly Robinson, Susanna Loeb, and Dorottya Demszky. 2024a · 2024
Closest in time.