Fetching the paper…
Reading the bibliography…
Evaluating the pedagogical capabilities of AI-based tutoring models is critical for making guided progress in the field.
CLASS: A design framework for building intelligent tutoring systems based on learning science principles
Shashank Sonkar, Naiming Liu, Debshila Mallick, and Richard Baraniuk. 2023 · 1961
Earlier work this paper cites.
Collaborative dialogue patterns in naturalistic one-to-one tutoring
Arthur C Graesser, Natalie K Person, and Joseph P Magliano. 1995 · 1995
Earlier work this paper cites.
Chapter 7 - the wisdom of practice: Lessons learned from the study of highly effective tutors
Mark R. Lepper and Maria Woolverton. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Scaffolding practices that enhance mathematics learning
Julia Anghileri. 2006 · 2006
Earlier work this paper cites.
Correlations between dialogue acts and learning in spoken tutoring dialogues
Diane Litman and Kate Forbes-Riley. 2006 · 2006
Earlier work this paper cites.
Balancing cognitive and motivational scaffolding in tutorial dialogue
Kristy Elizabeth Boyer, Robert Phillips, Michael Wallis, Mladen Vouk, and James Lester. 2008 · 2008
Earlier work this paper cites.
The icap framework: Linking cognitive engagement to active learning outcomes
Michelene TH Chi and Ruth Wylie. 2014 · 2014
Earlier work this paper cites.
Active learning increases student performance in science, engineering, and mathematics
Scott Freeman, Sarah L Eddy, Miles McDonough, Michelle K Smith, Nnadozie Okoroafor, Hannah Jordt, and Mary Pat Wenderoth. 2014 · 2014
Earlier work this paper cites.
Autotutor and family: A review of 17 years of natural language tutoring
Benjamin D Nye, Arthur C Graesser, and Xiangen Hu. 2014 · 2014
Earlier work this paper cites.
Examining productive failure, productive success, unproductive failure, and unproductive success in learning
Manu Kapur. 2016 · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, and 3 others. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, and 1 others. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Cited alongside, same era.
Measuring conversational uptake: A case study on student-teacher interactions
Dorottya Demszky, Jing Liu, Zid Mancenido, Julie Cohen, Heather Hill, Dan Jurafsky, and Tatsunori Hashimoto. 2021 · 2021
Cited alongside, same era.
Dialog inpainting: Turning documents into dialogs
Zhuyun Dai, Arun Tejasvi Chaganty, Vincent Y Zhao, Aida Amini, Qazi Mamunur Rashid, Mike Green, and Kelvin Guu. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022 · 2022
Cited alongside, same era.
Automatic generation of socratic subquestions for teaching math word problems
Kumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha, Manu Kapur, and Mrinmaya Sachan. 2022 · 2022
Towards responsible development of generative ai for education: An evaluation-driven approach
Irina Jurenka, Markus Kunesch, Kevin R McKee, Daniel Gillick, Shaojian Zhu, Sara Wiltberger, Shubham Milind Phal, Katherine Hermann, Daniel Kasenberg, Avishkar Bhoopchand, and 1 others. 2024 · 2024
Later among the works it cites.
Instruct, not assist: LLM-based multi-turn planning and hierarchical questioning for socratic code debugging
Priyanka Kargupta, Ishika Agarwal, Dilek Hakkani Tur, and Jiawei Han. 2024 · 2024
Later among the works it cites.
Prometheus 2: An open source language model specialized in evaluating other language models
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024 · 2024
Later among the works it cites.
Mathchat: Benchmarking mathematical reasoning and instruction following in multi-turn interactions
Zhenwen Liang, Dian Yu, Wenhao Yu, Wenlin Yao, Zhihan Zhang, Xiangliang Zhang, and Dong Yu. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The ai teacher test: Measuring the pedagogical ability of blender and gpt-3 in educational dialogues
Anaïs Tack and Chris Piech. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023 · 2023
Cited alongside, same era.
The NCTE transcripts: A dataset of elementary math classroom transcripts
Dorottya Demszky and Heather Hill. 2023 · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
MathDial: A dialogue tutoring dataset with rich pedagogical properties grounded in math reasoning problems
Jakub Macina, Nico Daheim, Sankalan Chowdhury, Tanmay Sinha, Manu Kapur, Iryna Gurevych, and Mrinmaya Sachan. 2023a · 2023
Cited alongside, same era.
The BEA 2023 shared task on generating AI teacher responses in educational dialogues
Anaïs Tack, Ekaterina Kochmar, Zheng Yuan, Serge Bibauw, and Chris Piech. 2023 · 2023
Cited alongside, same era.
SocraticLM: Exploring socratic personalized teaching with large language models
Jiayu Liu, Zhenya Huang, Tong Xiao, Jing Sha, Jinze Wu, Qi Liu, Shijin Wang, and Enhong Chen. 2024 · 2024
Later among the works it cites.
Ruffle&riley: Insights from designing and evaluating a large language model-based conversational tutoring system
Robin Schmucker, Meng Xia, Amos Azaria, and Tom Mitchell. 2024 · 2024
Later among the works it cites.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2024 · 2024
Later among the works it cites.
Book2Dial: Generating teacher student interactions from textbooks for cost-effective development of educational chatbots
Junling Wang, Jakub Macina, Nico Daheim, Sankalan Pal Chowdhury, and Mrinmaya Sachan. 2024a · 2024
Later among the works it cites.
Rose E. Wang, Qingyang Zhang, Carly Robinson, Susanna Loeb, and Dorottya Demszky. 2024b · 2024
Later among the works it cites.
The promises and pitfalls of using language models to measure instruction quality in education
Paiheng Xu, Jing Liu, Nathan Jones, Julie Cohen, and Wei Ai. 2024 · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, and 1 others. 2024 · 2024
Later among the works it cites.
RewardBench: Evaluating reward models for language modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi. 2025 · 2025
Closest in time.
Unifying AI tutor evaluation: An evaluation taxonomy for pedagogical ability assessment of LLM-powered AI tutors
Kaushal Kumar Maurya, Kv Aditya Srivatsa, Kseniia Petukhova, and Ekaterina Kochmar. 2025 · 2025
Closest in time.
Learnlm: Improving gemini for learning
LearnLM Team, Abhinit Modi, Aditya Srikanth Veerubhotla, Aliya Rysbek, Andrea Huber, Brett Wiltshire, Brian Veprek, Daniel Gillick, Daniel Kasenberg, Derek Ahmed, Irina Jurenka, James Cohan, Jennifer She, Julia Wilkowski, Kaiz Alarakyia, Kevin R. McKee, Lisa Wang, Markus Kunesch, Mike Schaekermann, and 27 others. 2025 · 2025
Closest in time.
Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference
Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, and Iacopo Poli. 2025 · 2025
Closest in time.