Fetching the paper…
Reading the bibliography…
How can we test whether state-of-the-art generative models, such as Blender and GPT-3, are good AI teachers, capable of replying to a student in an educational dialogue? Designing an AI teacher test is challenging: although evaluation methods are much-needed, there is no off-the-shelf solution to measuring pedagogical ability.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Hybrid Monte Carlo
S. Duane, A. D. Kennedy, B. J. Pendleton, and D. Roweth · 1987
Earlier work this paper cites.
Recipes for building an open-domain chatbot
S. Roller, E. Dinan, N. Goyal, D. Ju, M. Williamson, Y. Liu, J. Xu, M. Ott, K. Shuster, E. M. Smith, Y.-L. Boureau, and J. Weston · 2004
Earlier work this paper cites.
Measuring teacher effectiveness: Some methodological reflections
D. Muijs · 2006
Earlier work this paper cites.
Approaches to Evaluating Teacher Effectiveness: A Research Synthesis
L. Goe, C. Bell, and O. Little · 2008
Earlier work this paper cites.
Using the Method of Pairwise Comparison to Obtain Reliable Teacher Assessments
S. Heldsinger and S. Humphry · 2010
Earlier work this paper cites.
The No-U-turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo
M. D. Homan and A. Gelman · 2014
Earlier work this paper cites.
ParlAI: A dialog research software platform
A. Miller, W. Feng, D. Batra, A. Bordes, A. Fisch, J. Lu, D. Parikh, and J. Weston · 2014
Cited alongside, same era.
Comparative judgement as a promising alternative to score competences
M. Lesterhuis, S. Verhavert, L. Coertjens, V. Donche, and S. De Maeyer · 2017
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
The teacher-student chatroom corpus
A. Caines, H. Yannakoudakis, H. Edmondson, H. Allen, P. Pérez-Paredes, B. Byrne, and P. Buttery · 2020
Cited alongside, same era.
Measuring Conversational Uptake: A Case Study on Student-Teacher Interactions
D. Demszky, J. Liu, Z. Mancenido, J. Cohen, H. Hill, D. Jurafsky, and T. Hashimoto · 2021
Later among the works it cites.
MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
K. Pillutla, S. Swayamdipta, R. Zellers, J. Thickstun, S. Welleck, Y. Choi, and Z. Harchaoui · 2021
Later among the works it cites.
PyStan (3.3.0)
A. Riddell, A. Hartikainen, and M. Carter · 2021
Later among the works it cites.
Are We There Yet? - A Systematic Literature Review on Chatbots in Education
S. Wollny, J. Schneider, D. Di Mitri, J. Weidlich, M. Rittberger, and H. Drachsler · 2021
Later among the works it cites.
Dialogue systems for language learning: A meta-analysis
S. Bibauw, W. Van den Noortgate, T. François, and P. Desmet · 2022
Closest in time.
Stan Modeling Language Users Guide and Reference Manual (v2.29.0), 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can you put it all together: Evaluating conversational agents’ ability to blend skills
E. M. Smith, M. Williamson, K. Shuster, J. Weston, and Y.-L. Boureau · 2020
Cited alongside, same era.
On the Opportunities and Risks of Foundation Models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Kohd, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang
Cited in the paper.
Stan Development Team · 2022
Closest in time.