Fetching the paper…
Reading the bibliography…
Evaluating the quality of a dialogue system is an understudied problem.
Flesch, R.F.: How to write plain English : a book for lawyers and consumers (1979)
1979
Earlier work this paper cites.
Danieli, M., Gerbino, E.: Metrics for evaluating dialogue strategies in a spoken language system. In: Proceedings of the 1995 AAAI spring symposium on Empirical Methods in Discourse Interpretation and Generation
1995
Earlier work this paper cites.
Hirschberg, J., Nakatani, C.H.: A prosodic analysis of discourse segments in direction-giving monologues. In: 34th ACL, 24-27 June 1996
1996
Earlier work this paper cites.
Walker, M.A., Kamm, C.A., Litman, D.J.: Towards developing general models of usability with PARADISE. Nat. Lang. Eng. 6
2000
Earlier work this paper cites.
Papineni, K., Roukos, S., Ward, T., Zhu, W.: Bleu: a method for automatic evaluation of machine translation. In: 40th ACL, July 6-12, 2002
2002
Earlier work this paper cites.
Finch, A.M., Akiba, Y., Sumita, E.: How does automatic machine translation evaluation correlate with human scoring as the number of reference translations increases? In: LREC 2004, May 26-28, 2004
2004
Earlier work this paper cites.
Lin, C.Y.: ROUGE: A package for automatic evaluation of summaries. In: Text Summarization Branches Out (2004)
2004
Earlier work this paper cites.
Lin, C., Och, F.J.: Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In: 42nd ACL, 2004
2004
Earlier work this paper cites.
Banerjee, S., Lavie, A.: METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In: Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization@ACL, June 29, 2005
2005
Earlier work this paper cites.
López-Cózar, R., Callejas, Z., McTear, M.F.: Testing the performance of spoken dialogue systems by means of an artificially simulated user. Artif. Intell. Rev. 26
2006
Earlier work this paper cites.
Schatzmann, J., Weilhammer, K., Stuttle, M., Young, S.: A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies. The knowledge engineering review 21
2006
Earlier work this paper cites.
Schatzmann, J., Thomson, B., Weilhammer, K., Ye, H., Young, S.: Agenda-based user simulation for bootstrapping a pomdp dialogue system. In: NAACL-HLT 2007
2007
Earlier work this paper cites.
Hastie, H.: Metrics and evaluation of spoken dialogue systems pp. 131–150 (2012)
2012
Earlier work this paper cites.
Rus, V., Lintean, M.C.: A comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics. In: Proceedings of the 7th Workshop on Building Educational Applications Using NLP, June 7, 2012
2012
Earlier work this paper cites.
Sun, H., Zhou, M.: Joint learning of a dual SMT system for paraphrase generation. In: 50th ACL, July 8-14, 2012
2012
Earlier work this paper cites.
Pietquin, O., Hastie, H.F.: A survey on metrics for the evaluation of user simulations. Knowl. Eng. Rev. 28
2013
Earlier work this paper cites.
Young, S., Gašić, M., Thomson, B., Williams, J.D.: Pomdp-based statistical spoken dialog systems: A review. Proceedings of the IEEE 101
2013
Earlier work this paper cites.
Forgues, G., Pineau, J., Larchevêque, J.M., Tremblay, R.: Bootstrapping dialog systems with word embeddings. In: Nips, modern machine learning and natural language processing workshop (2014)
2014
Earlier work this paper cites.
Li, L., He, H., Williams, J.D.: Temporal supervised learning for inferring a dialog policy from example conversations. In: 2014 IEEE Spoken Language Technology Workshop, December 7-10, 2014
2014
Earlier work this paper cites.
Galley, M., Brockett, C., Sordoni, A., Ji, Y., Auli, M., Quirk, C., Mitchell, M., Gao, J., Dolan, B.: deltableu: A discriminative metric for generation tasks with intrinsically diverse targets. In: 53rd ACL, July 26-31, 2015
2015
Earlier work this paper cites.
Kiros, R., Zhu, Y., Salakhutdinov, R., Zemel, R.S., Urtasun, R., Torralba, A., Fidler, S.: Skip-thought vectors. In: NIPS 2015, December 7-12
2015
Earlier work this paper cites.
Vedantam, R., Zitnick, C.L., Parikh, D.: Cider: Consensus-based image description evaluation. In: CVPR, June 7-12, 2015
2015
Earlier work this paper cites.
Vinyals, O., Le, Q.V.: A neural conversational model. CoRR 1506.05869
2015
Earlier work this paper cites.
Wen, T., Gasic, M., Mrksic, N., Su, P., Vandyke, D., Young, S.J.: Semantically conditioned lstm-based natural language generation for spoken dialogue systems. In: EMNLP, September 17-21, 2015
2015
Earlier work this paper cites.
Li, J., Galley, M., Brockett, C., Gao, J., Dolan, B.: A diversity-promoting objective function for neural conversation models. In: NAACL-HLT, June 12-17, 2016
2016
Earlier work this paper cites.
Li, J., Galley, M., Brockett, C., Spithourakis, G.P., Gao, J., Dolan, W.B.: A persona-based neural conversation model. In: 54th ACL, August 7-12, 2016
2016
Earlier work this paper cites.
Li, J., Monroe, W., Ritter, A., Jurafsky, D., Galley, M., Gao, J.: Deep reinforcement learning for dialogue generation. In: EMNLP, November 1-4, 2016
2016
Cited alongside, same era.
Liu, C., Lowe, R., Serban, I., Noseworthy, M., Charlin, L., Pineau, J.: How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In: EMNLP, November 1-4, 2016
2016
Cited alongside, same era.
Su, P., Gasic, M., rt al.: On-line active reward learning for policy optimisation in spoken dialogue systems. In: 54th ACL, August 7-12, 2016
2016
Cited alongside, same era.
Wieting, J., Bansal, M., Gimpel, K., Livescu, K.: Towards universal paraphrastic sentence embeddings. In: 4th ICLR, May 2-4, 2016
2016
Cited alongside, same era.
Dusek, O., Novikova, J., Rieser, V.: Referenceless quality estimation for natural language generation. CoRR 1708.01759
Tao, C., Mou, L., Zhao, D., Yan, R.: RUBER: an unsupervised method for automatic evaluation of open-domain dialog systems. In: 32nd AAAI, Feb 2-7, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Clark, E., Celikyilmaz, A., Smith, N.A.: Sentence mover’s similarity: Automatic evaluation for multi-sentence texts. In: 57th ACL (2019)
2019
Later among the works it cites.
Deriu, J., Rodrigo, Á., Otegi, A., Echegoyen, G., Rosset, S., Agirre, E., Cieliebak, M.: Survey on evaluation methods for dialogue systems. CoRR 1905.04071
2019
Later among the works it cites.
Dziri, N., Kamalloo, E., Mathewson, K.W., Zaïane, O.R.: Evaluating coherence in dialogue systems using entailment. In: NAACL-HLT, June 2-7, 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Kannan, A., Vinyals, O.: Adversarial evaluation of dialogue models. CoRR 1701.08198
2017
Cited alongside, same era.
Li, X., Chen, Y., Li, L., Gao, J., Celikyilmaz, A.: End-to-end task-completion neural dialogue systems. In: 8th IJCNLP, November 27 - December 1, 2017
2017
Cited alongside, same era.
Lowe, R., Noseworthy, M., Serban, I.V., Angelard-Gontier, N., Bengio, Y., Pineau, J.: Towards an automatic turing test: Learning to evaluate dialogue responses. In: 55th ACL (2017)
2017
Cited alongside, same era.
Miller, A.H., Feng, W., Batra, D., Bordes, A., Fisch, A., Lu, J., Parikh, D., Weston, J.: Parlai: A dialog research software platform. In: 2017 EMNLP
2017
Cited alongside, same era.
Novikova, J., Dusek, O., Curry, A.C., Rieser, V.: Why we need new evaluation metrics for NLG. In: EMNLP, September 9-11, 2017
2017
Cited alongside, same era.
Novikova, J., Dusek, O., Rieser, V.: The E2E dataset: New challenges for end-to-end generation. In: 18th SIGdial, August 15-17, 2017
2017
Cited alongside, same era.
Peng, B., Li, X., Li, L., Gao, J., Celikyilmaz, A., Lee, S., Wong, K.: Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning. In: EMNLP, September 9-11, 2017
2017
Cited alongside, same era.
2019
Later among the works it cites.
Ghandeharioun, A., Shen, J.H., Jaques, N., Ferguson, C., Jones, N., Lapedriza, À., Picard, R.W.: Approximating interactive human evaluation with self-play for open-domain dialog systems. In: NeurIPS, December 8-14, 2019
2019
Later among the works it cites.
Gupta, P., Mehri, S., Zhao, T., Pavel, A., Eskénazi, M., Bigham, J.P.: Investigating evaluation of open-domain dialogue systems with human generated multiple references. In: SIGdial, September 11-13, 2019
2019
Later among the works it cites.
Hashimoto, T.B., Zhang, H., Liang, P.: Unifying human and statistical evaluation for natural language generation. In: NAACL-HLT, June 2-7, 2019
2019
Later among the works it cites.
Lee, S., Zhu, Q., Takanobu, R., Zhang, Z., Zhang, Y., Li, X., Li, J., Peng, B., Li, X., Huang, M., Gao, J.: Convlab: Multi-domain end-to-end dialog system platform. In: 57th ACL, July 28 - August 2, 2019
2019
Later among the works it cites.
Sai, A.B., Gupta, M.D., Khapra, M.M., Srinivasan, M.: Re-evaluating ADEM: A deeper look at scoring dialogue responses. In: AAAI, Jan 27 - Feb 1, 2019
2019
Later among the works it cites.
See, A., Roller, S., Kiela, D., Weston, J.: What makes a good conversation? how controllable attributes affect human judgments. In: NAACL-HLT, June 2-7, 2019
2019
Later among the works it cites.
Shi, W., Qian, K., Wang, X., Yu, Z.: How to build user simulators to train rl-based dialog systems. In: 2019 EMNLP-IJCNLP
2019
Later among the works it cites.
Wu, C.S., Socher, R., Xiong, C.: Global-to-local memory pointer networks for task-oriented dialogue. In: 7th ICLR (2019)
2019
Later among the works it cites.
Zhao, W., Peyrard, M., Liu, F., Gao, Y., Meyer, C.M., Eger, S.: Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance. In: EMNLP-IJCNLP 2019
2019
Later among the works it cites.
Howcroft, D.M., Belz, A., Clinciu, M., Gkatzia, D., Hasan, S.A., Mahamood, S., Mille, S., van Miltenburg, E., Santhanam, S., Rieser, V.: Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions. In: 13thINLG, December 15-18, 2020
2020
Later among the works it cites.
Mehri, S., Eskénazi, M.: USR: an unsupervised and reference free evaluation metric for dialog generation. In: 58th ACL, July 5-10, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Sellam, T., Das, D., Parikh, A.P.: Bleurt: Learning robust metrics for text generation. In: 58th ACL (2020)
2020
Later among the works it cites.
Takanobu, R., Zhu, Q., Li, J., Peng, B., Gao, J., Huang, M.: Is your goal-oriented dialog model performing really well? empirical analysis of system-wise evaluation. In: Proceedings of the 21st SIGDial (2020)
2020
Later among the works it cites.
Yuma, T., Yoshinaga, N., Toyoda, M.: ubleu: Uncertainty-aware automatic evaluation method for open-domain dialogue systems. In: 58th ACL: SRW (2020)
2020
Later among the works it cites.
Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., Artzi, Y.: Bertscore: Evaluating text generation with BERT. In: ICLR, April 26-30, 2020
2020
Later among the works it cites.
Zhao, T., Lala, D., Kawahara, T.: Designing precise and robust dialogue response evaluators. In: 58th ACL, July 5-10, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.