Universal Sentence Encoder
Original
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Later among the works it cites.
The price of debiasing automatic metrics in natural language evaluation
Original
Arun Tejasvi Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Later among the works it cites.
Learning to Evaluate Image Captioning. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018 . IEEE Computer Society, 5804–5812
Yin Cui, Guandao Yang, Andreas Veit, Xun Huang, and Serge J. Belongie. 2018 · 2018
Later among the works it cites.
Deep learning in natural language processing
Li Deng and Yang Liu. 2018 · 2018
Later among the works it cites.
Understanding Back-Translation at Scale. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics, 489–500
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Later among the works it cites.
Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer. 2018 · 2018
Later among the works it cites.
Meteor++: Incorporating Copy Knowledge into Machine Translation Evaluation. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers . Association for Computational Linguistics, Belgium, Brussels, 740–745
Yinuo Guo, Chong Ruan, and Junfeng Hu. 2018 · 2018
Later among the works it cites.
The NarrativeQA Reading Comprehension Challenge
Tomás Kociský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018 · 2018
Later among the works it cites.
The challenging task of summary evaluation: an overview
Elena Lloret, Laura Plaza, and Ahmet Aker. 2018 · 2018
Later among the works it cites.
An efficient framework for learning sentence representations. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net
Lajanugen Logeswaran and Honglak Lee. 2018 · 2018
Later among the works it cites.
Results of the WMT18 Metrics Shared Task: Both characters and embeddings achieve good performance. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers . Association for Computational Linguistics, Belgium, Brussels, 671–688
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Later among the works it cites.
Towards a Better Metric for Evaluating Question Generation Systems. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics, 3950–3959
Preksha Nema and Mitesh M. Khapra. 2018 · 2018
Later among the works it cites.
SemEval-2018 Task 11: Machine Comprehension Using Commonsense Knowledge. In Proceedings of The 12th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2018, New Orleans, Louisiana, USA, June 5-6, 2018 , Marianna Apidianaki, Saif M. Mohammad, Jonathan May, Ekaterina Shutova, Steven Bethard, and Marine Carpuat (Eds.). Association for Computational Linguistics, 747–757
Simon Ostermann, Michael Roth, Ashutosh Modi, Stefan Thater, and Manfred Pinkal. 2018 · 2018
Later among the works it cites.
ITER: Improving Translation Edit Rate through Optimizable Edit Costs. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, WMT 2018, Belgium, Brussels, October 31 - November 1, 2018 , Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno-Yepes, Philipp Koehn, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana L. Neves, Matt Post, Lucia Specia, Marco Turchi, and Karin Verspoor (Eds.). Association for Computational Linguistics, 746–750
Joybrata Panja and Sudip Kumar Naskar. 2018 · 2018
Later among the works it cites.
Deep Contextualized Word Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers) , Marilyn A. Walker, Heng Ji, and Amanda Stent (Eds.). Association for Computational Linguistics, 2227–2237
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
A Call for Clarity in Reporting BLEU Scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, WMT 2018, Belgium, Brussels, October 31 - November 1, 2018 , Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno-Yepes, Philipp Koehn, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana L. Neves, Matt Post, Lucia Specia, Marco Turchi, and Karin Verspoor (Eds.). Association for Computational Linguistics, 186–191
Matt Post. 2018 · 2018
Later among the works it cites.
A Structured Review of the Validity of BLEU
Ehud Reiter. 2018 · 2018
Later among the works it cites.
Learning-based Composite Metrics for Improved Caption Evaluation. In Proceedings of ACL 2018, Melbourne, Australia, July 15-20, 2018, Student Research Workshop , Vered Shwartz, Jeniya Tabassum, Rob Voigt, Wanxiang Che, Marie-Catherine de Marneffe, and Malvina Nissim (Eds.). Association for Computational Linguistics, 14–20
Naeha Sharif, Lyndon White, Mohammed Bennamoun, and Syed Afaq Ali Shah. 2018a · 2018
Later among the works it cites.
NNEval: Neural Network Based Evaluation Metric for Image Captioning. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part VIII (Lecture Notes in Computer Science, Vol. 11212) , Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss (Eds.). Springer, 39–55
Naeha Sharif, Lyndon White, Mohammed Bennamoun, and Syed Afaq Ali Shah. 2018b · 2018
Later among the works it cites.
RUSE: Regressor Using Sentence Embeddings for Automatic Machine Translation Evaluation. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, WMT 2018, Belgium, Brussels, October 31 - November 1, 2018 , Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno-Yepes, Philipp Koehn, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana L. Neves, Matt Post, Lucia Specia, Marco Turchi, and Karin Verspoor (Eds.). Association for Computational Linguistics, 751–758
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2018 · 2018
Later among the works it cites.
RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018 , Sheila A. McIlraith and Kilian Q. Weinberger (Eds.). AAAI Press, 722–729
Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018 · 2018
Later among the works it cites.
ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Melbourne, Australia, 451–462
John Wieting and Kevin Gimpel. 2018 · 2018
Later among the works it cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) . Association for Computational Linguistics, New Orleans, Louisiana, 1112–1122
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
Recent trends in deep learning based natural language processing
Tom Young, Devamanyu Hazarika, Soujanya Poria, and Erik Cambria. 2018 · 2018
Later among the works it cites.
Personalizing Dialogue Agents: I have a dog, do you have pets too?. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Melbourne, Australia, 2204–2213
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Later among the works it cites.
Evaluating Question Answering Evaluation. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, MRQA@EMNLP 2019, Hong Kong, China, November 4, 2019 , Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen (Eds.). Association for Computational Linguistics, 119–124
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Later among the works it cites.
WMDO: Fluency-based Word Mover’s Distance for Machine Translation Evaluation. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1) . Association for Computational Linguistics, Florence, Italy, 494–500
Julian Chow, Lucia Specia, and Pranava Madhyastha. 2019 · 2019
Later among the works it cites.
Sentence Mover’s Similarity: Automatic Evaluation for Multi-Sentence Texts. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , Anna Korhonen, David R. Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics, 2748–2760
Elizabeth Clark, Asli Çelikyilmaz, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Handling Divergent Reference Texts when Evaluating Table-to-Text Generation. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , Anna Korhonen, David R. Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics, 4884–4895
Bhuwan Dhingra, Manaal Faruqui, Ankur P. Parikh, Ming-Wei Chang, Dipanjan Das, and William W. Cohen. 2019 · 2019
Later among the works it cites.
Evaluating the state-of-the-art of End-to-End Natural Language Generation: The E2E NLG challenge
Ondrej Dusek, Jekaterina Novikova, and Verena Rieser. 2020 · 2019
Later among the works it cites.
Word Embedding-Based Automatic MT Evaluation Metric using Word Position Information. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, Minneapolis, Minnesota, 1874–1883
Hiroshi Echizen’ya, Kenji Araki, and Eduard Hovy. 2019 · 2019
Later among the works it cites.
Taking MT Evaluation Metrics to Extremes: Beyond Correlation with Human Judgments
Marina Fomicheva and Lucia Specia. 2019 · 2019
Later among the works it cites.
Approximating Interactive Human Evaluation with Self-Play for Open-Domain Dialog Systems. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8-14 December 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 13658–13669
Asma Ghandeharioun, Judy Hanwen Shen, Natasha Jaques, Craig Ferguson, Noah Jones, Àgata Lapedriza, and Rosalind W. Picard. 2019 · 2019
Later among the works it cites.
Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings
Original
Sarik Ghazarian, Johnny Tian-Zheng Wei, Aram Galstyan, and Nanyun Peng. 2019 · 2019
Later among the works it cites.
Meteor++ 2.0: Adopt Syntactic Level Paraphrase Knowledge into Machine Translation Evaluation. In Proceedings of the Fourth Conference on Machine Translation, WMT 2019, Florence, Italy, August 1-2, 2019 - Volume 2: Shared Task Papers, Day 1 , Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno-Yepes, Philipp Koehn, André Martins, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana L. Neves, Matt Post, Marco Turchi, and Karin Verspoor (Eds.). Association for Computational Linguistics, 501–506
Yinuo Guo and Junfeng Hu. 2019 · 2019
Later among the works it cites.
Evaluating Rewards for Question Generation Models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, 2278–2283
Tom Hosking and Sebastian Riedel. 2019 · 2019
Later among the works it cites.
Neural Text Summarization: A Critical Evaluation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . Association for Computational Linguistics, Hong Kong, China, 540–551
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Later among the works it cites.
Reasoning Over Paragraph Effects in Situations. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, MRQA@EMNLP 2019, Hong Kong, China, November 4, 2019 , Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen (Eds.). Association for Computational Linguistics, 58–62
Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner. 2019 · 2019
Later among the works it cites.
YiSi - a Unified Semantic MT Quality Evaluation and Estimation Metric for Languages with Different Levels of Available Resources. In Proceedings of the Fourth Conference on Machine Translation, WMT 2019, Florence, Italy, August 1-2, 2019 - Volume 2: Shared Task Papers, Day 1 , Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno-Yepes, Philipp Koehn, André Martins, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana L. Neves, Matt Post, Marco Turchi, and Karin Verspoor (Eds.). Association for Computational Linguistics, 507–513
Chi-kiu Lo. 2019 · 2019
Later among the works it cites.
Results of the WMT19 Metrics Shared Task: Segment-Level and Strong MT Systems Pose Big Challenges. In Proceedings of the Fourth Conference on Machine Translation, WMT 2019, Florence, Italy, August 1-2, 2019 - Volume 2: Shared Task Papers, Day 1 , Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno-Yepes, Philipp Koehn, André Martins, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana L. Neves, Matt Post, Marco Turchi, and Karin Verspoor (Eds.). Association for Computational Linguistics, 62–90
Qingsong Ma, Johnny Wei, Ondrej Bojar, and Yvette Graham. 2019 · 2019
Later among the works it cites.
Putting Evaluation in Context: Contextual Embeddings Improve Machine Translation Evaluation. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , Anna Korhonen, David R. Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics, 2799–2808
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019 · 2019
Later among the works it cites.
Re-Evaluating ADEM: A Deeper Look at Scoring Dialogue Responses. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019 . AAAI Press, 6220–6227
Ananya B. Sai, Mithun Das Gupta, Mitesh M. Khapra, and Mukundhan Srinivasan. 2019 · 2019
Later among the works it cites.
What makes a good conversation? How controllable attributes affect human judgments. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, 1702–1723
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 2019
Later among the works it cites.
Machine Translation Evaluation with BERT Regressor
Original
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2019 · 2019
Later among the works it cites.
Webnlg challenge: Human evaluation results
Anastasia Shimorina, Claire Gardent, Shashi Narayan, and Laura Perez-Beltrachini. 2019 · 2019
Later among the works it cites.
EED: Extended Edit Distance Measure for Machine Translation. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1) . Association for Computational Linguistics, Florence, Italy, 514–520
Peter Stanchev, Weiyue Wang, and Hermann Ney. 2019 · 2019
Later among the works it cites.
Beyond BLEU: Training Neural Machine Translation with Semantic Similarity. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers , Anna Korhonen, David R. Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics, 4344–4355
John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, and Graham Neubig. 2019 · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8-14 December 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 5754–5764
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
BERTScore: Evaluating Text Generation with BERT
Original
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
Linguistic Appropriateness and Pedagogic Usefulness of Reading Comprehension Questions. In Proceedings of The 12th Language Resources and Evaluation Conference . European Language Resources Association, Marseille, France, 1753–1762
Andrea Horbach, Itziar Aldabe, Marie Bexte, Oier Lopez de Lacalle, and Montse Maritxalar. 2020 · 2020
Closest in time.
Towards a Decomposable Metric for Explainable Evaluation of Text Generation from AMR
Original
Juri Opitz and Anette Frank. 2020 · 2020
Closest in time.
How To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics
Original
Prasanna Parthasarathi, Joelle Pineau, and Sarath Chandar. 2020 · 2020
Closest in time.
Query-based summarization of discussion threads
Suzan Verberne, Emiel Krahmer, Sander Wubben, and Antal van den Bosch. 2020 · 2020
Closest in time.
Better Hypothesis Testing for Statistical Machine Translation: Controlling for Optimizer Instability. In The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 19-24 June, 2011, Portland, Oregon, USA - Short Papers . The Association for Computer Linguistics, 176–181
Jonathan H. Clark, Chris Dyer, Alon Lavie, and Noah A. Smith. 2011 · 2031
Closest in time.
Adversarial Examples for Evaluating Reading Comprehension Systems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017 , Martha Palmer, Rebecca Hwa, and Sebastian Riedel (Eds.). Association for Computational Linguistics, 2021–2031
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.
Discrete vs. Continuous Rating Scales for Language Evaluation in NLP. In The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 19-24 June, 2011, Portland, Oregon, USA - Short Papers . The Association for Computer Linguistics, 230–235
Anja Belz and Eric Kow. 2011 · 2040
Closest in time.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 (JMLR Workshop and Conference Proceedings, Vol. 37) , Francis R. Bach and David M. Blei (Eds.). JMLR.org, 2048–2057
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.
A New Quantitative Quality Measure for Machine Translation Systems. In COLING 1992 Volume 2: The 15th International Conference on Computational Linguistics
Keh-Yih Su, Ming-Wen Wu, and Jing-Shin Chang. 1992 · 2067
Closest in time.
deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing, ACL 2015, July 26-31, 2015, Beijing, China, Volume 2: Short Papers . The Association for Computer Linguistics, 445–450
Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk, Margaret Mitchell, Jianfeng Gao, and Bill Dolan. 2015 · 2073
Closest in time.
Comparing Automatic Evaluation Measures for Image Description. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, ACL 2014, June 22-27, 2014, Baltimore, MD, USA, Volume 2: Short Papers . The Association for Computer Linguistics, 452–457
Desmond Elliott and Frank Keller. 2014 · 2074
Closest in time.