Fetching the paper…
Reading the bibliography…
Recent text generation research has increasingly focused on open-ended domains such as story and poetry generation.
Statistical power analysis for the behavioral sciences
Jacob Cohen. 1988 · 1988
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluation of text generation: A survey
A. Çelikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Non-expert evaluation of summarization systems is risky
Dan Gillick and Yang Liu. 2010 · 2010
Earlier work this paper cites.
Amazon mechanical turk: Gold mine or coal mine?
Karën Fort, Gilles Adda, and K. Bretonnel Cohen. 2011 · 2011
Earlier work this paper cites.
Evaluating online labor markets for experimental research: Amazon.com's mechanical turk
Adam J. Berinsky, Gregory A. Huber, and Gabriel S. Lenz. 2012 · 2012
Earlier work this paper cites.
Human evaluation of grammatical error correction systems
Roman Grundkiewicz, Marcin Junczys-Dowmunt, and Edward Gillian. 2015 · 2015
Earlier work this paper cites.
How do humans evaluate machine translation
Francisco Guzmán, Ahmed Abdelali, Irina Temnikova, Hassan Sajjad, and Stephan Vogel. 2015 · 2015
Earlier work this paper cites.
Measuring the prevalence of problematic respondent behaviors among MTurk, campus, and community participants
Elizabeth A. Necka, Stephanie Cacioppo, Greg J. Norman, and John T. Cacioppo. 2016 · 2016
Earlier work this paper cites.
Can machine translation systems be evaluated by the crowd alone
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2017 · 2017
Earlier work this paper cites.
Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017 · 2017
Earlier work this paper cites.
Mturk character misrepresentation: Assessment and solutions
K. Wessling, J. Huber, and O. Netzer. 2017 · 2017
Earlier work this paper cites.
Discourse-aware neural rewards for coherent text generation
Antoine Bosselut, Asli Celikyilmaz, Xiaodong He, Jianfeng Gao, Po-Sen Huang, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
Neural text generation in stories using entity representations as context
Elizabeth Clark, Yangfeng Ji, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
What’s this movie about? a joint neural network architecture for movie content analysis
Philip John Gorinski and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Learning to write with cooperative discriminators
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
Delete, retrieve, generate: a simple approach to sentiment and style transfer
Juncen Li, Robin Jia, He He, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Embedding multimodal relational data for knowledge base completion
Pouya Pezeshkpour, Liyan Chen, and Sameer Singh. 2018 · 2018
Earlier work this paper cites.
Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer
Sudha Rao and Joel Tetreault. 2018 · 2018
Earlier work this paper cites.
A structured review of the validity of BLEU
Ehud Reiter. 2018 · 2018
Earlier work this paper cites.
Attaining the unattainable? reassessing claims of human parity in neural machine translation
Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018 · 2018
Earlier work this paper cites.
A neural approach to pun generation
Zhiwei Yu, Jiwei Tan, and Xiaojun Wan. 2018 · 2018
Earlier work this paper cites.
Strategies for structuring story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2019 · 2019
Earlier work this paper cites.
Judge the judges: A large-scale evaluation study of neural language models for online review generation
Cristina Garbacea, Samuel Carton, Shiyan Yan, and Qiaozhu Mei. 2019 · 2019
Earlier work this paper cites.
Pun generation with surprise
He He, Nanyun Peng, and Percy Liang. 2019 · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Visual story post-editing
Ting-Yao Hsu, Chieh-Yang Huang, Yen-Chia Hsu, and Ting-Hao Huang. 2019 · 2019
Cited alongside, same era.
Unsupervised hierarchical story infilling
Daphne Ippolito, David Grangier, Chris Callison-Burch, and Douglas Eck. 2019 · 2019
Cited alongside, same era.
Complexity-weighted loss and diverse reranking for sentence simplification
Reno Kriz, João Sedoc, Marianna Apidianaki, Carolina Zheng, Gaurav Kumar, Eleni Miltsakaki, and Chris Callison-Burch. 2019 · 2019
Cited alongside, same era.
Towards explainable NLP: A generative explanation framework for text classification
Hui Liu, Qingyu Yin, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
Improving neural story generation by targeted common sense grounding
Huanru Henry Mao, Bodhisattwa Prasad Majumder, Julian McAuley, and Garrison Cottrell. 2019 · 2019
Enabling language models to fill in the blanks
Chris Donahue, Mina Lee, and Percy Liang. 2020 · 2020
Later among the works it cites.
Video2Commonsense: Generating commonsense descriptions to enrich video captioning
Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2020 · 2020
Later among the works it cites.
Content planning for neural story generation with aristotelian rescoring
Seraphina Goldfarb-Tarrant, Tuhin Chakrabarty, Ralph Weischedel, and Nanyun Peng. 2020 · 2020
Later among the works it cites.
Neural syntactic preordering for controlled paraphrase generation
Tanya Goyal and Greg Durrett. 2020 · 2020
Later among the works it cites.
A knowledge-enhanced pretraining model for commonsense story generation
Jian Guan, Fei Huang, Zhihao Zhao, Xiaoyan Zhu, and Minlie Huang. 2020 · 2020
Later among the works it cites.
Substance over Style: Document-Level Targeted Content Transfer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A panel for lemons? positivity bias, reputation systems and data quality on MTurk
Ted Matherly. 2019 · 2019
Cited alongside, same era.
Evaluating style transfer for text
Remi Mir, Bjarke Felbo, Nick Obradovich, and Iyad Rahwan. 2019 · 2019
Cited alongside, same era.
The steep road to happily ever after: an analysis of current visual storytelling models
Yatri Modi and Natalie Parde. 2019 · 2019
Cited alongside, same era.
Toward a better story end: Collecting human evaluation with reasons
Yusuke Mori, Hiroaki Yamane, Yusuke Mukuta, and Tatsuya Harada. 2019 · 2019
Cited alongside, same era.
Counterfactual story reasoning and generation
Lianhui Qin, Antoine Bosselut, Ari Holtzman, Chandra Bhagavatula, Elizabeth Clark, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Allison Hegel, Sudha Rao, Asli Celikyilmaz, and Bill Dolan. 2020 · 2020
Later among the works it cites.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Later among the works it cites.
Automatic detection of generated text is easiest when humans are fooled
Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2020 · 2020
Later among the works it cites.
Best practices for crowd-based evaluation of German summarization: Comparing crowd, expert and automatic evaluation
Neslihan Iskender, Tim Polzehl, and Sebastian Möller. 2020 · 2020
Later among the works it cites.
Neural CRF model for sentence alignment in text simplification
Chao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong, and Wei Xu. 2020 · 2020
Later among the works it cites.
The shape of and solutions to the mturk quality crisis
Ryan Kennedy, Scott Clifford, Tyler Burleigh, Philip D. Waggoner, Ryan Jewell, and Nicholas J. G. Winter. 2020 · 2020
Later among the works it cites.
Reformulating unsupervised style transfer as paraphrase generation
Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020 · 2020
Later among the works it cites.
Learning to generate multiple style transfer outputs for an input sentence
Kevin Lin, Ming-Yu Liu, Ming-Ting Sun, and Jan Kautz. 2020 · 2020
Later among the works it cites.
Zero-shot crosslingual sentence simplification
Jonathan Mallinson, Rico Sennrich, and Mirella Lapata. 2020 · 2020
Later among the works it cites.
Sparse text generation
Pedro Henrique Martins, Zita Marinho, and André F. T. Martins. 2020 · 2020
Later among the works it cites.
Back to the future: Unsupervised backprop-based decoding for counterfactual and abductive commonsense reasoning
Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras, Antoine Bosselut, and Yejin Choi. 2020 · 2020
Later among the works it cites.
PlotMachines: Outline-conditioned generation with dynamic plot state tracking
Hannah Rashkin, Asli Celikyilmaz, Yejin Choi, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
MEGATRON-CNTRL: Controllable story generation with external knowledge using large-scale language models
Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri, Pascale Fung, Anima Anandkumar, and Bryan Catanzaro. 2020 · 2020
Later among the works it cites.
Routing enforced generative model for recipe generation
Zhiwei Yu, Hongyu Zang, and Xiaojun Wan. 2020 · 2020
Later among the works it cites.
Towards document-level human MT evaluation: On the issues of annotator agreement, effort and misevaluation
Sheila Castilho. 2021 · 2021
Closest in time.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021 · 2021
Closest in time.
Genie: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, and Daniel S. Weld. 2021 · 2021
Closest in time.
After the bot scare: Understanding what’s been happening with data collection on mturk and how to stop it
Aaron Moss and Leib Litman · 2021
Closest in time.
Human evaluation of automatically generated text: Current trends and best practice guidelines
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2021 · 2021
Closest in time.
Towards generating long and coherent text with multi-level latent variable models
Dinghan Shen, Asli Celikyilmaz, Yizhe Zhang, Liqun Chen, Xin Wang, Jianfeng Gao, and Lawrence Carin. 2019 · 2089
Closest in time.