Fetching the paper…
Reading the bibliography…
Storytelling is an integral part of human experience and plays a crucial role in social interactions.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
CTRL: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
A new measure of rank correlation
Maurice G. Kendall. 1938 · 1938
Earlier work this paper cites.
Regression Analysis
Evan J. Williams. 1959 · 1959
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
The proof and measurement of association between two things
Charles Spearman. 1961 · 1961
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Joseph L. Fleiss. 1971 · 1971
Earlier work this paper cites.
Issues in psychophysical measurement
Stanley S. Stevens. 1971 · 1971
Earlier work this paper cites.
Tests for comparing elements of a correlation matrix
James H. Steiger. 1980 · 1980
Earlier work this paper cites.
What makes a good story
Allyssa McCabe and Carole Peterson. 1984 · 1984
Earlier work this paper cites.
Controlling the false discovery rate: A practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg. 1995 · 1995
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
The four elements of every successful story
Robert Dickman. 2003 · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Storytelling in science
Stephen Rowcliffe. 2004 · 2004
Earlier work this paper cites.
Interpreting BLEU/NIST scores: How much improvement do we need to have a better system?
Ying Zhang, Stephan Vogel, and Alex Waibel. 2004 · 2004
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Answering the call for a standard reliability measure for coding data
Andrew F. Hayes and Klaus Krippendorff. 2007 · 2007
Earlier work this paper cites.
Computing inter-rater reliability and its variance in the presence of high agreement
Kilem Li Gwet. 2008 · 2008
Earlier work this paper cites.
The power of story: Using storytelling to improve literacy learning
Sara Miller and Lisa Pennycuff. 2008 · 2008
Earlier work this paper cites.
Thinking, Fast and Slow
Daniel Kahneman. 2011 · 2011
Earlier work this paper cites.
Computing inter-rater reliability for observational data: An overview and tutorial
Kevin A. Hallgren. 2012 · 2012
Earlier work this paper cites.
Storytelling on mobile devices for cultural heritage
Vincenzo Lombardo and Rossana Damiano. 2012 · 2012
Earlier work this paper cites.
Story generation with crowdsourced plot graphs
Boyang Li, Stephen Lee-Urban, George Johnston, and Mark Riedl. 2013 · 2013
Earlier work this paper cites.
A Comparison of Cohen’s Kappa and Gwet’s AC1 when calculating inter-rater reliability coefficients: A study conducted with personality disorder samples
Nahathai Wongpakaran, Tinakon Wongpakaran, Danny Wedding, and Kilem L. Gwet. 2013 · 2013
Earlier work this paper cites.
How a creative storytelling intervention can improve medical student attitude towards persons with dementia: A mixed methods study
Daniel R. George, Heather L. Stuckey, and Megan M. Whitehead. 2014 · 2014
Earlier work this paper cites.
Testing for significance of increased correlation with human judgment
Yvette Graham and Timothy Baldwin. 2014 · 2014
Earlier work this paper cites.
The Creative Process: A Computer Model of Storytelling and Creativity
Scott R. Turner. 2014 · 2014
Earlier work this paper cites.
chrF: Character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
A bayesian perspective on Likert scales and central tendency
Igor Douven. 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Cited alongside, same era.
UMAP: Uniform manifold approximation and projection
Leland McInnes, John Healy, Nathaniel Saul, and Lukas Großberger. 2018 · 2018
Cited alongside, same era.
Pingouin: Statistics in Python
Raphael Vallat. 2018 · 2018
Cited alongside, same era.
Scientists rise up against statistical significance
Valentin Amrhein, Sander Greenland, and Blake McShane. 2019 · 2019
Cited alongside, same era.
Why, when and how to adjust your P values?
Mohieddin Jafari and Naser Ansari-Pour. 2019 · 2019
Cited alongside, same era.
Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Cited alongside, same era.
Want to reduce labeling cost? GPT-3 can help
Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021 · 2021
Later among the works it cites.
A temporal variational model for story generation
David Wilmot and Frank Keller. 2021 · 2021
Later among the works it cites.
BARTScore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Later among the works it cites.
Of human criteria and automatic metrics: A benchmark of the evaluation of story generation
Cyril Chhun, Pierre Colombo, Fabian M. Suchanek, and Chloé Clavel. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abandon statistical significance
Blakeley B. McShane, David Gal, Andrew Gelman, Christian Robert, and Jennifer L. Tackett. 2019 · 2019
Cited alongside, same era.
Significance test of increase in correlation for NLP evaluations in Python
Jihyung Moon. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Moving to a world beyond “p<0.05”
Ronald L. Wasserstein, Allen L. Schirm, and Nicole A. Lazar. 2019 · 2019
Cited alongside, same era.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022 · 2022
Later among the works it cites.
Is GPT-3 text indistinguishable from human text? Scarecrow: A framework for scrutinizing machine text
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Chance-corrected agreement coefficients
Aris Fergadis and Benedikt Scheffler. 2022 · 2022
Later among the works it cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Later among the works it cites.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz. 2022 · 2022
Later among the works it cites.
Rewriting results sections in the language of evidence
Stefanie Muff, Erlend B. Nilsen, Robert B. O’Hara, and Chloé R. Nater. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
A novel auto-annotation technique for aspect level sentiment analysis
Muhammad Aasim Qureshi, Muhammad Asif, Mohd Fadzil Hassan, Ghulam Mustafa, Muhammad Khurram Ehsan, Aasim Ali, and Unaza Sajid. 2022 · 2022
Later among the works it cites.
LaMDA: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022 · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022a · 2022
Later among the works it cites.
Ask me anything: A simple strategy for prompting language models
Simran Arora, Avanika Narayan, Mayee F Chen, Laurel Orr, Neel Guha, Kush Bhatia, Ines Chami, and Christopher Re. 2023 · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023 · 2023
Later among the works it cites.
Art or artifice? Large language models and the false promise of creativity
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2023 · 2023
Later among the works it cites.
PaLM: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2023 · 2023
Later among the works it cites.
The glass ceiling of automatic evaluation in natural language generation
Pierre Colombo, Maxime Peyrard, Nathan Noiry, Robert West, and Pablo Piantanida. 2023 · 2023
Later among the works it cites.
Is GPT-3 a good data annotator?
Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia, Boyang Li, Shafiq Joty, and Lidong Bing. 2023 · 2023
Later among the works it cites.
Concours de nouvelles 2023
Edilivre. 2023 · 2023
Later among the works it cites.
A story to sell: The influence of storytelling on consumers’ purchasing behavior
João Ricardo de Oliveira Júnior, Ricardo Limongi, Weng Marc Lim, Jacqueline K. Eastman, and Satish Kumar. 2023 · 2023
Later among the works it cites.
OpenOrca: An open dataset of GPT augmented FLAN reasoning traces
Wing Lian, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and "Teknium". 2023 · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of GPT-4
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. 2023 · 2023
Later among the works it cites.
A prompt pattern catalog to enhance prompt engineering with ChatGPT
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023 · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi. 2023a · 2023
Later among the works it cites.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2023b · 2023
Later among the works it cites.
Dissociating language and thought in large language models
Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2024 · 2024
Closest in time.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. 2024 · 2024
Closest in time.