Fetching the paper…
Reading the bibliography…
As text generated by large language models proliferates, it becomes vital to understand how humans engage with such text, and whether or not they are able to detect when the text they are reading did not originate with a human writer.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 1901
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Keskar, N. S.; McCann, B.; Varshney, L. R.; Xiong, C.; and Socher, R. 2019 · 1909
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
The new york times annotated corpus
Sandhaus, E. 2008 · 2008
Earlier work this paper cites.
Corpus of Presidential Speeches
Brown, D. W. 2016 · 2016
Earlier work this paper cites.
Hierarchical Neural Story Generation
Fan, A.; Lewis, M.; and Dauphin, Y. 2018 · 2018
Earlier work this paper cites.
GLTR: Statistical Detection and Visualization of Generated Text
Gehrmann, S.; Strobelt, H.; and Rush, A. M. 2019 · 2019
Earlier work this paper cites.
Recipe1M+: A Dataset for Learning Cross-Modal Embeddings for Cooking Recipes and Food Images
Marin, J.; Biswas, A.; Ofli, F.; Hynes, N.; Salvador, A.; Aytar, Y.; Weber, I.; and Torralba, A. 2019 · 2019
Earlier work this paper cites.
Towards understanding and detecting fake reviews in app stores
Martens, D.; and Maalej, W. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Cited alongside, same era.
Defending Against Neural Fake News
Zellers, R.; Holtzman, A.; Rashkin, H.; Bisk, Y.; Farhadi, A.; Roesner, F.; and Choi, Y. 2019 · 2019
Cited alongside, same era.
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems
Deriu, J.; Tuggener, D.; von Däniken, P.; Campos, J. A.; Rodrigo, A.; Belkacem, T.; Soroa, A.; Agirre, E.; and Cieliebak, M. 2020 · 2020
Cited alongside, same era.
RoFT: A Tool for Evaluating Human Detection of Machine-Generated Text
Dugan, L.; Ippolito, D.; Kirubarajan, A.; and Callison-Burch, C. 2020 · 2020
Cited alongside, same era.
The Curious Case of Neural Text Degeneration
Scarecrow: A framework for scrutinizing machine text
Dou, Y.; Forbes, M.; Koncel-Kedziorski, R.; Smith, N. A.; and Choi, Y. 2021 · 2021
Later among the works it cites.
TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language Models
He, J.; Peng, B.; Liao, Y.; Liu, Q.; and Xiong, D. 2021 · 2021
Later among the works it cites.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Ouyang, L.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; Jiang, X.; Cobbe, K.; Eloundou, T.; Krueger, G.; Button, K.; Knight, M.; Chess, B.; and Schulman, J. 2021 · 2021
Later among the works it cites.
Ethical and social risks of harm from Language Models
Weidinger, L.; Mellor, J.; Rauh, M.; Griffin, C.; Uesato, J.; Huang, P.-S.; Cheng, M.; Glaese, M.; Balle, B.; Kasirzadeh, A.; Kenton, Z.; Brown, S.; Hawkins, W.; Stepleton, T.; Biles, C.; Birhane, A.; Haas, J.; Rimell, L.; Hendricks, L. A.; Isaac, W.; Legassick, S.; Irving, G.; and Gabriel, I. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; and Choi, Y. 2020 · 2020
Cited alongside, same era.
Automatic Detection of Generated Text is Easiest when Humans are Fooled
Ippolito, D.; Duckworth, D.; Callison-Burch, C.; and Eck, D. 2020 · 2020
Cited alongside, same era.
All That’s ‘Human’Is Not Gold: Evaluating Human Evaluation of Generated Text
Clark, E.; August, T.; Serrano, S.; Haduong, N.; Gururangan, S.; and Smith, N. A. 2021 · 2021
Cited alongside, same era.
Trading Off Diversity and Quality in Natural Language Generation
Zhang, H.; Duckworth, D.; Ippolito, D.; and Neelakantan, A. 2021 · 2021
Later among the works it cites.
How Human is Human Evaluation? Improving the Gold Standard for NLG with Utility Theory
Ethayarajh, K.; and Jurafsky, D. 2022 · 2022
Closest in time.
Training Compute-Optimal Large Language Models
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; Casas, D. d. L.; Hendricks, L. A.; Welbl, J.; Clark, A.; et al. 2022 · 2022
Closest in time.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2022 · 2022
Closest in time.