Fetching the paper…
Reading the bibliography…
As LLMs become commonplace, machine-generated text has the potential to flood the internet with spam, social media bots, and valueless content.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
The χ \chi 2 test of goodness of fit
William G Cochran · 1952
Earlier work this paper cites.
Distribution-free Statistical Tests
James V. Bradley · 1960
Earlier work this paper cites.
Electronic marking and identification techniques to discourage document copying
Jack T Brassil, Steven Low, Nicholas F Maxemchuk, and Lawrence O’Gorman · 1995
Earlier work this paper cites.
Natural Language Watermarking: Design, Analysis, and a Proof-of-Concept Implementation
Mikhail J. Atallah, Victor Raskin, Michael Crogan, Christian Hempelmann, Florian Kerschbaum, Dina Mohamed, and Sanket Naik · 2001
Earlier work this paper cites.
Attacking Neural Text Detectors
Max Wolff and Stuart Wolff · 2002
Earlier work this paper cites.
Natural Language Watermarking Using Semantic Substitution for Chinese Text
Yuei-Lin Chiang, Lu-Ping Chang, Wen-Tai Hsieh, and Wen-Chih Chen · 2004
Earlier work this paper cites.
The hiding virtues of ambiguity: Quantifiably resilient watermarking of natural language text through synonym substitutions
Umut Topkara, Mercan Topkara, and Mikhail J. Atallah · 2006
Earlier work this paper cites.
Identifying real or fake articles: Towards better language modeling
Sameer Badaskar, Sachin Agarwal, and Shilpa Arora · 2008
Earlier work this paper cites.
Detecting fake content with relative entropy scoring
Thomas Lavergne, Tanguy Urvoy, and François Yvon · 2008
Earlier work this paper cites.
Detection of artificial texts
EA Grechnikov, GG Gusev, AA Kustarev, and AM Raigorodsky · 2009
Earlier work this paper cites.
Watermarking the Outputs of Structured Prediction with an application in Statistical Machine Translation
Ashish Venugopal, Jakob Uszkoreit, David Talbot, Franz Och, and Juri Ganitkevitch · 2011
Earlier work this paper cites.
Computer-generated text detection using machine learning: A systematic review
Daria Beresneva · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Generating Steganographic Text with LSTMs
Tina Fang, Martin Jaggi, and Katerina Argyraki · 2017
Earlier work this paper cites.
Towards Near-imperceptible Steganographic Text
Falcon Dai and Zheng Cai · 2019
Earlier work this paper cites.
Gltr: Statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander M Rush · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Earlier work this paper cites.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston · 2019
Cited alongside, same era.
Neural Linguistic Steganography
Zachary Ziegler, Yuntian Deng, and Alexander Rush · 2019
Cited alongside, same era.
How Effectively Can Machines Defend Against Machine-Generated Fake News? An Empirical Study
Meghana Moorthy Bhat and Srinivasan Parthasarathy · 2020
Cited alongside, same era.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Limits of Detecting Text Generated by Large-Scale Language Models
Lav R. Varshney, Nitish Shirish Keskar, and Richard Socher · 2020
Cited alongside, same era.
Three bricks to consolidate watermarks for large language models
Pierre Fernandez, Antoine Chaffin, Karim Tit, Vivien Chappelier, and Teddy Furon · 2023
Closest in time.
Radar: Robust ai-text detection via adversarial learning
Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho · 2023
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Closest in time.
A Watermark for Large Language Models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Closest in time.
Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data Hiding
Sahar Abdelnabi and Mario Fritz · 2021
Cited alongside, same era.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
Meteor: Cryptographically Secure Steganography for Realistic Distributions, 2021
Gabriel Kaptchuk, Tushar M. Jois, Matthew Green, and Aviel Rubin · 2021
Cited alongside, same era.
MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui · 2021
Cited alongside, same era.
Using AI to Scale Spear Phishing, August 2021
Bruce Schneier · 2021
Cited alongside, same era.
Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods
Evan Crothers, Nathalie Japkowicz, and Herna Viktor · 2022
Cited alongside, same era.
The Ethical Need for Watermarks in Machine-Generated Language
Alexei Grinbaum and Laurynas Adomaitis · 2022
Cited alongside, same era.
Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer · 2023
Closest in time.
Who wrote this code? watermarking for code generation
Taehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, and Gunhee Kim · 2023
Closest in time.
GPT detectors are biased against non-native English writers
Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou · 2023
Closest in time.
DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn · 2023
Closest in time.
People are using A.I. chatbots to write Amazon reviews
Annie Palmer · 2023
Closest in time.
Can AI-Generated Text be Reliably Detected?
Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi · 2023
Closest in time.
The Curse of Recursion: Training on Generated Data Makes Models Forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2023
Closest in time.
The Science of Detecting LLM-Generated Texts
Ruixiang Tang, Yu-Neng Chuang, and Xia Hu · 2023
Closest in time.
Gptzero update v1, January 2023
Edward Tian · 2023
Closest in time.
The alignment handbook
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Alexander M. Rush, and Thomas Wolf · 2023
Closest in time.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality, March 2023
Vicuna-Team · 2023
Closest in time.
AI Is Tearing Wikipedia Apart, May 2023
Claire Woodcock · 2023
Closest in time.
Robust natural language watermarking through invariant features
KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak · 2023
Closest in time.