Fetching the paper…
Reading the bibliography…
The burgeoning generative capabilities of large language models (LLMs) have raised growing concerns about abuse, demanding automatic machine-generated text detectors.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
WordNet: A lexical database for English
George A. Miller. 1992 · 1992
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Universal mobile information retrieval
David Machado, Tiago Barbosa, Sebastião Pais, Bruno Martins, and Gaël Dias. 2009 · 2009
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Extractive summarization: Limits, compression, generalized model and heuristics
Rakesh Verma and Daniel Lee. 2017 · 2017
Earlier work this paper cites.
Switchout: an efficient data augmentation algorithm for neural machine translation
Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Gltr: Statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander M Rush. 2019 · 2019
Earlier work this paper cites.
Gpt-2 output dataset
OpenAI. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Generalized reference evapotranspiration models with limited climatic data based on random forest and gene expression programming in guangxi, china
Sheng Wang, Jinjiao Lian, Yuzhong Peng, Baoqing Hu, and Hongsong Chen. 2019 · 2019
Earlier work this paper cites.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. 2019 · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Earlier work this paper cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau Alexis, Khandelwal Kartikay, Goyal Naman, Chaudhary Vishrav, Wenzek Guillaume, Guzmán Francisco, Grave Edouard, Ott Myle, Zettlemoyer Luke, and Stoyanov Veselin. 2020 · 2020
Earlier work this paper cites.
Yake! keyword extraction from single documents using multiple local features
Ricardo Campos, Vítor Mangaravite, Arian Pasquali, Alípio Jorge, Célia Nunes, and Adam Jatowt. 2020 · 2020
Earlier work this paper cites.
Mixtext: Linguistically-informed interpolation of hidden space for semi-supervised text classification
Jiaao Chen, Zichao Yang, and Diyi Yang. 2020 · 2020
Earlier work this paper cites.
Robust spammer detection by nash reinforcement learning
Yingtong Dou, Guixiang Ma, Philip S Yu, and Sihong Xie. 2020 · 2020
Earlier work this paper cites.
Understanding the limitations of conditional generative models
Ethan Fetaya, Joern-Henrik Jacobsen, Will Grathwohl, and Richard Zemel. 2020 · 2020
Earlier work this paper cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis Mike, Liu Yinhan, Goyal Naman, Ghazvininejad Marjan, Mohamed Abdelrahman, Levy Omer, Stoyanov Ves, and Zettlemoyer Luke. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Csi: Novelty detection via contrastive learning on distributionally shifted instances
Jihoon Tack, Sangwoo Mo, Jongheon Jeong, and Jinwoo Shin. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020 · 2020
Cited alongside, same era.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020 · 2020
Cited alongside, same era.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023 · 2023
Later among the works it cites.
Understanding the effects of rlhf on llm generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. 2023 · 2023
Later among the works it cites.
Coco: Coherence-enhanced machine-generated text detection under low resource with contrastive learning
Xiaoming Liu, Zhaohan Zhang, Yichen Wang, Hang Pu, Yu Lan, and Chao Shen. 2023 · 2023
Later among the works it cites.
Smaller language models are better black-box machine-generated text detectors
Fatemehsadat Mireshghallah, Justus Mattern, Sicun Gao, Reza Shokri, and Taylor Berg-Kirkpatrick. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Supervised contrastive learning for pre-trained language model fine-tuning
Beliz Gunel, Jingfei Du, Alexis Conneau, and Ves Stoyanov. 2021 · 2021
Cited alongside, same era.
Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX
Ben Wang. 2021 · 2021
Cited alongside, same era.
Few-shot text classification with triplet networks, data augmentation, and curriculum learning
Jason Wei, Chengyu Huang, Soroush Vosoughi, Yu Cheng, and Shiqi Xu. 2021 · 2021
Cited alongside, same era.
Contrastive out-of-distribution detection for pretrained transformers
Wenxuan Zhou, Fangyu Liu, and Muhao Chen. 2021 · 2021
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. 2022 · 2022
Cited alongside, same era.
Mask-then-fill: A flexible and effective data augmentation framework for event extraction
Jun Gao, Changlong Yu, Wei Wang, Huan Zhao, and Ruifeng Xu. 2022 · 2022
Cited alongside, same era.
Alp: Data augmentation using lexicalized pcfgs for few-shot text classification
Hazel H Kim, Daecheol Woo, Seong Joon Oh, Jeong-Won Cha, and Yo-Sub Han. 2022 · 2022
Cited alongside, same era.
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Ai text classifier
OpenAI. 2023 · 2023
Later among the works it cites.
Automatic prompt augmentation and selection with chain-of-thought from labeled data
KaShun Shum, Shizhe Diao, and Tong Zhang. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Christoforos Vasilatos, Manaar Alam, Talal Rahwan, Yasir Zaki, and Michail Maniatakos. 2023 · 2023
Later among the works it cites.
Gpt-who: An information density-based machine-generated text detector
Saranya Venkatraman, Adaku Uchendu, and Dongwon Lee. 2023 · 2023
Later among the works it cites.
Ghostbuster: Detecting text ghostwritten by large language models
Vivek Verma, Eve Fleisig, Nicholas Tomlin, and Dan Klein. 2023 · 2023
Later among the works it cites.
SeqXGPT: Sentence-level AI-generated text detection
Pengyu Wang, Linyang Li, Ke Ren, Botian Jiang, Dong Zhang, and Xipeng Qiu. 2023 · 2023
Later among the works it cites.
LLMDet: A third party large language models generated text detection tool
Kangxi Wu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua. 2023b · 2023
Later among the works it cites.
Dna-gpt: Divergent n-gram analysis for training-free detection of gpt-generated text
Xianjun Yang, Wei Cheng, Linda Petzold, William Yang Wang, and Haifeng Chen. 2023 · 2023
Later among the works it cites.
Twhin-bert: A socially-enriched pre-trained language model for multilingual tweet representations at twitter
Xinyang Zhang, Yury Malkov, Omar Florez, Serim Park, Brian McWilliams, Jiawei Han, and Ahmed El-Kishky. 2023 · 2023
Later among the works it cites.
Fast-detectGPT: Efficient zero-shot detection of machine-generated text via conditional probability curvature
Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. 2024 · 2024
Closest in time.
Raidar: generative AI detection via rewriting
Chengzhi Mao, Carl Vondrick, Hao Wang, and Junfeng Yang. 2024 · 2024
Closest in time.
Yuhui Shi, Qiang Sheng, Juan Cao, Hao Mi, Beizhe Hu, and Danding Wang. 2024 · 2024
Closest in time.
Stumbling blocks: Stress testing the robustness of machine-generated text detectors under attacks
Yichen Wang, Shangbin Feng, Abe Bohan Hou, Xiao Pu, Chao Shen, Xiaoming Liu, Yulia Tsvetkov, and Tianxing He. 2024 · 2024
Closest in time.