Fetching the paper…
Reading the bibliography…
With the rapid development and widespread application of Large Language Models (LLMs), the use of Machine-Generated Text (MGT) has become increasingly common, bringing with it potential risks, especially in terms of quality and integrity in fields like news, education, and science.
Real or fake? learning to discriminate machine from human generated text
Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur Szlam. 2019 · 1906
Earlier work this paper cites.
Gltr: Statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander M Rush. 2019 · 1906
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019 · 1908
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Automatic detection of generated text is easiest when humans are fooled
Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2019 · 1911
Earlier work this paper cites.
A mathematical theory of communication
Claude Elwood Shannon. 1948 · 1948
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions and reversals
Vladimir Iosifovich Levenshtein. 1966 · 1966
Earlier work this paper cites.
Practical solutions to the problem of diagonal dominance in kernel document clustering
et al. Greene, Derek. 2006 · 2006
Earlier work this paper cites.
Effects of age and gender on blogging
Jonathan Schler, Moshe Koppel, Shlomo Argamon, and James W Pennebaker. 2006 · 2006
Earlier work this paper cites.
Natural language processing with python: Analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Enron email dataset
CMU. 2015 · 2015
Earlier work this paper cites.
news-please: A generic news crawler and extractor
Felix Hamborg, Norman Meuschke, Corinna Breitinger, and Bela Gipp. 2017 · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Tomáš Kočiskỳ, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018 · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Earlier work this paper cites.
Metaphoria: An algorithmic companion for metaphor creation
Katy Ilonka Gero and Lydia B Chilton. 2019 · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Multi-style generative reading comprehension
Kyosuke Nishida, Itsumi Saito, Kosuke Nishida, Kazutoshi Shinoda, Atsushi Otsuka, Hisako Asano, and Junji Tomita. 2019 · 2019
Earlier work this paper cites.
An overview of overfitting and its solutions
Xue Ying. 2019 · 2019
Earlier work this paper cites.
Covid-qa: A question answering dataset for covid-19
Timo Möller, Anthony Reina, Raghavan Jayakumar, and Malte Pietsch. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Authorship attribution for neural text generation
Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee. 2020 · 2020
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Earlier work this paper cites.
Steam reviews dataset 2021
Najzeko. 2021 · 2021
Cited alongside, same era.
Ted talk transcripts (2006 - 2021)
TheDataBeast. 2021 · 2021
Cited alongside, same era.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2021 · 2021
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. 2022 · 2022
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022 · 2022
Cited alongside, same era.
Authentigpt: Detecting machine-generated text via black-box language models denoising
Zhen Guo and Shangdi Yu. 2023 · 2023
Later among the works it cites.
Mgtbench: Benchmarking machine-generated text detection
Xinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023 · 2023
Later among the works it cites.
Chatgpt is going to change education, not destroy it
Will Douglas Heavenarchive. 2023 · 2023
Later among the works it cites.
Radar: Robust ai-text detection via adversarial learning
Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho. 2023 · 2023
Later among the works it cites.
Open llm leaderboard
Hugging Face. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Protoformer: Embedding prototypes for transformers
Ashkan Farhangi, Ning Sui, Nan Hua, Haiyan Bai, Arthur Huang, and Zhishan Guo. 2022 · 2022
Cited alongside, same era.
Sparks: Inspiration for science writing using language models
Katy Ilonka Gero, Vivian Liu, and Lydia Chilton. 2022 · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022 · 2022
Cited alongside, same era.
Openai models - gpt3.5
OpenAI. 2022 · 2022
Cited alongside, same era.
Gpt4all: Training an assistant-style chatbot with large scale data distillation from gpt-3.5-turbo
Yuvanesh Anand, Zach Nussbaum, Brandon Duderstadt, Benjamin Schmidt, and Andriy Mulyar. 2023 · 2023
Cited alongside, same era.
Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. 2023 · 2023
Cited alongside, same era.
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023 · 2023
Later among the works it cites.
Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense
Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. 2023 · 2023
Later among the works it cites.
Smaller language models are better black-box machine-generated text detectors
Fatemehsadat Mireshghallah, Justus Mattern, Sicun Gao, Reza Shokri, and Taylor Berg-Kirkpatrick. 2023 · 2023
Later among the works it cites.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Can ai-generated text be reliably detected?
Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi. 2023 · 2023
Later among the works it cites.
Rewritelm: An instruction-tuned large language model for text rewriting
Lei Shu, Liangchen Luo, Jayakumar Hoskere, Yun Zhu, Canoee Liu, Simon Tong, Jindong Chen, and Lei Meng. 2023 · 2023
Later among the works it cites.
Alejo Jose G Sison, Marco Tulio Daza, Roberto Gozalo-Brizuela, and Eduardo C Garrido-Merchán. 2023 · 2023
Later among the works it cites.
Stablelm
StabilityAI. 2023 · 2023
Later among the works it cites.
Detectllm: Leveraging log rank information for zero-shot detection of machine-generated text
Jinyan Su, Terry Yue Zhuo, Di Wang, and Preslav Nakov. 2023 · 2023
Later among the works it cites.
Gptzero: An ai text detector
Edward Tian. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Ghostbuster: Detecting text ghostwritten by large language models
Vivek Verma, Eve Fleisig, Nicholas Tomlin, and Dan Klein. 2023 · 2023
Later among the works it cites.
Llmdet: A third party large language models generated text detection tool
Kangxi Wu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua. 2023 · 2023
Later among the works it cites.
On the generalization of training-based chatgpt detection methods
Han Xu, Jie Ren, Pengfei He, Shenglai Zeng, Yingqian Cui, Amy Liu, Hui Liu, and Jiliang Tang. 2023 · 2023
Later among the works it cites.
Gpt paternity test: Gpt generated text detection with gpt genetic inheritance
Xiao Yu, Yuang Qi, Kejiang Chen, Guoqiang Chen, Xi Yang, Pengyuan Zhu, Weiming Zhang, and Nenghai Yu. 2023 · 2023
Later among the works it cites.
The next chapter: A study of large language models in storytelling
Jey Han Lau Zhuohan Xie, Trevor Cohn. 2023 · 2023
Later among the works it cites.
Monitoring ai-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews
Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel A. McFarland, and James Y. Zou. 2024 · 2024
Closest in time.