Fetching the paper…
Reading the bibliography…
Watermarking has offered an effective approach to distinguishing text generated by large language models (LLMs) from human-written text.
Statistical theory of extreme values and some practical applications: A series of lectures , volume 33
Emil Julius Gumbel · 1948
Earlier work this paper cites.
A statistical problem arising in the theory of detection of signals in the presence of noise in a multi-channel system and leading to stable distribution laws
RL Dobrusin · 1958
Earlier work this paper cites.
Efficient estimates and optimum inference procedures in large samples
C Radhakrishna Rao · 1962
Earlier work this paper cites.
Boundary-value problems for random walks and large deviations in function spaces
Aleksandr A Borovkov · 1967
Earlier work this paper cites.
On asymptotically optimal non-parametric criteria
AA Borokov and NM Sycheva · 1968
Earlier work this paper cites.
Martingale approach in the theory of goodness-of-fit tests
Estate V Khmaladze · 1982
Earlier work this paper cites.
A learning algorithm for Boltzmann machines
David H Ackley, Geoffrey E Hinton, and Terrence J Sejnowski · 1985
Earlier work this paper cites.
Hodges-Lehmann asymptotic efficiency of the Kolmogorov and Smirnov goodness-of-fit tests
Ya Yu Nikitin · 1987
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
Large deviations and bahadur efficiency of the Khmaladze-Aki statistic
OA Podkorytova · 1994
Earlier work this paper cites.
WordNet: A lexical database for English
George A Miller · 1995
Earlier work this paper cites.
Asymptotic efficiency of nonparametric tests
ÍÀkov ÍÙr’evich Nikitin · 1995
Earlier work this paper cites.
Convex analysis and minimization algorithms I: Fundamentals , volume 305
Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal · 1996
Earlier work this paper cites.
Applied Cryptography
Bruce Schneier · 1996
Earlier work this paper cites.
Some problems of hypothesis testing leading to infinitely divisible distributions
Yuri I Ingster · 1997
Earlier work this paper cites.
A note on the asymptotic distribution of Berk—Jones type statistics under the null hypothesis
Jon A Wellner and Vladimir Koltchinskii · 2003
Earlier work this paper cites.
Higher criticism for detecting sparse heterogeneous mixtures
David Donoho and Jiashun Jin · 2004
Earlier work this paper cites.
Goodness-of-fit tests via phi-divergences
Leah Jager and Jon A Wellner · 2007
Earlier work this paper cites.
Properties of higher criticism under strong dependence
Peter Hall and Jiashun Jin · 2008
Earlier work this paper cites.
Innovated higher criticism for detecting sparse signals in correlated noise
Peter Hall and Jiashun Jin · 2010
Earlier work this paper cites.
Optimal detection of heterogeneous and heteroscedastic mixtures
T Tony Cai, X Jessie Jeng, and Jiashun Jin · 2011
Earlier work this paper cites.
Robust statistics
Peter J Huber and Elvezio M Ronchetti · 2011
Earlier work this paper cites.
Perturb-and-map random fields: Using discrete optimization to learn and sample from energy models
George Papandreou and Alan L Yuille · 2011
Earlier work this paper cites.
Optimal detection of sparse mixtures against a given null distribution
T Tony Cai and Yihong Wu · 2014
Earlier work this paper cites.
New inequalities for Gamma and Digamma functions
M. R. Farhangdoost and M. Kargar Dolatabadi · 2014
Earlier work this paper cites.
Covariance assisted screening and estimation
Tracy Ke, Jiashun Jin, and Jianqing Fan · 2014
Earlier work this paper cites.
A* sampling
Chris J Maddison, Daniel Tarlow, and Tom Minka · 2014
Earlier work this paper cites.
Higher criticism for large-scale inference, especially for rare and weak effects
David Donoho and Jiashun Jin · 2015
Cited alongside, same era.
The intermediates take it all: Asymptotics of higher criticism statistics and a powerful alternative based on equal local levels
Veronika Gontscharuk, Sandra Landwehr, and Helmut Finner · 2015
Cited alongside, same era.
Higher criticism: p-values and criticism
Jian Li and David Siegmund · 2015
Cited alongside, same era.
Categorical reparameterization with Gumbel-Softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
Rare and weak effects in large-scale inference: Methods and phase diagrams
Jiashun Jin and Zheng Tracy Ke · 2016
Cited alongside, same era.
Human behavior and the principle of least effort: An introduction to human ecology
George Kingsley Zipf · 2016
Mark my words: Analyzing and evaluating language model watermarks
Julien Piet, Chawin Sitawarin, Vivian Fang, Norman Mu, and David Wagner · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
DiPmark: A stealthy, efficient and resilient watermark for large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distribution-free tests for sparse heterogeneous mixtures
Ery Arias-Castro and Meng Wang · 2017
Cited alongside, same era.
Goodness-of-fit-techniques
Ralph B D’Agostino · 2017
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Disinformation’s spread: Bots, trolls and all of us
Kate Starbird · 2019
Cited alongside, same era.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Yihan Wu, Zhengmian Hu, Hongyang Zhang, and Heng Huang · 2023
Later among the works it cites.
Sheared LLaMA: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen · 2023
Later among the works it cites.
Robust multi-bit natural language watermarking through invariant features
KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak · 2023
Later among the works it cites.
Towards better statistical understanding of watermarking LLMs
Zhongze Cai, Shang Liu, Hanzhao Wang, Huaiyang Zhong, and Xiaocheng Li · 2024
Closest in time.
Pseudorandom error-correcting codes
Miranda Christ and Sam Gunn · 2024
Closest in time.
Undetectable watermarks for language models
Miranda Christ, Sam Gunn, and Or Zamir · 2024
Closest in time.
Scalable watermarking for identifying large language model outputs
Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, Rob McAdam, Johannes Welbl, Vandana Bachani, Alex Kaskasoli, Robert Stanforth, Tatiana Matejovicova, Jamie Hayes, Nidhi Vyas, Majd Al Merey, Jonah Brown-Cohen, Rudy Bunel, Borja Balle, Taylan Cemgil, Zahra Ahmed, Kitty Stacpoole, Ilia Shumailov, Ciprian Baetu, Sven Gowal, Demis Hassabis, and Pushmeet Kohli · 2024
Closest in time.
AI watermarking must be watertight to be effective
Nature editorial · 2024
Closest in time.
GumbelSoft: Diversified language model watermarking via the GumbelMax-trick
Jiayi Fu, Xuandong Zhao, Ruihan Yang, Yuansen Zhang, Jiangjie Chen, and Yanghua Xiao · 2024
Closest in time.
WaterMax: Breaking the LLM watermark detectability-robustness-quality trade-off
Eva Giboulot and Furon Teddy · 2024
Closest in time.
Edit distance robust watermarks for language models
Noah Golowich and Ankur Moitra · 2024
Closest in time.
SemStamp: A semantic watermark with paraphrastic robustness for text generation
Abe Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, and Yulia Tsvetkov · 2024
Closest in time.
Unbiased watermark for large language models
Zhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu, Hongyang Zhang, and Heng Huang · 2024
Closest in time.
On the reliability of watermarks for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein · 2024
Closest in time.
Robust distortion-free watermarks for language models
Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang · 2024
Closest in time.
A statistical framework of watermarks for large language models: Pivot, detection efficiency and optimal rules
Xiang Li, Feng Ruan, Huiyuan Wang, Qi Long, and Weijie J. Su · 2024
Closest in time.
A semantic invariant robust watermark for large language models
Aiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng, and Lijie Wen · 2024
Closest in time.
Adaptive text watermark for large language models
Yepeng Liu and Yuheng Bu · 2024
Closest in time.
Understanding the source of what we see and hear online, May 2024
OpenAI · 2024
Closest in time.
A robust semantics-based watermark for large language model against paraphrasing
Jie Ren, Han Xu, Yiding Liu, Yingqian Cui, Shuaiqiang Wang, Dawei Yin, and Jiliang Tang · 2024
Closest in time.
Debiasing watermarks for large language models via maximal coupling
Yangxinyu Xie, Xiang Li, Tanwi Mallick, Weijie J. Su, and Ruixun Zhang · 2024
Closest in time.
Duwak: Dual watermarks in large language models
Chaoyi Zhu, Jeroen Galjaard, Pin-Yu Chen, and Lydia Y Chen · 2024
Closest in time.