Fetching the paper…
Reading the bibliography…
There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data $C$ that was in their training set.
Calibrating noise to sensitivity in private data analysis
C. Dwork, F. McSherry, K. Nissim, and A. D. Smith · 2006
Earlier work this paper cites.
Why copyright law excludes systems and processes from the scope of its protection
P. Samuelson · 2007
Earlier work this paper cites.
The algorithmic foundations of differential privacy
C. Dwork, A. Roth, et al · 2014
Earlier work this paper cites.
Artificial intelligence and the copyright dilemma
K. Hristov · 2016
Earlier work this paper cites.
Privacy-preserving prediction
C. Dwork and V. Feldman · 2018
Earlier work this paper cites.
The new legal landscape for text mining and machine learning
M. Sag · 2018
Earlier work this paper cites.
Artificial intelligence’s fair use crisis
B. L. Sobel · 2018
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Earlier work this paper cites.
Do we train on test data? purging CIFAR of near-duplicates
B. Barz and J. Denzler · 2020
Earlier work this paper cites.
Copyright infringement in ai-generated artworks
J. Gillotte · 2020
Earlier work this paper cites.
The trade-offs of private prediction
L. van der Maaten and A. Hannun · 2020
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al · 2021
Cited alongside, same era.
Deep learning with label differential privacy
B. Ghazi, N. Golowich, R. Kumar, P. Manurangsi, and C. Zhang · 2021
Cited alongside, same era.
On density estimation with diffusion models
D. P. Kingma, T. Salimans, B. Poole, and J. Ho · 2021
Cited alongside, same era.
Large language models can be strong differentially private learners
X. Li, F. Tramer, P. Liang, and T. Hashimoto · 2021
Cited alongside, same era.
Text and data mining of in-copyright works: is it legal?
P. Samuelson · 2021
Cited alongside, same era.
Differentially private learning needs better features (or much more data)
F. Tramer and D. Boneh · 2021
Cited alongside, same era.
Deduplicating training data makes language models better
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini · 2022
Later among the works it cites.
Training text-to-text transformers with privacy guarantees
N. Ponomareva, J. Bastings, and S. Vassilvitskii · 2022
Later among the works it cites.
Formalizing human ingenuity: A quantitative framework for copyright law’s substantial similarity
S. Scheffler, E. Tromer, and M. Varia · 2022
Later among the works it cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
K. Tirumala, A. H. Markosyan, L. Zettlemoyer, and A. Aghajanyan · 2022
Later among the works it cites.
Second Request for Reconsideration for Refusal to Register A Recent Entrance to Paradise
C. O. R. Board · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quantifying memorization across neural language models
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramèr, and C. Zhang · 2022
Cited alongside, same era.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Cited alongside, same era.
Preventing verbatim memorization in language models gives a false sense of privacy
D. Ippolito, F. Tramèr, M. Nasr, C. Zhang, M. Jagielski, K. Lee, C. A. Choquette-Choo, and N. Carlini · 2022
Cited alongside, same era.
Deduplicating training data mitigates privacy risks in language models
N. Kandpal, E. Wallace, and C. Raffel · 2022
Cited alongside, same era.
Do language models plagiarize?
J. Lee, T. Le, J. Chen, and D. Lee
Cited in the paper.
N. Carlini, J. Hayes, M. Nasr, M. Jagielski, V. Sehwag, F. Tramèr, B. Balle, D. Ippolito, and E. Wallace · 2023
Closest in time.
Can copyright be reduced to privacy?
N. Elkin-Koren, U. Hacohen, R. Livni, and S. Moran · 2023
Closest in time.
Mosaic Large Language Models
Mosaic ML · 2023
Closest in time.
Circular 33: Works not protected by copyright
U.S. Copyright Office · 2023
Closest in time.
pytorch-ddpm : Unofficial PyTorch implementation of Denoising Diffusion Probabilistic Models
Yi-Lun Wu · 2023
Closest in time.