Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are advancing at a remarkable pace, with myriad applications under development.
Inference in an authorship problem: A comparative study of discrimination methods applied to the authorship of the disputed Federalist Papers
F. Mosteller and D. L. Wallace · 1963
Earlier work this paper cites.
Direct and indirect effects
J. Pearl · 2001
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Differential privacy
C. Dwork · 2006
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
C. Dwork, F. McSherry, K. Nissim, and A. Smith · 2006
Earlier work this paper cites.
Membership privacy: A unifying framework for privacy definitions
N. Li, W. Qardaji, D. Su, Y. Wu, and W. Yang · 2013
Earlier work this paper cites.
Cyberspace extortion: North Korea versus the United States
G. Siboni and D. Siman-Tov · 2014
Earlier work this paper cites.
Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers
G. Ateniese, L. V. Mancini, A. Spognardi, A. Villani, D. Vitali, and G. Felici · 2015
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures
M. Fredrikson, S. Jha, and T. Ristenpart · 2015
Earlier work this paper cites.
Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation)
European Union · 2016
Earlier work this paper cites.
Who did what: A large-scale person-centered cloze dataset
T. Onishi, H. Wang, M. Bansal, K. Gimpel, and D. McAllester · 2016
Earlier work this paper cites.
A methodology for formalizing model-inversion attacks
X. Wu, M. Fredrikson, S. Jha, and J. F. Naughton · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
P. W. Koh and P. Liang · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
R. Shokri, M. Stronati, C. Song, and V. Shmatikov · 2017
Earlier work this paper cites.
Privacy amplification by iteration
V. Feldman, I. Mironov, K. Talwar, and A. Thakurta · 2018
Earlier work this paper cites.
Property inference attacks on fully connected neural networks using permutation invariant representations
K. Ganju, Q. Wang, W. Yang, C. A. Gunter, and N. Borisov · 2018
Earlier work this paper cites.
California consumer privacy act (CCPA), 2018
S. of California Department of Justice · 2018
Earlier work this paper cites.
Improving language understanding by Generative Pre-Training, 2018
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
S. Yeom, I. Giacomelli, M. Fredrikson, and S. Jha · 2018
Earlier work this paper cites.
FLAIR: An easy-to-use framework for state-of-the-art NLP
A. Akbik, T. Bergmann, D. Blythe, K. Rasul, S. Schweter, and R. Vollgraf · 2019
Earlier work this paper cites.
M. Arjovsky, L. Bottou, I. Gulrajani, and D. Lopez-Paz · 2019
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
N. Carlini, C. Liu, Ú. Erlingsson, J. Kos, and D. Song · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Handling divergent reference texts when evaluating table-to-text generation
B. Dhingra, M. Faruqui, A. Parikh, M.-W. Chang, D. Das, and W. W. Cohen · 2019
Earlier work this paper cites.
Natural Questions: A benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Earlier work this paper cites.
How bad can it Git? Characterizing secret leakage in public GitHub repositories
M. Meli, M. R. McNiece, and B. Reaves · 2019
Earlier work this paper cites.
Exploiting unintended feature leakage in collaborative learning
L. Melis, C. Song, E. De Cristofaro, and V. Shmatikov · 2019
Earlier work this paper cites.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. S. H. Lewis, A. Bakhtin, Y. Wu, and A. H. Miller · 2019
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
D. Saxton, E. Grefenstette, F. Hill, and P. Kohli · 2019
Earlier work this paper cites.
Cross-domain authorship attribution using pre-trained language models
G. Barlas and E. Stamatatos · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
BertAA: BERT fine-tuning for authorship attribution
M. Fabien, E. Villatoro-Tello, P. Motlicek, and S. Parida · 2020
Earlier work this paper cites.
Does learning require memorization? A short tale about a long tail
V. Feldman · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
V. Feldman and C. Zhang · 2020
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2020
Earlier work this paper cites.
spaCy: Industrial-strength natural language processing in Python, 2020
M. Honnibal, I. Montani, S. Van Landeghem, and A. Boyd · 2020
Earlier work this paper cites.
How can we know what language models know?
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis · 2020
Earlier work this paper cites.
Fair learning
M. A. Lemley and B. Casey · 2020
Earlier work this paper cites.
Dataset inference: Ownership resolution in machine learning
P. Maini, M. Yaghini, and N. Papernot · 2020
Earlier work this paper cites.
E-BERT: Efficient-yet-effective entity embeddings for BERT
N. Pörner, U. Waltinger, and H. Schütze · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
On exposure bias, hallucination and domain shift in neural machine translation
C. Wang and R. Sennrich · 2020
Earlier work this paper cites.
Revisiting challenges in data-to-text generation with fact grounding
H. Wang · 2020
Earlier work this paper cites.
Addressing “documentation debt” in machine learning: A retrospective datasheet for BookCorpus
J. Bandy and N. Vincent · 2021
Earlier work this paper cites.
GPT-Neo: Large scale autoregressive language modeling with Mesh-Tensorflow, 2021
S. Black, L. Gao, P. Wang, C. Leahy, and S. Biderman · 2021
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the Colossal Clean Crawled Corpus
J. Dodge, A. Marasović, G. Ilharco, D. Groeneveld, M. Mitchell, and M. Gardner · 2021
Earlier work this paper cites.
Neural Path Hunter: Reducing hallucination in dialogue systems via path grounding
N. Dziri, A. Madotto, O. Zaïane, and A. J. Bose · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Revisiting membership inference under realistic assumptions
B. Jayaraman, L. Wang, K. Knipmeyer, Q. Gu, and D. Evans · 2021
Earlier work this paper cites.
Does BERT pretrained on clinical notes reveal sensitive data?
E. Lehman, S. Jain, K. Pichotta, Y. Goldberg, and B. C. Wallace · 2021
Earlier work this paper cites.
Learning with user-level privacy
D. Levy, Z. Sun, K. Amin, S. Kale, A. Kulesza, M. Mohri, and A. T. Suresh · 2021
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
S. C. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
What’s in the box? An analysis of undesirable content in the Common Crawl corpus
A. S. Luccioni and J. D. Viviano · 2021
Earlier work this paper cites.
Style pooling: Automatic text style obfuscation for improved classification fairness
F. Mireshghallah and T. Berg-Kirkpatrick · 2021
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Earlier work this paper cites.
The curious case of hallucinations in neural machine translation
V. Raunak, A. Menezes, and M. Junczys-Dowmunt · 2021
Earlier work this paper cites.
Process for adapting language models to society (PALMS) with values-targeted datasets
I. Solaiman and C. Dennison · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 billion parameter autoregressive language model, 2021
B. Wang and A. Komatsuzaki · 2021
Earlier work this paper cites.
Improving robustness to model inversion attacks via mutual information regularization
T. Wang, Y. Zhang, and R. Jia · 2021
Earlier work this paper cites.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Earlier work this paper cites.
Counterfactual memorization in neural language models
C. Zhang, D. Ippolito, K. Lee, M. Jagielski, F. Tramèr, and N. Carlini · 2021
Cited alongside, same era.
Leakage of dataset properties in multi-party machine learning
W. Zhang, S. Tople, and O. Ohrimenko · 2021
Cited alongside, same era.
A Review on Language Models as Knowledge Bases
B. AlKhamissi, M. Li, A. Celikyilmaz, M. Diab, and M. Ghazvininejad · 2022
Cited alongside, same era.
Language models as agent models
J. Andreas · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung · 2023
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Closest in time.
Flows: Building blocks of reasoning and collaborating AI
M. Josifoski, L. Klein, M. Peyrard, Y. Li, S. Geng, J. P. Schnitzler, Y. Yao, J. Wei, D. Paul, and R. West · 2023
Closest in time.
Language model decoding as likelihood–utility alignment
M. Josifoski, M. Peyrard, F. Rajic, J. Wei, D. Paul, V. Hartmann, B. Patra, V. Chaudhary, E. Kıcıman, B. Faltings, and R. West · 2023
Closest in time.
User inference attacks on large language models
N. Kandpal, K. Pillutla, A. Oprea, P. Kairouz, C. A. Choquette-Choo, and Z. Xu · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Constitutional AI: Harmlessness from AI feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
Overview of PAN 2022: Authorship verification, profiling irony and stereotype spreaders, and style change detection
J. Bevendorff, B. Chulvi, E. Fersini, A. Heini, M. Kestemont, K. Kredens, M. Mayerl, R. Ortega-Bueno, P. Pęzik, M. Potthast, et al · 2022
Cited alongside, same era.
Some misconceptions about software in the copyright literature
J. Bloch and P. Samuelson · 2022
Cited alongside, same era.
What does it mean for a language model to preserve privacy?
H. Brown, K. Lee, F. Mireshghallah, R. Shokri, and F. Tramèr · 2022
Cited alongside, same era.
Quantifying memorization across neural language models
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang · 2022
Cited alongside, same era.
On the impossible safety of large AI models
E.-M. El-Mhamdi, S. Farhadkhani, R. Guerraoui, N. Gupta, L. N. Hoang, R. Pinot, and J. Stephan · 2022
Cited alongside, same era.
Preventing verbatim memorization in language models gives a false sense of privacy
D. Ippolito, F. Tramèr, M. Nasr, C. Zhang, M. Jagielski, K. Lee, C. A. Choquette-Choo, and N. Carlini · 2022
Cited alongside, same era.
Closest in time.
ProPILE: Probing privacy leakage in large language models
S. Kim, S. Yun, H. Lee, M. Gubri, S. Yoon, and S. J. Oh · 2023
Closest in time.
The Stack: 3 TB of permissively licensed source code
D. Kocetkov, R. Li, L. B. allal, J. LI, C. Mou, Y. Jernite, M. Mitchell, C. M. Ferrandis, S. Hughes, T. Wolf, D. Bahdanau, L. V. Werra, and H. de Vries · 2023
Closest in time.
Pretraining language models with human preferences
T. Korbak, K. Shi, A. Chen, R. Bhalerao, C. L. Buckley, J. Phang, S. R. Bowman, and E. Perez · 2023
Closest in time.
GitHub Copilot AI is leaking functional API keys
A. Kulkarni · 2023
Closest in time.
Do language models plagiarize?
J. Lee, T. Le, J. Chen, and D. Lee · 2023
Closest in time.
Multi-step jailbreaking privacy attacks on ChatGPT
H. Li, D. Guo, W. Fan, M. Xu, and Y. Song · 2023
Closest in time.
Circuit breaking: Removing model behaviors with targeted ablation
M. Li, X. Davies, and M. Nadeau · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Y. Li, Y. Li, and A. Risteski · 2023
Closest in time.
The Flan Collection: Designing data and methods for effective instruction tuning
S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, et al · 2023
Closest in time.
Bounding the capabilities of large language models in open text generation with prompt constraints
A. Lu, H. Zhang, Y. Zhang, X. Wang, and D. Yang · 2023
Closest in time.
Analyzing leakage of personally identifiable information in language models
N. Lukas, A. Salem, R. Sim, S. Tople, L. Wutschitz, and S. Zanella-Béguelin · 2023
Closest in time.
Can neural network memorization be localized?
P. Maini, M. C. Mozer, H. Sedghi, Z. C. Lipton, J. Z. Kolter, and C. Zhang · 2023
Closest in time.
Developing methods for identifying and removing copyrighted content from generative AI models
K. S. I. Mantri and N. Sasikumatm · 2023
Closest in time.
Membership inference attacks against language models via neighbourhood comparison
J. Mattern, F. Mireshghallah, Z. Jin, B. Schoelkopf, M. Sachan, and T. Berg-Kirkpatrick · 2023
Closest in time.
Whoops, Samsung workers accidentally leaked trade secrets via ChatGPT
C. Mauran · 2023
Closest in time.
R. T. McCoy, S. Yao, D. Friedman, M. Hardy, and T. L. Griffiths · 2023
Closest in time.
Inverse scaling: When bigger isn’t better
I. R. McKenzie, A. Lyzhov, M. Pieler, A. Parrish, A. Mueller, A. Prabhu, E. McLean, A. Kirtland, A. Ross, A. Liu, et al · 2023
Closest in time.
Mass-editing memory in a transformer
K. Meng, A. S. Sharma, A. Andonian, Y. Belinkov, and D. Bau · 2023
Closest in time.
SILO language models: Isolating legal risk in a nonparametric datastore
S. Min, S. Gururangan, E. Wallace, H. Hajishirzi, N. A. Smith, and L. Zettlemoyer · 2023
Closest in time.
Auditing large language models: a three-layered approach
J. Mökander, J. Schuett, H. R. Kirk, and L. Floridi · 2023
Closest in time.
Use of LLMs for illicit purposes: Threats, prevention measures, and vulnerabilities
M. Mozes, X. He, B. Kleinberg, and L. D. Griffin · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Liberum, J. Smith, and J. Steinhardt · 2023
Closest in time.
The Waluigi effect
C. Nardo · 2023
Closest in time.
CodexLeaks: Privacy leaks from code generation language models in GitHub Copilot
L. Niu, S. Mirza, Z. Maradni, and C. Pöpper · 2023
Closest in time.
OpenAI · 2023
Closest in time.
GPTBot, 2023
OpenAI · 2023
Closest in time.
Controlling the extraction of memorized data from large language models via prompt-tuning
M. Ozdayi, C. Peris, J. FitzGerald, C. Dupuy, J. Majmudar, H. Khan, R. Parikh, and R. Gupta · 2023
Closest in time.
Unifying large language models and knowledge graphs: A roadmap
S. Pan, L. Luo, Y. Wang, C. Chen, J. Wang, and X. Wu · 2023
Closest in time.
TRAK: Attributing model behavior at scale
S. M. Park, K. Georgiev, A. Ilyas, G. Leclerc, and A. Madry · 2023
Closest in time.
Can sensitive information be deleted from LLMs? Objectives for defending against extraction attacks, 2023
V. Patil, P. Hase, and M. Bansal · 2023
Closest in time.
Scammers use AI to enhance their family emergency schemes
A. Puig · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Closest in time.
Beyond fair use: Legal risk evaluation for training LLMs on copyrighted text
N. Rahman and E. Santacana · 2023
Closest in time.
In-context retrieval-augmented language models
O. Ram, Y. Levine, I. Dalmedigos, D. Muhlgay, A. Shashua, K. Leyton-Brown, and Y. Shoham · 2023
Closest in time.
Measuring attribution in natural language generation models
H. Rashkin, V. Nikolaev, M. Lamm, L. Aroyo, M. Collins, D. Das, S. Petrov, G. S. Tomar, I. Turc, and D. Reitter · 2023
Closest in time.
I’m afraid I can’t do that: Predicting prompt refusal in black-box generative language models
M. Reuter and W. Schulze · 2023
Closest in time.
SoK: Let the privacy games begin! A unified treatment of data inference privacy in machine learning
A. Salem, G. Cherubin, D. Evans, B. Köpf, A. Paverd, A. Suri, S. Tople, and S. Zanella-Béguelin · 2023
Closest in time.
How your data is used to improve model performance
M. Schade · 2023
Closest in time.
REPLUG: Retrieval-augmented black-box language models
W. Shi, S. Min, M. Yasunaga, M. Seo, R. James, M. Lewis, L. Zettlemoyer, and W.-t. Yih · 2023
Closest in time.
Identifying and mitigating privacy risks stemming from language models: A survey
V. Smith, A. S. Shamsabadi, C. Ashurst, and A. Weller · 2023
Closest in time.
Announcing Microsoft 365 Copilot general availability and Microsoft 365 Chat
J. Spataro · 2023
Closest in time.
Beyond memorization: Violating privacy via inference with large language models
R. Staab, M. Vero, M. Balunović, and M. Vechev · 2023
Closest in time.
Understanding arithmetic reasoning in language models using causal mediation analysis
A. Stolfo, Y. Belinkov, and M. Sachan · 2023
Closest in time.
Dissecting distribution inference
A. Suri, Y. Lu, Y. Chen, and D. Evans · 2023
Closest in time.
On the robustness of dataset inference
S. Szyller, R. Zhang, J. Liu, and N. Asokan · 2023
Closest in time.
Newspapers want payment for articles used to power ChatGPT
N. Tiku · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
From theories on styles to their transfer in text: Bridging the gap with a hierarchical survey
E. Troiano, A. Velutharambath, and R. Klinger · 2023
Closest in time.
Provable copyright protection for generative models
N. Vyas, S. Kakade, and B. Barak · 2023
Closest in time.
DecodingTrust: A comprehensive assessment of trustworthiness in GPT models
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, et al · 2023
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2023
Closest in time.
KGA: A general machine unlearning framework based on knowledge gap alignment
L. Wang, T. Chen, W. Yuan, X. Zeng, K.-F. Wong, and H. Yin · 2023
Closest in time.
Aligning large language models with human: A survey
Y. Wang, W. Zhong, L. Li, F. Mi, X. Zeng, W. Huang, L. Shang, X. Jiang, and Q. Liu · 2023
Closest in time.
The New York Times prohibits using its content to train AI models, 2023
J. Weatherbed · 2023
Closest in time.
“According to…” prompting language models improves quoting from pre-training data
O. Weller, M. Marone, N. Weir, D. Lawrie, D. Khashabi, and B. Van Durme · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Y. Wolf, N. Wies, Y. Levine, and A. Shashua · 2023
Closest in time.
CodeIPPrompt: Intellectual property infringement assessment of code language models
Z. Yu, Y. Wu, N. Zhang, C. Wang, Y. Vorobeychik, and C. Xiao · 2023
Closest in time.
The wisdom of hindsight makes language models better instruction followers
T. Zhang, F. Liu, J. Wong, P. Abbeel, and J. E. Gonzalez · 2023
Closest in time.