Fetching the paper…
Reading the bibliography…
In this work, we carry out a data archaeology to infer books that are known to ChatGPT and GPT-4 using a name cloze membership inference query.
Nation, ethnicity, and the geography of British fiction, 1880–1940
Elizabeth F. Evans and Matthew Wilkens. 2018 · 1940
Earlier work this paper cites.
Distinction: A social critique of the judgement of taste
Pierre Bourdieu. 1987 · 1987
Earlier work this paper cites.
Michel Foucault’s Archaeology of Scientific Reason: Science and the History of Reason
Gary Gutting. 1989 · 1989
Earlier work this paper cites.
Cultural capital the problem of literary canon formation
John Guillory. 1993 · 1993
Earlier work this paper cites.
Temporal language models for the disclosure of historical text
FM Jong, Henning Rode, and Djoerd Hiemstra. 2005 · 2005
Earlier work this paper cites.
The computational turn: Thinking about the digital humanities
David M. Berry. 2011 · 2011
Earlier work this paper cites.
The humanities, done digitally
Kathleen Fitzpatrick. 2012 · 2012
Earlier work this paper cites.
How We Think: Transforming Power and Digital Technologies
N. Katherine Hayles. 2012 · 2012
Earlier work this paper cites.
Developing things: Notes toward an epistemology of building in the digital humanities
Stephen Ramsay and Geoffrey Rockwell. 2012 · 2012
Earlier work this paper cites.
The Goldilocks principle: Reading children’s books with explicit memory representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Who did what: A large-scale person-centered cloze dataset
Takeshi Onishi, Hai Wang, Mohit Bansal, Kevin Gimpel, and David McAllester. 2016 · 2016
Earlier work this paper cites.
Estimating the date of first publication in a large-scale digital library
David Bamman, Michelle Carney, Jon Gillick, Cody Hennesy, and Vijitha Sridhar. 2017 · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Earlier work this paper cites.
Why literary time is measured in minutes
Ted Underwood. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Cultural entrenchment of folktales is encoded in language
Folgert Karsdorp and Lauren Fonteyn. 2019 · 2019
Earlier work this paper cites.
Living machines: A study of atypical animacy
Mariona Coll Ardanuy, Federico Nanni, Kaspar Beelen, Kasra Hosseini, Ruth Ahnert, Jon Lawrence, Katherine McDonough, Giorgia Tolfo, Daniel CS Wilson, and Barbara McGillivray. 2020 · 2020
Earlier work this paper cites.
Can GPT-3 pass a writer’s Turing test?
Katherine Elkins and Jon Chun. 2020 · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Cited alongside, same era.
Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation
Aparna Elangovan, Jiayuan He, and Karin Verspoor. 2021 · 2021
Cited alongside, same era.
Preventing verbatim memorization in language models gives a false sense of privacy
Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A Choquette-Choo, and Nicholas Carlini. 2022 · 2022
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022 · 2022
Later among the works it cites.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz. 2022 · 2022
Later among the works it cites.
An empirical analysis of memorization in fine-tuned autoregressive language models
Fatemehsadat Mireshghallah, Archit Uniyal, Tianhao Wang, David Evans, and Taylor Berg-Kirkpatrick. 2022 · 2022
Later among the works it cites.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021 · 2021
Cited alongside, same era.
AI and the Human
Lauren M E Goodlad and Wai Chee Dimock. 2021 · 2021
Cited alongside, same era.
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
Cited alongside, same era.
Gender and representation bias in GPT-3 generated stories
Li Lucy and David Bamman. 2021 · 2021
Cited alongside, same era.
Narrative theory for computational narrative understanding
Andrew Piper, Richard Jean So, and David Bamman. 2021 · 2021
Cited alongside, same era.
The Goodreads “Classics”: A computational study of readers, Amazon, and crowdsourced amateur criticism
Melanie Walsh and Maria Antoniak. 2021 · 2021
Cited alongside, same era.
FanfictionNLP: A text processing pipeline for fanfiction
Michael Yoder, Sopan Khosla, Qinlan Shen, Aakanksha Naik, Huiming Jin, Hariharan Muralidharan, and Carolyn Rosé. 2021 · 2021
Cited alongside, same era.
Peeking Inside the DH Toolbox – Detection and Classification of Software Tools in DH Publications
Nicolas Ruth, Andreas Niekler, and Manuel Burghardt. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Later among the works it cites.
Heroes, villains, and victims, and GPT-3: Automated extraction of character roles without training data
Dominik Stammbach, Maria Antoniak, and Elliott Ash. 2022 · 2022
Later among the works it cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022 · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022 · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023 · 2023
Closest in time.
Extracting training data from diffusion models
Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace. 2023 · 2023
Closest in time.
Poetry Will Not Optimize, or What Is Literature to AI?
M Elam. 2023 · 2023
Closest in time.
Foundation models and fair use
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang. 2023 · 2023
Closest in time.
Data portraits: Recording foundation model training data
Marc Marone and Benjamin Van Durme. 2023 · 2023
Closest in time.
From the archive to the computer: Michel Foucault and the digital humanities
Henning Schmidgen, Bernhard Dotzler, and Benno Stein. 2023 · 2023
Closest in time.
Why open-source generative AI models are an ethical way forward for science
Arthur Spirling. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Closest in time.
Using GPT-4 to measure the passage of time in fiction
Ted Underwood. 2023 · 2023
Closest in time.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2023 · 2023
Closest in time.
Supervised language modeling for temporal resolution of texts
Abhimanu Kumar, Matthew Lease, and Jason Baldridge. 2011 · 2072
Closest in time.