Fetching the paper…
Reading the bibliography…
The goal of data attribution is to trace model predictions back to training data.
“The Proof and Measurement of Association between Two Things”
Charles Spearman · 1904
Earlier work this paper cites.
“The principle of minimized iterations in the solution of the matrix eigenvalue problem”
Walter Arnoldi · 1951
Earlier work this paper cites.
“Notes on the n-Person Game—II: The Value of an n-Person Game, The RAND Corporation, The RAND Corporation”
LS Shapley · 1951
Earlier work this paper cites.
“Logistic Regression Diagnostics”
Daryl Pregibon · 1981
Earlier work this paper cites.
“Residuals and influence in regression”
R Cook and Sanford Weisberg · 1982
Earlier work this paper cites.
“Extensions of Lipschitz mappings into a Hilbert space”
William Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
“De-noising by soft-thresholding”
David Donoho · 1995
Earlier work this paper cites.
“Okapi at TREC-3”
Stephen Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu and Mike Gatford · 1995
Earlier work this paper cites.
“Random projection, margins, kernels, and feature-selection”
Avrim Blum · 2006
Earlier work this paper cites.
“Random features for large-scale kernel machines”
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
“Learning Multiple Layers of Features from Tiny Images”
Alex Krizhevsky · 2009
Earlier work this paper cites.
“Compressed least-squares regression”
Odalric Maillard and Rémi Munos · 2009
Earlier work this paper cites.
“Robust statistics: the approach based on influence functions”
Frank Hampel, Elvezio Ronchetti, Peter Rousseeuw and Werner Stahel · 2011
Earlier work this paper cites.
“Integrating NLP using linked data”
Sebastian Hellmann, Jens Lehmann, Sören Auer and Martin Brümmer · 2013
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár and C Zitnick · 2014
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2015
Earlier work this paper cites.
“ImageNet Large Scale Visual Recognition Challenge”
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander. Berg and Li Fei-Fei · 2015
Earlier work this paper cites.
“What makes ImageNet good for transfer learning?”
Minyoung Huh, Pulkit Agrawal and Alexei Efros · 2016
Earlier work this paper cites.
“Influence sketching: Finding influential samples in large-scale regressions”
Mike Wojnowicz, Ben Cruz, Xuan Zhao, Brian Wallace, Matt Wolff, Jay Luan and Caleb Crable · 2016
Earlier work this paper cites.
“Second-order stochastic optimization for machine learning in linear time”
Naman Agarwal, Brian Bullins and Elad Hazan · 2017
Earlier work this paper cites.
“Badnets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain”
Tianyu Gu, Brendan Dolan-Gavitt and Siddharth Garg · 2017
Earlier work this paper cites.
“Understanding Black-box Predictions via Influence Functions”
Pang Koh and Percy Liang · 2017
Earlier work this paper cites.
“A unified approach to interpreting model predictions”
Scott Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
“Empirical analysis of the hessian of over-parametrized neural networks”
Levent Sagun, Utku Evci, V Güney, Yann Dauphin and Léon Bottou · 2017
Earlier work this paper cites.
“Random projections for large-scale regression”
Gian-Andrea Thanei, Christina Heinze and Nicolai Meinshausen · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Holographic Feature Representations of Deep Networks.”
Martin Zinkevich, Alex Davies and Dale Schuurmans · 2017
Earlier work this paper cites.
“T-rex: A large scale alignment of natural language with knowledge base triples”
Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest and Elena Simperl · 2018
Earlier work this paper cites.
“Neural Tangent Kernel: Convergence and Generalization in Neural Networks”
Arthur Jacot, Franck Gabriel and Clement Hongler · 2018
Earlier work this paper cites.
“A scalable estimate of the extra-sample prediction error via approximate leave-one-out”
Kamiar Rad and Arian Maleki · 2018
Earlier work this paper cites.
“GLUE: A multi-task benchmark and analysis platform for natural language understanding”
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy and Samuel Bowman · 2018
Earlier work this paper cites.
“Representer Point Selection for Explaining Deep Neural Networks”
Chih-Kuan Yeh, Joon Kim, Ian.. Yen and Pradeep Ravikumar · 2018
Earlier work this paper cites.
“The unreasonable effectiveness of deep features as a perceptual metric”
Richard Zhang, Phillip Isola, Alexei Efros, Eli Shechtman and Oliver Wang · 2018
Earlier work this paper cites.
“Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks”
Sanjeev Arora, Simon. Du, Wei Hu, Zhiyuan Li and Ruosong Wang · 2019
Earlier work this paper cites.
“Second-Order Group Influence Functions for Black-Box Predictions”
Samyadeep Basu, Xuchen You and Soheil Feizi · 2019
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Cited alongside, same era.
“Convergence of Adversarial Training in Overparametrized Networks”
Ruiqi Gao, Tianle Cai, Haochuan Li, Liwei Wang, Cho-Jui Hsieh and Jason Lee · 2019
Cited alongside, same era.
“ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness.”
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix. Wichmann and Wieland Brendel · 2019
Cited alongside, same era.
“Data shapley: Equitable valuation of data for machine learning”
Amirata Ghorbani and James Zou · 2019
Cited alongside, same era.
“Learning transferable visual models from natural language supervision”
Alec Radford, Jong Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin and Jack Clark · 2021
Later among the works it cites.
“Breeds: Benchmarks for subpopulation shift”
Shibani Santurkar, Dimitris Tsipras and Aleksander Madry · 2021
Later among the works it cites.
“Interactive label cleaning with example-based explanations”
Stefano Teso, Andrea Bontempelli, Fausto Giunchiglia and Andrea Passerini · 2021
Later among the works it cites.
“mT5: A massively multilingual pre-trained text-to-text transformer”
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua and Colin Raffel · 2021
Later among the works it cites.
“Tensor Programs IIb: Architectural Universality Of Neural Tangent Kernel Training Dynamics”
Greg Yang and Etai Littwin · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Adversarial Examples Are Not Bugs, They Are Features”
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran and Aleksander Madry · 2019
Cited alongside, same era.
“Towards Efficient Data Valuation Based on the Shapley Value”
Ruoxi Jia, David Dao, Boxin Wang, Frances Hubis, Nick Hynes, Nezihe Gürel, Bo Li, Ce Zhang, Dawn Song and Costas. Spanos · 2019
Cited alongside, same era.
“On the accuracy of influence functions for measuring group effects”
Pang Koh, Kai-Siang Ang, Hubert Teo and Percy Liang · 2019
Cited alongside, same era.
“Interpreting black box predictions using fisher kernels”
Rajiv Khanna, Been Kim, Joydeep Ghosh and Sanmi Koyejo · 2019
Cited alongside, same era.
“Language Models as Knowledge Bases?”
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu and Alexander Miller · 2019
Cited alongside, same era.
“Regularization matters: Generalization and optimization of neural nets vs their induced kernel”
Colin Wei, Jason Lee, Qiang Liu and Tengyu Ma · 2019
Cited alongside, same era.
“Discriminative jackknife: Quantifying uncertainty in deep learning via higher-order influence functions”
Ahmed Alaa and Mihaela Van · 2020
Cited alongside, same era.
“Towards Tracing Factual Knowledge in Language Models Back to the Training Data”
Ekin Akyurek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas and Kelvin Guu · 2022
Later among the works it cites.
“Neural networks as kernel learners: The silent alignment effect”
Alexander Atanasov, Blake Bordelon and Cengiz Pehlevan · 2022
Later among the works it cites.
“Generalization through the lens of leave-one-out error”
Gregor Bachmann, Thomas Hofmann and Aurélien Lucchi · 2022
Later among the works it cites.
“If Influence Functions are the Answer, Then What is the Question?”
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi and Roger Grosse · 2022
Later among the works it cites.
“Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models”
Bernd Bohnet, Vinh Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Soares, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig and Kai Hui · 2022
Later among the works it cites.
“Reproducible scaling laws for contrastive language-image learning”
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt and Jenia Jitsev · 2022
Later among the works it cites.
“Quantifying memorization across neural language models”
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer and Chiyuan Zhang · 2022
Later among the works it cites.
“Careful Data Curation Stabilizes In-context Learning”
Ting-Yun Chang and Robin Jia · 2022
Later among the works it cites.
“Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)”
Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan, Vaishaal Shankar, Achal Dave and Ludwig Schmidt · 2022
Later among the works it cites.
“Training compute-optimal large language models”
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego Casas, Lisa Hendricks, Johannes Welbl and Aidan Clark · 2022
Later among the works it cites.
“Identifying a Training-Set Attack’s Target Using Renormalized Influence Estimation”
Zayd Hammoudeh and Daniel Lowd · 2022
Later among the works it cites.
“Training Data Influence Analysis and Estimation: A Survey”
Zayd Hammoudeh and Daniel Lowd · 2022
Later among the works it cites.
“A framework and benchmark for deep batch active learning for regression”
David Holzmüller, Viktor Zaverkin, Johannes Kästner and Ingo Steinwart · 2022
Later among the works it cites.
“Datamodels: Predicting Predictions from Training Data”
Andrew Ilyas, Sung Park, Logan Engstrom, Guillaume Leclerc and Aleksander Madry · 2022
Later among the works it cites.
“Resolving Training Biases via Influence-based Data Relabeling”
Shuming Kong, Yanyan Shen and Linpeng Huang · 2022
Later among the works it cites.
“Deduplicating Training Data Makes Language Models Better”
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch and Nicholas Carlini · 2022
Later among the works it cites.
“Measuring the Effect of Training Data on Deep Learning Predictions via Randomized Experiments”
Jinkun Lin, Anqi Zhang, Mathias Lecuyer, Jinyang Li, Aurojit Panda and Siddhartha Sen · 2022
Later among the works it cites.
“Behind the Scenes of Gradient Descent: A Trajectory Analysis via Basis Function Decomposition”
Jianhao Ma, Lingjun Guo and Salar Fattahi · 2022
Later among the works it cites.
“A kernel-based view of language model fine-tuning”
Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen and Sanjeev Arora · 2022
Later among the works it cites.
“High-resolution image synthesis with latent diffusion models”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer · 2022
Later among the works it cites.
“Is a caption worth a thousand images? a controlled study for representation learning”
Shibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang and Tatsunori Hashimoto · 2022
Later among the works it cites.
“ModelDiff: A Framework for Comparing Learning Algorithms”
Harshay Shah, Sung Park, Andrew Ilyas and Aleksander Madry · 2022
Later among the works it cites.
“Scaling up influence functions”
Andrea Schioppa, Polina Zablotskaia, David Vilar and Artem Sokolov · 2022
Later among the works it cites.
“Lamda: Language models for dialog applications”
Romal Thoppilan, Daniel De, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker and Yu Du · 2022
Later among the works it cites.
“More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations Generalize”
Alexander Wei, Wei Hu and Jacob Steinhardt · 2022
Later among the works it cites.
“TCT: Convexifying federated learning using bootstrapped neural tangent kernels”
Yaodong Yu, Alexander Wei, Sai Karimireddy, Yi Ma and Michael Jordan · 2022
Later among the works it cites.
“Rethinking Influence Functions of Neural Networks in the Over-Parameterized Regime”
Rui Zhang and Shihua Zhang · 2022
Later among the works it cites.
“The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich Regimes”
Alexander Atanasov, Blake Bordelon, Sabarish Sainathan and Cengiz Pehlevan · 2023
Closest in time.
“Understanding Influence Functions and Datamodels via Harmonic Analysis”
Nikunj Saunshi, Arushi Gupta, Mark Braverman and Sanjeev Arora · 2023
Closest in time.