Fetching the paper…
Reading the bibliography…
Training data attribution (TDA) methods offer to trace a model's prediction on any given example back to specific influential training examples.
A value for n-person games
Lloyd S Shapley et al · 1953
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
Arthur E Hoerl and Robert W Kennard · 1970
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R Hampel · 1974
Earlier work this paper cites.
Characterizations of an empirical influence function for detecting influential cases in regression
R Dennis Cook and Sanford Weisberg · 1980
Earlier work this paper cites.
Estimating training data influence by tracking gradient descent
Garima Pruthi, Frederick Liu, Mukund Sundararajan, and Satyen Kale · 2002
Earlier work this paper cites.
Submodular functions and optimization
Satoru Fujishige · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2006
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
The unreasonable effectiveness of data
Alon Halevy, Peter Norvig, and Fernando Pereira · 2009
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Yang Fan, Fei Tian, Tao Qin, Jiang Bian, and Tie-Yan Liu · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei · 2018
Cited alongside, same era.
Screenernet: Learning self-paced curriculum for deep neural networks
Tae-Hoon Kim and Jonghyun Choi · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Influence tuning: Demoting spurious correlations via instance attribution and instance-driven updates
Xiaochuang Han and Yulia Tsvetkov · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Later among the works it cites.
Revisiting methods for finding influential examples
Anders Søgaard et al · 2021
Later among the works it cites.
A survey on curriculum learning
Xin Wang, Yudong Chen, and Wenwu Zhu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Sik Kim, Ian En-Hsu Yen, and Pradeep Ravikumar · 2018
Cited alongside, same era.
On the accuracy of influence functions for measuring group effects
Pang Wei W Koh, Kai-Siang Ang, Hubert Teo, and Percy S Liang · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Explaining black box predictions and unveiling data artifacts through influence functions
Xiaochuang Han, Byron C Wallace, and Yulia Tsvetkov · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi · 2020
Cited alongside, same era.
Fastif: Scalable influence functions for efficient model interpretation and debugging
Han Guo, Nazneen Rajani, Peter Hase, Mohit Bansal, and Caiming Xiong · 2021
Cited alongside, same era.
Xiaochuang Han and Yulia Tsvetkov · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel · 2022
Later among the works it cites.
Combining feature and instance attribution to detect artifacts
Pouya Pezeshkpour, Sarthak Jain, Sameer Singh, and Byron C Wallace · 2022
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Tali Bers, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush · 2022
Later among the works it cites.
Scaling up influence functions
Andrea Schioppa, Polina Zablotskaia, David Vilar, and Artem Sokolov · 2022
Later among the works it cites.
How many and which training points would need to be removed to flip this prediction?
Jinghan Yang, Sarthak Jain, and Byron C. Wallace · 2023
Closest in time.