Fetching the paper…
Reading the bibliography…
Multimodal datasets are a critical component in recent breakthroughs such as Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms.
Semantic redundancies in image-classification datasets: The 10% you don’t need
Vighnesh Birodkar, Hossein Mobahi, and Samy Bengio · 1901
Earlier work this paper cites.
Repair: Removing representation bias by dataset resampling
Yi Li and Nuno Vasconcelos · 1904
Earlier work this paper cites.
On the accuracy of influence functions for measuring group effects
Pang Wei Koh, Kai-Siang Ang, Hubert Teo, and Percy S Liang · 1905
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing · 1905
Earlier work this paper cites.
Selection via proxy: Efficient data selection for deep learning
C Coleman, C Yeh, S Mussmann, B Mirzasoleiman, P Bailis, P Liang, J Leskovec, and M Zaharia · 1906
Earlier work this paper cites.
Coresets for data-efficient training of machine learning models
Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec · 1906
Earlier work this paper cites.
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song · 1907
Earlier work this paper cites.
Kimmo Karkkainen and Jungseock Joo · 1908
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 1910
Earlier work this paper cites.
The visual task adaptation benchmark, 2019
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, André Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Neil Houlsby · 1910
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 1911
Earlier work this paper cites.
Towards a critical race methodology in algorithmic fairness
A. Hanna, Emily L. Denton, Andrew Smart, and Jamila Smith-Loud · 1912
Earlier work this paper cites.
Kaiyu Yang, Klint Qinami, Li Fei-Fei, Jia Deng, and Olga Russakovsky · 1912
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R Hampel · 1974
Earlier work this paper cites.
Detection of influential observation in linear
R Dennis Cook · 1977
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
Yann LeCun · 1998
Earlier work this paper cites.
Two-phase clustering process for outliers detection
Mon-Fong Jiang, Shian-Shyong Tseng, and Chih-Ming Su · 2001
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Fast and three-rious: Speeding up weak supervision with triplet methods
Daniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper, Kayvon Fatahalian, and Christopher Ré · 2002
Earlier work this paper cites.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew Peters, Ashish Sabharwal, and Yejin Choi · 2002
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan · 2002
Earlier work this paper cites.
Findout: Finding outliers in very large datasets
Dantong Yu, Gholamhosein Sheikholeslami, and Aidong Zhang · 2002
Earlier work this paper cites.
Approximating extent measures of points
Pankaj K. Agarwal, Sariel Har-Peled, and Kasturi R. Varadarajan · 2004
Earlier work this paper cites.
The iwildcam 2020 competition dataset, 2020
Sara Beery, Elijah Cole, and Arvi Gjoka · 2004
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2004
Earlier work this paper cites.
On coresets for k-means and k-median clustering
Sariel Har-Peled and Soham Mazumdar · 2004
Earlier work this paper cites.
Tao: A large-scale benchmark for tracking any object
Achal Dave, Tarasha Khurana, Pavel Tokmakov, Cordelia Schmid, and Deva Ramanan · 2005
Earlier work this paper cites.
Explaining black box predictions and unveiling data artifacts through influence functions, 2020
Xiaochuang Han, Byron C Wallace, and Yulia Tsvetkov · 2005
Earlier work this paper cites.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer · 2006
Earlier work this paper cites.
Large image datasets: A pyrrhic win for computer vision?
Vinay Uday Prabhu and Abeba Birhane · 2006
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results, 2007
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Generalized expectation criteria for semi-supervised learning with weakly labeled data
Gideon S Mann and Andrew McCallum · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee · 2011
Earlier work this paper cites.
Scalable training of mixture models via coresets
Dan Feldman, Matthew Faulkner, and Andreas Krause · 2011
Earlier work this paper cites.
Knowledge-based weak supervision for information extraction of overlapping relations
Raphael Hoffmann, Congle Zhang, Xiao Ling, Luke Zettlemoyer, and Daniel S Weld · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg · 2011
Earlier work this paper cites.
Robust statistics for outlier detection
Peter J Rousseeuw and Mia Hubert · 2011
Earlier work this paper cites.
The german traffic sign recognition benchmark: a multi-class classification competition
Johannes Stallkamp, Marc Schlipsing, Jan Salmen, and Christian Igel · 2011
Earlier work this paper cites.
A survey of crowdsourcing systems
Man-Ching Yuen, Irwin King, and Kwong-Sak Leung · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Fastif: Scalable influence functions for efficient model interpretation and debugging, 2020
Han Guo, Nazneen Fatema Rajani, Peter Hase, Mohit Bansal, and Caiming Xiong · 2012
Earlier work this paper cites.
WILDS: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Cats and dogs
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft, 2013
S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Cited alongside, same era.
Coresets for nonparametric estimation - the case of dp-means
Olivier Bachem, Mario Lucic, and Andreas Krause · 2015
Cited alongside, same era.
Microsoft COCO captions: Data collection and evaluation server, 2015
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Submodularity in data subset selection and active learning
Kai Wei, Rishabh Iyer, and Jeff Bilmes · 2015
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Later among the works it cites.
Data-centric ai competition, 2021
Andrew Ng, Dillon Laird, and Lynn He · 2021
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Data programming: Creating large training sets, quickly
Alexander J Ratner, Christopher M De Sa, Sen Wu, Daniel Selsam, and Christopher Ré · 2016
Cited alongside, same era.
YFCC100M: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Cited alongside, same era.
Sun database: Exploring a large collection of scene categories
Jianxiong Xiao, Krista A Ehinger, James Hays, Antonio Torralba, and Aude Oliva · 2016
Cited alongside, same era.
Apache spark: a unified engine for big data processing
Matei Zaharia, Reynold S Xin, Patrick Wendell, Tathagata Das, Michael Armbrust, Ankur Dave, Xiangrui Meng, Josh Rosen, Shivaram Venkataraman, Michael J Franklin, et al · 2016
Cited alongside, same era.
Remote sensing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu · 2017
Cited alongside, same era.
Input sparsity time low-rank approximation via ridge leverage score sampling
Michael B. Cohen, Cameron Musco, and Christopher Musco · 2017
Cited alongside, same era.
Hieu Pham, Zihang Dai, Golnaz Ghiasi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, and Quoc V. Le · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
LAION-400M: Open dataset of clip-filtered 400 million image-text pairs, 2021
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Later among the works it cites.
How much can clip benefit vision-and-language tasks?, 2021
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer · 2021
Later among the works it cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork · 2021
Later among the works it cites.
Shuhei Yokoo · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision, 2021
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Later among the works it cites.
WRENCH: A comprehensive benchmark for weak supervision
Jieyu Zhang, Yue Yu, Yinghao Li, Yujing Wang, Yaming Yang, Mao Yang, and Alexander Ratner · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al · 2022
Later among the works it cites.
WinoGAViL: Gamified association benchmark to challenge vision-and-language models, 2022
Yonatan Bitton, Nitzan Bitton Guetta, Ron Yosef, Yuval Elovici, Mohit Bansal, Gabriel Stanovsky, and Roy Schwartz · 2022
Later among the works it cites.
Coyo-700m: Image-text pair dataset
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim · 2022
Later among the works it cites.
Data distributional properties drive emergent in-context learning in transformers
Stephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang, Aaditya Singh, Pierre H. Richemond, Jay McClelland, and Felix Hill · 2022
Later among the works it cites.
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut · 2022
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning, 2022
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev · 2022
Later among the works it cites.
Understanding dataset difficulty with v-usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta · 2022
Later among the works it cites.
dcbench: a benchmark for data-centric ai systems
Sabri Eyuboglu, Bojan Karlaš, Christopher Ré, Ce Zhang, and James Zou · 2022
Later among the works it cites.
Data determines distributional robustness in contrastive language image pre-training (clip)
Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan, Vaishaal Shankar, Achal Dave, and Ludwig Schmidt · 2022
Later among the works it cites.
Deepcore: A comprehensive library for coreset selection in deep learning, 2022
Chengcheng Guo, Bo Zhao, and Yanbing Bai · 2022
Later among the works it cites.
Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Robots enact malignant stereotypes
Andrew Hundt, William Agnew, Vicky Zeng, Severin Kacianka, and Matthew Gombolay · 2022
Later among the works it cites.
Patching open-vocabulary models by interpolating weights
Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song, Hannaneh Hajishirzi, Simon Kornblith, Ali Farhadi, and Ludwig Schmidt · 2022
Later among the works it cites.
Datamodels: Predicting predictions from training data, 2022
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Later among the works it cites.
Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre-training
Yulong Liu, Guibo Zhu, Bin Zhu, Qi Song, Guojing Ge, Haoran Chen, GuanHui Qiao, Ru Peng, Lingxiang Wu, and Jinqiao Wang · 2022
Later among the works it cites.
Dataperf: Benchmarks for data-centric ai development, 2022
Mark Mazumder, Colby Banbury, Xiaozhe Yao, Bojan Karlaš, William Gaviria Rojas, Sudnya Diamos, Greg Diamos, Lynn He, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Juan Ciro, Lora Aroyo, Bilge Acun, Sabri Eyuboglu, Amirata Ghorbani, Emmett Goodman, Tariq Kane, Christine R. Kirkpatrick, Tzu-Sheng Kuo, Jonas Mueller, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung, Newsha Ardalani, Praveen Paritosh, Ce Zhang, James Zou, Carole-Jean Wu, Cody Coleman, Andrew Ng, Peter Mattson, and Vijay Janapa Reddi · 2022
Later among the works it cites.
Quality not quantity: On the interaction between dataset design and robustness of clip
Thao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh, and Ludwig Schmidt · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision, 2022
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents, 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world
William A Gaviria Rojas, Sudnya Diamos, Keertan Ranjan Kini, David Kanter, Vijay Janapa Reddi, and Cody Coleman · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Extending the wilds benchmark for unsupervised adaptation
Shiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao, Sang Michael Xie, Kendrick Shen, Ananya Kumar, Weihua Hu, Michihiro Yasunaga, Henrik Marklund, Sara Beery, Etienne David, Ian Stavness, Wei Guo, Jure Leskovec, Kate Saenko, Tatsunori Hashimoto, Sergey Levine, Chelsea Finn, and Percy Liang · 2022
Later among the works it cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Later among the works it cites.
Universalizing weak supervision
Changho Shin, Winfred Li, Harit Vishwakarma, Nicholas Roberts, and Frederic Sala · 2022
Later among the works it cites.
CLIP models are few-shot learners: Empirical studies on VQA and visual entailment
Haoyu Song, Li Dong, Weinan Zhang, Ting Liu, and Furu Wei · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos · 2022
Later among the works it cites.
A study of face obfuscation in ImageNet
Kaiyu Yang, Jacqueline H Yau, Li Fei-Fei, Jia Deng, and Olga Russakovsky · 2022
Later among the works it cites.
Filip: Fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2022
Later among the works it cites.
A survey on programmatic weak supervision, 2022
Jieyu Zhang, Cheng-Yu Hsieh, Yue Yu, Chao Zhang, and Alexander Ratner · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, and Ishan Misra · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication, 2023
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Morcos · 2023
Closest in time.
Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger · 2023
Closest in time.
Poisoning web-scale training datasets is practical, 2023
Nicholas Carlini, Matthew Jagielski, Christopher A Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr · 2023
Closest in time.
Simfluence: Modeling the influence of individual training examples by simulating training runs, 2023
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Filtering, distillation, and hard negatives for vision-language pre-training
Filip Radenovic, Abhimanyu Dubey, Abhishek Kadian, Todor Mihaylov, Simon Vandenhende, Yash Patel, Yi Wen, Vignesh Ramanathan, and Dhruv Mahajan · 2023
Closest in time.
Beyond web-scraping: Crowd-sourcing a geodiverse datase, 2023
Vikram V. Ramaswamy, Sing Yu Lin, Dora Zhao, Aaron B. Adcock, Laurens van der Maaten, Deepti Ghadiyaram, and Olga Russakovsky · 2023
Closest in time.
On the de-duplication of laion-2b, 2023
Ryan Webster, Julien Rabin, Loic Simon, and Frederic Jurie · 2023
Closest in time.