Fetching the paper…
Reading the bibliography…
Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems.
Switchboard: telephone speech corpus for research and development
J. Godfrey, E. Holliman, and J. McDaniel · 1992
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge
K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Hidden technical debt in machine learning systems
D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J.-F. Crespo, and D. Dennison · 2015
Earlier work this paper cites.
Torchvision: Pytorch’s computer vision library
T. maintainers and contributors · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Serverless computing: Current trends and open problems
I. Baldini, P. Castro, K. Chang, P. Cheng, S. Fink, V. Ishakian, N. Mitchell, V. Muthusamy, R. Rabbah, A. Slominski, et al · 2017
Earlier work this paper cites.
Dawnbench: An end-to-end deep learning benchmark and competition
C. Coleman, D. Narayanan, D. Kang, T. Zhao, J. Zhang, L. Nardi, P. Bailis, K. Olukotun, C. Ré, and M. Zaharia · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
S. M. Lundberg and S.-I. Lee · 2017
Earlier work this paper cites.
Making neural QA as simple as possible but not simpler
D. Weissenborn, G. Wiese, and L. Seiffe · 2017
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
J. Buolamwini and T. Gebru · 2018
Earlier work this paper cites.
Aibench: towards scalable and comprehensive datacenter ai benchmarking
W. Gao, C. Luo, L. Wang, X. Xiong, J. Chen, T. Hao, Z. Jiang, F. Fan, M. Du, Y. Huang, et al · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
S. Gururangan, S. Swayamdipta, O. Levy, R. Schwartz, S. R. Bowman, and N. A. Smith · 2018
Earlier work this paper cites.
Reproducible, reusable, and robust reinforcement learning, 2018
J. Pineau · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
A. Poliak, J. Naradowsky, A. Haldar, R. Rudinger, and B. Van Durme · 2018
Earlier work this paper cites.
Semantically equivalent adversarial rules for debugging nlp models
M. T. Ribeiro, S. Singh, and C. Guestrin · 2018
Earlier work this paper cites.
Performance impact caused by hidden bias of training data for recognizing textual entailment
M. Tsuchiya · 2018
Earlier work this paper cites.
Tbd: Benchmarking and analyzing deep neural network training
H. Zhu, M. Akrout, B. Zheng, A. Pelegris, A. Phanishayee, B. Schroeder, and G. Pekhimenko · 2018
Earlier work this paper cites.
Don’t take the premise for granted: Mitigating artifacts in natural language inference
Y. Belinkov, A. Poliak, S. M. Shieber, B. Van Durme, and A. M. Rush · 2019
Earlier work this paper cites.
Selection via proxy: Efficient data selection for deep learning
C. Coleman, C. Yeh, S. Mussmann, B. Mirzasoleiman, P. Bailis, P. Liang, J. Leskovec, and M. Zaharia · 2019
Earlier work this paper cites.
Excavating ai: The politics of training sets for machine learning, September 2019
K. Crawford and T. Paglen · 2019
Earlier work this paper cites.
M. Geva, Y. Goldberg, and J. Berant · 2019
Cited alongside, same era.
Crowdsourcing with fairness, diversity and budget constraints | proceedings of the 2019 aaai/acm conference on ai, ethics, and society
N. Goel and B. Faltings · 2019
Cited alongside, same era.
Revealing the dark secrets of bert, 2019
O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky · 2019
Cited alongside, same era.
Human uncertainty makes classification more robust
J. C. Peterson, R. M. Battleday, T. L. Griffiths, and O. Russakovsky · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing nlp
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2019
Cited alongside, same era.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai
N. Sambasivan, S. Kapania, H. Highfill, D. Akrong, P. Paritosh, and L. M. Aroyo · 2021
Later among the works it cites.
Aibench training: Balanced industry-standard ai training benchmarking
F. Tang, W. Gao, J. Zhan, C. Lan, X. Wen, L. Wang, C. Luo, Z. Cao, X. Xiong, Z. Jiang, et al · 2021
Later among the works it cites.
Benchmark and survey of automated machine learning frameworks
M.-A. Zöller and M. F. Huber · 2021
Later among the works it cites.
Data excellence for ai: why should you care?
L. Aroyo, M. Lease, P. Paritosh, and M. Schaekermann · 2022
Closest in time.
Benchmark of filter methods for feature selection in high-dimensional gene expression survival data
A. Bommert, T. Welchowski, M. Schmid, and J. Rahnenführer · 2022
Closest in time.
Dcbench: A benchmark for data-centric ai systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Welty, P. Paritosh, and L. Aroyo · 2019
Cited alongside, same era.
Bringing the people back in: Contesting benchmark machine learning datasets
E. Denton, A. Hanna, R. Amironesei, A. Smart, H. Nicole, and M. K. Scheuerman · 2020
Cited alongside, same era.
The open images dataset v4
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al · 2020
Cited alongside, same era.
Mlperf training benchmark
P. Mattson, C. Cheng, G. Diamos, C. Coleman, P. Micikevicius, D. Patterson, H. Tang, G.-Y. Wei, P. Bailis, V. Bittorf, D. Brooks, D. Chen, D. Dutta, U. Gupta, K. Hazelwood, A. Hock, X. Huang, D. Kang, D. Kanter, N. Kumar, J. Liao, D. Narayanan, T. Oguntebi, G. Pekhimenko, L. Pentecost, V. Janapa Reddi, T. Robie, T. St John, C.-J. Wu, L. Xu, C. Young, and M. Zaharia · 2020
Cited alongside, same era.
A comprehensive benchmark framework for active learning methods in entity matching
V. V. Meduri, L. Popa, P. Sen, and M. Sarwat · 2020
Cited alongside, same era.
Mlperf inference benchmark
V. J. Reddi, C. Cheng, D. Kanter, P. Mattson, G. Schmuelling, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou, R. Chukka, C. Coleman, S. Davis, P. Deng, G. Diamos, J. Duke, D. Fick, J. S. Gardner, I. Hubara, S. Idgunji, T. B. Jablin, J. Jiao, T. S. John, P. Kanwar, D. Lee, J. Liao, A. Lokhmotov, F. Massa, P. Meng, P. Micikevicius, C. Osborne, G. Pekhimenko, A. T. R. Rajan, D. Sequeira, A. Sirasao, F. Sun, H. Tang, M. Thomson, F. Wei, E. Wu, L. Xu, K. Yamada, B. Yu, G. Yuan, A. Zhong, P. Zhang, and Y. Zhou · 2020
Cited alongside, same era.
Adversarial test set for image classification: Lessons learned from cats4ml data challenge
L. Aroyo, P. Paritosh, S. Ibtasam, D. Bansal, K. Rong, and K. Wong · 2021
Cited alongside, same era.
S. Eyuboglu, B. Karlaš, C. Ré, C. Zhang, and J. Zou · 2022
Closest in time.
Domino: Discovering systematic errors with cross-modal embeddings
S. Eyuboglu, M. Varma, K. K. Saab, J.-B. Delbrouck, C. Lee-Messer, J. Dunnmon, J. Zou, and C. Re · 2022
Closest in time.
The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world
W. Gaviria Rojas, S. Diamos, K. Kini, D. Kanter, V. Janapa Reddi, and C. Coleman · 2022
Closest in time.
Data debugging with shapley importance over end-to-end machine learning pipelines
B. Karlaš, D. Dao, M. Interlandi, B. Li, S. Schelter, W. Wu, and C. Zhang · 2022
Closest in time.
Handling and presenting harmful text in nlp research
H. Kirk, A. Birhane, B. Vidgen, and L. Derczynski · 2022
Closest in time.
Red-teaming the stable diffusion safety filter
J. Rando, D. Paleka, D. Lindner, L. Heim, and F. Tramèr · 2022
Closest in time.
Is one annotation enough?-a data-centric image classification benchmark for noisy and ambiguous label estimation
L. Schmarje, V. Grossmann, C. Zelenka, S. Dippel, R. Kiko, M. Oszust, M. Pastell, J. Stracke, A. Valros, N. Volkmann, et al · 2022
Closest in time.
Usb: A unified semi-supervised learning benchmark for classification
Y. Wang, H. Chen, Y. Fan, W. Sun, R. Tao, W. Hou, R. Wang, L. Yang, Z. Zhou, L.-Z. Guo, et al · 2022
Closest in time.
Amazon aws data exchange, 2023
Amazon · 2023
Closest in time.
Bloomberg api, 2023
Bloomberg · 2023
Closest in time.
Databricks data marketplace, 2023
Databricks · 2023
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2023
Closest in time.
Taus data marketplace, BloombergAPI
TAUS · 2023
Closest in time.
Twitter api, 2023
Twitter · 2023
Closest in time.
Data-centric artificial intelligence: A survey
D. Zha, Z. P. Bhat, K.-H. Lai, F. Yang, Z. Jiang, S. Zhong, and X. Hu · 2023
Closest in time.