Fetching the paper…
Reading the bibliography…
Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs.
A basis for analyzing test-retest reliability
Guttman, L · 1945
Earlier work this paper cites.
Convergent and discriminant validation by the multitrait-multimethod matrix
Campbell, D. T. and Fiske, D. W · 1959
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Cohen, J · 1960
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Fleiss, J. L · 1971
Earlier work this paper cites.
Sun and skin
Fitzpatrick, T · 1975
Earlier work this paper cites.
An ‘other-race effect’for categorizing faces by sex
O’Toole, A. J., Peterson, J., and Deffenbacher, K. A · 1996
Earlier work this paper cites.
Masc: The manually annotated sub-corpus of american english
Ide, N., Baker, C., Fellbaum, C., Fillmore, C., and Passonneau, R · 2008
Earlier work this paper cites.
How to analyze political attention with minimal assumptions and costs
Quinn, K. M., Monroe, B. L., Colaresi, M., Crespin, M. H., and Radev, D. R · 2010
Earlier work this paper cites.
Research methods in education
Check, J. and Schutt, R. K · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A. and Efros, A. A · 2011
Earlier work this paper cites.
Social science research: Principles, methods, and practices
Bhattacherjee, A · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Silberman, N., Hoiem, D., Kohli, P., and Fergus, R · 2012
Earlier work this paper cites.
Detecting hate speech on the world wide web
Warner, W. and Hirschberg, J · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
Song, S., Lichtenberg, S. P., and Xiao, J · 2015
Earlier work this paper cites.
Celis, L. E., Deshpande, A., Kathuria, T., and Vishnoi, N. K · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Broad twitter corpus: A diverse named entity recognition resource
Derczynski, L., Bontcheva, K., and Roberts, I · 2016
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
Lebret, R., Grangier, D., and Auli, M · 2016
Earlier work this paper cites.
A large contextual dataset for classification, detection and counting of cars with deep learning
Mundhenk, T. N., Konjevod, G., Sakla, W. A., and Boakye, K · 2016
Earlier work this paper cites.
Abusive language detection in online user content
Nobata, C., Tetreault, J., Thomas, A., Mehdad, Y., and Chang, Y · 2016
Earlier work this paper cites.
The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes
Ros, G., Sellart, L., Materzynska, J., Vazquez, D., and Lopez, A. M · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Sigurdsson, G. A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R · 2017
Earlier work this paper cites.
Semi-supervised sequence tagging with bidirectional language models
Peters, M. E., Ammar, W., Bhagavatula, C., and Power, R · 2017
Earlier work this paper cites.
A multi-view stereo benchmark with high-resolution images and multi-camera videos
Schops, T., Schonberger, J. L., Galliani, S., Sattler, T., Schindler, K., Pollefeys, M., and Geiger, A · 2017
Earlier work this paper cites.
Shankar, S., Halpern, Y., Breck, E., Atwood, J., Wilson, J., and Sculley, D · 2017
Earlier work this paper cites.
Understanding abuse: A typology of abusive language detection subtasks
Waseem, Z., Davidson, T., Warmsley, D., and Weber, I · 2017
Earlier work this paper cites.
Do artifacts have politics?
Winner, L · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A · 2017
Earlier work this paper cites.
Object recognition with and without objects
Zhu, Z., Xie, L., and Yuille, A · 2017
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A · 2018
Earlier work this paper cites.
Summarizing opinions: Aspect extraction meets sentiment prediction and they are both weakly supervised
Angelidis, S. and Lapata, M · 2018
Earlier work this paper cites.
Measurement theory and applications for the social sciences
Bandalos, D. L · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Bender, E. M. and Friedman, B · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Buolamwini, J. and Gebru, T · 2018
Earlier work this paper cites.
Learning to act properly: Predicting and explaining affordances from images
Chuang, C.-Y., Li, J., Torralba, A., and Fidler, S · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Damen, D., Doughty, H., Farinella, G. M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Dixon, L., Li, J., Sorensen, J., Thain, N., and Vasserman, L · 2018
Earlier work this paper cites.
Inferring shared attention in social scene videos
Fan, L., Chen, Y., Wei, P., Wang, W., and Zhu, S.-C · 2018
Earlier work this paper cites.
Gender recognition or gender reductionism? the social implications of embedded gender recognition systems
Hamidi, F., Scheuerman, M. K., and Branham, S. M · 2018
Earlier work this paper cites.
Composition loss for counting, density map estimation and localization in dense crowds
Idrees, H., Tayyab, M., Athrey, K., Zhang, D., Al-Maadeed, S., Rajpoot, N., and Shah, M · 2018
Earlier work this paper cites.
Classification of moral foundations in microblog political discourse
Johnson, K. and Goldwasser, D · 2018
Earlier work this paper cites.
Trackingnet: A large-scale dataset and benchmark for object tracking in the wild
Müller, M., Bibi, A., Giancola, S., Alsubaihi, S., and Ghanem, B · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rudinger, R., Naradowsky, J., Leonard, B., and Van Durme, B · 2018
Earlier work this paper cites.
Tackling the story ending biases in the story cloze test
Sharma, R., Allen, J., Bakhshandeh, O., and Mostafazadeh, N · 2018
Earlier work this paper cites.
Crrn: Multi-scale guided concurrent reflection removal network
Wan, R., Shi, B., Duan, L.-Y., Tan, A.-H., and Kot, A. C · 2018
Earlier work this paper cites.
Mind the gap: A balanced corpus of gendered ambiguous pronouns
Webster, K., Recasens, M., Axelrod, V., and Baldridge, J · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.-W · 2018
Earlier work this paper cites.
Sketchyscene: Richly-annotated scene sketches
Zou, C., Yu, Q., Du, R., Mo, H., Song, Y.-Z., Xiang, T., Gao, C., Chen, B., and Zhang, H · 2018
Earlier work this paper cites.
Big bird: A large, fine-grained, bigram relatedness dataset for examining semantic composition
Asaadi, S., Mohammad, S., and Kiritchenko, S · 2019
Earlier work this paper cites.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B · 2019
Earlier work this paper cites.
Taskmaster-1: Toward a realistic and diverse dialog dataset
Byrne, B., Krishnamoorthi, K., Sankar, C., Neelakantan, A., Goodrich, B., Duckworth, D., Yavuz, S., Dubey, A., Kim, K.-Y., and Cedilnik, A · 2019
Earlier work this paper cites.
Racial bias in hate speech and abusive language detection datasets
Davidson, T., Bhattacharya, D., and Weber, I · 2019
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
De-Arteaga, M., Romanov, A., Wallach, H., Chayes, J., Borgs, C., Chouldechova, A., Geyik, S., Kenthapadi, K., and Kalai, A. T · 2019
Earlier work this paper cites.
The role of pragmatic and discourse context in determining argument impact
Durmus, E., Ladhak, F., and Cardie, C · 2019
Earlier work this paper cites.
Simple dynamic word embeddings for mapping perceptions in the public sphere
Gillani, N. and Levy, R · 2019
Earlier work this paper cites.
People + AI Guidebook
Google PAIR · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Abstractive summarization of reddit posts with multi-level memory networks
Kim, B., Kim, H., and Kim, G · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Driv&act: A multi-modal dataset for fine-grained driver behavior recognition in autonomous vehicles
Martin, M., Roitberg, A., Haurilet, M., Horne, M., ReiB, S., Voit, M., and Stiefelhagen, R · 2019
Earlier work this paper cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
Ni, J., Li, J., and McAuley, J · 2019
Earlier work this paper cites.
Multi-domain goal-oriented dialogues (multidogo): Strategies toward curating and annotating large scale dialogue data
Peskov, D., Clarke, N., Krone, J., Fodor, B., Zhang, Y., Youssef, A., and Diab, M · 2019
Earlier work this paper cites.
Human uncertainty makes classification more robust
Peterson, J. C., Battleday, R. M., Griffiths, T. L., and Russakovsky, O · 2019
Earlier work this paper cites.
How computers see gender: An evaluation of gender classification in commercial facial analysis services
Scheuerman, M. K., Paul, J. M., and Brubaker, J. R · 2019
Earlier work this paper cites.
Analysis of automatic annotation suggestions for hard discourse-level tasks in expert domains
Schulz, C., Meyer, C. M., Kiesewetter, J., Sailer, M., Bauer, E., Fischer, M. R., Fischer, F., and Gurevych, I · 2019
Earlier work this paper cites.
Pushing the frontiers of unconstrained crowd counting: New dataset and benchmark method
Sindagi, V. A., Yasarla, R., and Patel, V. M · 2019
Earlier work this paper cites.
Proactive human-machine conversation with explicit conversation goal
Wu, W., Guo, Z., Zhou, X., Wu, H., Zhang, X., Lian, R., and Wang, H · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Freihand: A dataset for markerless capture of hand pose and shape from single rgb images
Zimmermann, C., Ceylan, D., Yang, J., Russell, B., Argus, M. J., and Brox, T · 2019
Cited alongside, same era.
The practice of social research
Babbie, E. R · 2020
Cited alongside, same era.
Cmu-moseas: A multimodal language dataset for spanish, portuguese, german and french
Bagher Zadeh, A., Cao, Y., Hessner, S., Liang, P. P., Poria, S., and Morency, L.-P · 2020
Cited alongside, same era.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H · 2020
Cited alongside, same era.
Toward gender-inclusive coreference resolution
Cao, Y. T. and Daumé III, H · 2020
Solid: A large-scale semi-supervised dataset for offensive language identification
Rosenthal, S., Atanasova, P., Karadzhov, G., Zampieri, M., and Nakov, P · 2021
Later among the works it cites.
Visual semantic role labeling for video understanding
Sadhu, A., Gupta, T., Yatskar, M., Nevatia, R., and Kembhavi, A · 2021
Later among the works it cites.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai
Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., and Aroyo, L. M · 2021
Later among the works it cites.
Do datasets have politics? disciplinary values in computer vision dataset development
Scheuerman, M. K., Hanna, A., and Denton, E · 2021
Later among the works it cites.
Beyond fair pay: Ethical implications of NLP crowdsourcing
Shmueli, B., Fell, J., Ray, S., and Ku, L.-W · 2021
Later among the works it cites.
Image representations learned with unsupervised pre-training contain human-like biases
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tao: A large-scale benchmark for tracking any object
Dave, A., Khurana, T., Tokmakov, P., Schmid, C., and Ramanan, D · 2020
Cited alongside, same era.
Bringing the people back in: Contesting benchmark machine learning datasets
Denton, E., Hanna, A., Amironesei, R., Smart, A., Nicole, H., and Scheuerman, M. K · 2020
Cited alongside, same era.
The mapillary traffic sign dataset for detection and classification on a global scale
Ertler, C., Mislej, J., Ollmann, T., Porzi, L., Neuhold, G., and Kuang, Y · 2020
Cited alongside, same era.
Taking a deeper look at co-salient object detection
Fan, D.-P., Lin, Z., Ji, G.-P., Zhang, D., Fu, H., and Cheng, M.-M · 2020
Cited alongside, same era.
From arabic sentiment analysis to sarcasm detection: The arsarcasm dataset
Farha, I. A. and Magdy, W · 2020
Cited alongside, same era.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Cited alongside, same era.
Steed, R. and Caliskan, A · 2021
Later among the works it cites.
Adding chit-chat to enhance task-oriented dialogues
Sun, K., Moon, S., Crook, P., Roller, S., Silvert, B., Liu, B., Wang, Z., Liu, H., Cho, E., and Cardie, C · 2021
Later among the works it cites.
Nutrition5k: Towards automatic nutritional understanding of generic food
Thames, Q., Karpur, A., Norris, W., Xia, F., Panait, L., Weyand, T., and Sim, J · 2021
Later among the works it cites.
Cropharvest: A global dataset for crop-type classification
Tseng, G., Zvonkov, I., Nakalembe, C. L., and Kerner, H · 2021
Later among the works it cites.
There is more than meets the eye: Self-supervised multi-object detection and tracking with sound by distilling multimodal knowledge
Valverde, F. R., Valeria Hurtado, J., and Valada, A · 2021
Later among the works it cites.
Benchmarking representation learning for natural world image collections
Van Horn, G., Cole, E., Beery, S., Wilber, K., Belongie, S., and MacAodha, O · 2021
Later among the works it cites.
Quantifying social organization and political polarization in online platforms
Waller, I. and Anderson, A · 2021
Later among the works it cites.
Implicitly abusive comparisons – a new dataset and linguistic analysis
Wiegand, M., Geulig, M., and Ruppenhofer, J · 2021
Later among the works it cites.
Fake it till you make it: face analysis in the wild using synthetic data alone
Wood, E., Baltrušaitis, T., Hewitt, C., Dziadzio, S., Cashman, T. J., and Shotton, J · 2021
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C · 2021
Later among the works it cites.
Ap-10k: A benchmark for animal pose estimation in the wild
Yu, H., Xu, Y., Zhang, J., Zhao, W., Guan, Z., and Tao, D · 2021
Later among the works it cites.
Synthbio: A case study in faster curation of text datasets
Yuan, A., Ippolito, D., Nikolaev, V., Callison-Burch, C., Coenen, A., and Gehrmann, S · 2021
Later among the works it cites.
Mr. tydi: A multi-lingual benchmark for dense retrieval
Zhang, X., Ma, X., Shi, P., and Lin, J · 2021
Later among the works it cites.
Understanding and evaluating racial biases in image captioning
Zhao, D., Wang, A., and Russakovsky, O · 2021
Later among the works it cites.
Wikibias: Detecting multi-span subjective biases in language
Zhong, Y., Yang, J., Xu, W., and Yang, D · 2021
Later among the works it cites.
URL https://datatopics.worldbank.org/world-development-indicators/the-world-by-income-and-region.html
The world by income and region, 2022 · 2022
Later among the works it cites.
The values encoded in machine learning research
Birhane, A., Kalluri, P., Card, D., Agnew, W., Dotan, R., and Bao, M · 2022
Later among the works it cites.
Responsible language technologies: Foreseeing and mitigating harms
Blodgett, S. L., Liao, Q. V., Olteanu, A., Mihalcea, R., Muller, M., Scheuerman, M. K., Tan, C., and Yang, Q · 2022
Later among the works it cites.
Fiber: Fill-in-the-blanks as a challenging video understanding evaluation framework
Castro, S., Wang, R., Huang, P., Stewart, I., Ignat, O., Liu, N., Stroud, J., and Mihalcea, R · 2022
Later among the works it cites.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Davani, A. M., Díaz, M., and Prabhakaran, V · 2022
Later among the works it cites.
Leveraging wikipedia article evolution for promotional tone detection
De Kock, C. and Vlachos, A · 2022
Later among the works it cites.
Ithaca365: Dataset and driving perception under repeated and challenging weather conditions
Diaz, C. A., Xia, Y., You, Y., Nino, J., Chen, J., Monica, J., Chen, X., Luo, K. Z., Wang, Y., Emond, M., et al · 2022
Later among the works it cites.
Jury learning: Integrating dissenting voices into machine learning models
Gordon, M. L., Lam, M. S., Park, J. S., Patel, K., Hancock, J., Hashimoto, T., and Bernstein, M. S · 2022
Later among the works it cites.
Lot: A story-centric benchmark for evaluating chinese long text understanding and generation
Guan, J., Feng, Z., Chen, Y., He, R., Mao, X., Fan, C., and Huang, M · 2022
Later among the works it cites.
Understanding machine learning practitioners’ data documentation perceptions, needs, challenges, and desiderata
Heger, A. K., Marquis, L. B., Vorvoreanu, M., Wallach, H., and Wortman Vaughan, J · 2022
Later among the works it cites.
NLP’s clever hans moment has arrived, Jan 2022
Heinzerling, B · 2022
Later among the works it cites.
Does recommend-revise produce reliable annotations? an analysis on missing instances in docred
Huang, Q., Hao, S., Ye, Y., Zhu, S., Feng, Y., and Zhao, D · 2022
Later among the works it cites.
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Liu, Y., Liu, Y., Jiang, C., Lyu, K., Wan, W., Shen, H., Liang, B., Fu, Z., Wang, H., and Yi, L · 2022
Later among the works it cites.
Dad-3dheads: A large-scale dense, accurate and diverse dataset for 3d head alignment from a single image
Martyniuk, T., Kupyn, O., Kurlyak, Y., Krashenyi, I., Matas, J., and Sharmanska, V · 2022
Later among the works it cites.
Mitchell, M., Luccioni, A. S., Lambert, N., Gerchick, M., McMillan-Major, A., Ozoani, E., Rajani, N., Thrush, T., Jernite, Y., and Kiela, D · 2022
Later among the works it cites.
It is okay to not be okay: Overcoming emotional bias in affective image captioning by contrastive data collection
Mohamed, Y., Khan, F. F., Haydarov, K., and Elhoseiny, M · 2022
Later among the works it cites.
Multilingual event linking to wikidata
Pratapa, A., Gupta, R., and Mitamura, T · 2022
Later among the works it cites.
Data cards: Purposeful and transparent dataset documentation for responsible ai
Pushkarna, M., Zaldivar, A., and Kjartansson, O · 2022
Later among the works it cites.
The dollar street dataset: Images representing the geographic and socioeconomic diversity of the world
Rojas, W. A. G., Diamos, S., Kini, K. R., Kanter, D., Reddi, V. J., and Coleman, C · 2022
Later among the works it cites.
Assessing annotator identity sensitivity via item response theory: A case study in a hate speech corpus
Sachdeva, P. S., Barreto, R., von Vacano, C., and Kennedy, C. J · 2022
Later among the works it cites.
Vila: Improving structured content extraction from scientific pdfs using visual layout groups
Shen, Z., Lo, K., Wang, L. L., Kuehl, B., Weld, D. S., and Downey, D · 2022
Later among the works it cites.
Mad: A scalable dataset for language grounding in videos from movie audio descriptions
Soldan, M., Pardo, A., Alcazar, J. L., Heilbron, F. C., Zhao, C., Giancola, S., and Ghanem, B · 2022
Later among the works it cites.
What makes reading comprehension questions difficult?
Sugawara, S., Nangia, N., Warstadt, A., and Bowman, S · 2022
Later among the works it cites.
Revise: A tool for measuring and mitigating bias in visual datasets
Wang, A., Liu, A., Zhang, R., Kleiman, A., Kim, L., Zhao, D., Shirai, I., Narayanan, A., and Russakovsky, O · 2022
Later among the works it cites.
The exploited labor behind artificial intelligence
Williams, A., Miceli, M., and Gebru, T · 2022
Later among the works it cites.
Enhancing fairness in face detection in computer vision systems by demographic bias mitigation
Yang, Y., Gupta, A., Feng, J., Singhal, P., Yadav, V., Wu, Y., Natarajan, P., Hedau, V., and Joo, J · 2022
Later among the works it cites.
Wild-time: A benchmark of in-the-wild distribution shift over time
Yao, H., Choi, C., Cao, B., Lee, Y., Koh, P. W. W., and Finn, C · 2022
Later among the works it cites.
Unifying panoptic segmentation for autonomous driving
Zendel, O., Schorghuber, M., Rainer, B., Murschitz, M., and Beleznai, C · 2022
Later among the works it cites.
Visible-thermal uav tracking: A large-scale benchmark and new baseline
Zhang, P., Zhao, J., Wang, D., Lu, H., and Ruan, X · 2022
Later among the works it cites.
Deconstructing nlg evaluation: Evaluation practices, assumptions, and their implications
Zhou, K., Blodgett, S. L., Trischler, A., Daumé III, H., Suleman, K., and Olteanu, A · 2022
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
Ethical considerations for collecting human-centric image datasets
Andrews, J. T., Zhao, D., Thong, W., Modas, A., Papakyriakopoulos, O., Nagpal, S., and Xiang, A · 2023
Later among the works it cites.
Representation in ai evaluations
Bergman, A. S., Hendricks, L. A., Rauh, M., Wu, B., Agnew, W., Kunesch, M., Duan, I., Gabriel, I., and Isaac, W · 2023
Later among the works it cites.
On hate scaling laws for data-swamps
Birhane, A., Prabhu, V., Han, S., and Boddeti, V. N · 2023
Later among the works it cites.
Making intelligence: Ethical values in iq and ml benchmarks
Blili-Hamelin, B. and Hancox-Li, L · 2023
Later among the works it cites.
The foundation model transparency index
Bommasani, R., Klyman, K., Longpre, S., Kapoor, S., Maslej, N., Xiong, B., Zhang, D., and Liang, P · 2023
Later among the works it cites.
Diaz, F. and Madaio, M · 2023
Later among the works it cites.
The vendi score: A diversity evaluation metric for machine learning
Friedman, D. and Dieng, A. B · 2023
Later among the works it cites.
Segment anything
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Later among the works it cites.
An american puzzle: Fitting race in a box
Lai, K. R. and Medina, J · 2023
Later among the works it cites.
“There’s no data like more data” automatic speech recognition and the making of algorithmic culture
Li, X · 2023
Later among the works it cites.
How hard are computer vision datasets? calibrating dataset difficulty to viewing time
Mayo, D., Cummings, J., Lin, X., Gutfreund, D., Katz, B., and Barbu, A · 2023
Later among the works it cites.
Flickr africa: Examining geo-diversity in large-scale, human-centric visual data
Naggita, K., LaChance, J., and Xiang, A · 2023
Later among the works it cites.
It takes two to tango: Navigating conceptualizations of nlp tasks and measurements of performance
Subramonian, A., Yuan, X., Daumé III, H., and Blodgett, S. L · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory
Xiao, Z., Zhang, S., Lai, V., and Liao, Q. V · 2023
Later among the works it cites.
Habitat-matterport 3d semantics dataset
Yadav, K., Ramrakhya, R., Ramakrishnan, S. K., Gervet, T., Turner, J., Gokaslan, A., Maestre, N., Chang, A. X., Batra, D., Savva, M., Clegg, A. W., and Chaplot, D. S · 2023
Later among the works it cites.
Addressing” documentation debt” in machine learning research: A retrospective datasheet for bookcorpus
Bandy, J. and Vincent, N · 2024
Closest in time.
Twigma: A dataset of ai-generated images with metadata from twitter
Chen, Y. and Zou, J · 2024
Closest in time.
Geode: a geographically diverse evaluation dataset for object recognition
Ramaswamy, V. V., Lin, S. Y., Zhao, D., Adcock, A. B., van der Maaten, L., Ghadiyaram, D., and Russakovsky, O · 2024
Closest in time.