Fetching the paper…
Reading the bibliography…
Large datasets underlying much of current machine learning raise serious issues concerning inappropriate content such as offensive, insulting, threatening, or might otherwise cause anxiety.
The Open Images Dataset V4
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper R. R. Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. 2020 · 1981
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR) . 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Semi-supervised Sequence Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . 3079–3087
Andrew M. Dai and Quoc V. Le. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, (CVPR) . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
OpenImages: A public dataset for large-scale multi-label and multi-class image classification
Ivan Krasin, Tom Duerig, Neil Alldrin, Andreas Veit, Sami Abu-El-Haija, Serge Belongie, David Cai, Zheyun Feng, Vittorio Ferrari, Victor Gomes, Abhinav Gupta, Dhyanesh Narayanan, Chen Sun, Gal Chechik, and Kevin Murphy. 2016 · 2016
Earlier work this paper cites.
The Socio-Moral Image Database (SMID): A novel stimulus set for the study of social, moral and affective processes
Damien L. Crone, Stefan Bode, Carsten Murawski, and Simon M. Laham. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Semantics Derived Automatically from Language Corpora Contain Human-like Moral Choices. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES) . 37–44
Sophie Jentzsch, Patrick Schramowski, Constantin A. Rothkopf, and Kristian Kersting. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 2463–2473
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019 · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning, (ICML) (Proceedings of Machine Learning Research, Vol. 97) . 6105–6114
Mingxing Tan and Quoc V. Le. 2019 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . 1–25
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Scalable Detection of Offensive and Non-compliant Content / Logo in Product Images. In Proceedings of IEEE Winter Conference on Applications of Computer Vision (WACV) . 2236–2245
Shreyansh Gandhi, Samrat Kokkula, Abon Chaudhuri, Alessandro Magnani, Theban Stanley, Behzad Ahmadi, Venkatesh Kandaswamy, Omer Ovenc, and Shie Mannor. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings (EMNLP) . 3356–3369
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
XHate-999: Analyzing and Detecting Abusive Language Across Domains and Languages. In Proceedings of the 28th International Conference on Computational Linguistics . International Committee on Computational Linguistics, 6350–6365
Goran Glavaš, Mladen Karan, and Ivan Vulić. 2020 · 2020
Earlier work this paper cites.
Exploring Hate Speech Detection in Multimodal Publications. In Proceedings of IEEE Winter Conference on Applications of Computer Vision (WACV) . 1459–1467
Raul Gomez, Jaume Gibert, Lluís Gómez, and Dimosthenis Karatzas. 2020 · 2020
Earlier work this paper cites.
Fortifying Toxic Speech Detectors Against Veiled Toxicity. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 7732–7739
Xiaochuang Han and Yulia Tsvetkov. 2020 · 2020
Earlier work this paper cites.
Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis
Agostina J. Larrazabal, Nicolás Nieto, Victoria Peterson, Diego H. Milone, and Enzo Ferrante. 2020 · 2020
Cited alongside, same era.
How Much Knowledge Can You Pack Into the Parameters of a Language Model?. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 5418–5426
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Cited alongside, same era.
Social Bias Frames: Reasoning about Social and Power Implications of Language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 5477–5490
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
The Moral Choice Machine
Patrick Schramowski, Cigdem Turan, Sophie Jentzsch, Constantin A. Rothkopf, and Kristian Kersting. 2020 · 2020
Cited alongside, same era.
REVISE: A Tool for Measuring and Mitigating Bias in Visual Datasets. In Proceedings of 16th European Conference of Computer Vision (ECCV) . 733–751
WARP: Word-level Adversarial ReProgramming
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021 · 2021
Later among the works it cites.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. In Proceedings of the International Conference on Machine Learning, (ICML) . 4904–4916
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Later among the works it cites.
The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Online and Punta Cana, Dominican Republic, 3045–3059
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Later among the works it cites.
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Angelina Wang, Arvind Narayanan, and Olga Russakovsky. 2020 · 2020
Cited alongside, same era.
Towards fairer datasets: filtering and balancing the distribution of the people subtree in the ImageNet hierarchy. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAccT) . 547–558
Kaiyu Yang, Klint Qinami, Li Fei-Fei, Jia Deng, and Olga Russakovsky. 2020 · 2020
Cited alongside, same era.
Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus. In Proceedings of NeurIPS Datasets and Benchmarks . 1–13
Jack Bandy and Nicholas Vincent. 2021 · 2021
Cited alongside, same era.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of ACM Conference on Fairness, Accountability, and Transparency (FAccT) . 610–623
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Large image datasets: A pyrrhic win for computer vision?. In Proceedings of IEEE Winter Conference on Applications of Computer Vision (WACV) . 1536–1546
Abeba Birhane and Vinay Uday Prabhu. 2021 · 2021
Cited alongside, same era.
High-Performance Large-Scale Image Recognition Without Normalization. In Proceedings of the 38th International Conference on Machine Learning, (ICML) . 1059–1071
Andy Brock, Soham De, Samuel L. Smith, and Karen Simonyan. 2021 · 2021
Cited alongside, same era.
Transformer Interpretability Beyond Attention Visualization. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 782–791
Hila Chefer, Shir Gur, and Lior Wolf. 2021 · 2021
Cited alongside, same era.
On the genealogy of machine learning datasets: A critical history of ImageNet
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole. 2021 · 2021
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP) . 5356–5371
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Later among the works it cites.
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021 · 2021
Later among the works it cites.
Learning How to Ask: Querying LMs with Mixtures of Soft Prompts. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, (NAACL-HLT) . 5203–5212
Guanghui Qin and Jason Eisner. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the International Conference on Machine Learning (ICML) . 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Zero-Shot Text-to-Image Generation. In Proceedings of the International Conference on Machine Learning (ICML) . 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases. In Proceedings of ACM Conference on Fairness, Accountability, and Transparency (FAccT) . 701–713
Ryan Steed and Aylin Caliskan. 2021 · 2021
Later among the works it cites.
EfficientNetV2: Smaller Models and Faster Training. In Proceedings of the 38th International Conference on Machine Learning, (ICML) (Proceedings of Machine Learning Research, Vol. 139) . 10096–10106
Mingxing Tan and Quoc V. Le. 2021 · 2021
Later among the works it cites.
Multimodal Few-Shot Learning with Frozen Language Models. In Advances in Neural Information Processing Systems
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. 2021 · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Later among the works it cites.
Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022 · 2022
Closest in time.
Large pre-trained language models contain human-like biases of what is right and wrong to do
Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A. Rothkopf, and Kristian Kersting. 2022 · 2022
Closest in time.