Fetching the paper…
Reading the bibliography…
Vision-and-language tasks have increasingly drawn more attention as a means to evaluate human-like reasoning in machine learning models.
Pushing the frontiers of unconstrained face detection and recognition: IARPA Janus Benchmark A. In CVPR . IEEE Computer Society, 1931–1939
Brendan F. Klare, Ben Klein, Emma Taborsky, Austin Blanton, Jordan Cheney, Kristen Allen, Patrick Grother, Alan Mah, Mark James Burge, and Anil K. Jain. 2015 · 1939
Earlier work this paper cites.
MUREL: Multimodal Relational Reasoning for Visual Question Answering. In CVPR . Computer Vision Foundation / IEEE, 1989–1998
Rémi Cadène, Hedi Ben-younes, Matthieu Cord, and Nicolas Thome. 2019a · 1998
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In ECCV (5) (Lecture Notes in Computer Science, Vol. 8693) . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering. In ICCV . IEEE Computer Society, 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai. 2016 · 2016
Earlier work this paper cites.
U.S. government to stop using these words to refer to minorities
Madison Park. 2016 · 2016
Earlier work this paper cites.
YFCC100M: the new data in multimedia research
Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. 2016 · 2016
Earlier work this paper cites.
Boys’ and Girls’ Educational Choices in Secondary Education. The Role of Gender Ideology
Maaike Van der Vleuten, Eva Jaspers, Ineke Maas, and Tanja van der Lippe. 2016 · 2016
Earlier work this paper cites.
Image Captioning with Semantic Attention. In CVPR . IEEE Computer Society, 4651–4659
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. 2016 · 2016
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images. In CVPR . IEEE Computer Society, 4995–5004
Yuke Zhu, Oliver Groth, Michael S. Bernstein, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
MUTAN: Multimodal Tucker Fusion for Visual Question Answering. In ICCV . IEEE Computer Society, 2631–2639
Hedi Ben-younes, Rémi Cadène, Matthieu Cord, and Nicolas Thome. 2017 · 2017
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In CVPR . IEEE Computer Society, 6325–6334
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
No Classification Without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World
Shreya Shankar, Yoni Halpern, Eric Breck, James Atwood, Jimbo Wilson, and D Sculley. 2017 · 2017
Earlier work this paper cites.
Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2017 · 2017
Earlier work this paper cites.
Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints. In EMNLP . Association for Computational Linguistics, 2979–2989
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017 · 2017
Earlier work this paper cites.
Don’t Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering. In CVPR . Computer Vision Foundation / IEEE Computer Society, 4971–4980
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018 · 2018
Earlier work this paper cites.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR . Computer Vision Foundation / IEEE Computer Society, 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. In FAT (Proceedings of Machine Learning Research, Vol. 81) . PMLR, 77–91
Joy Buolamwini and Timnit Gebru. 2018 · 2018
Cited alongside, same era.
VizWiz Grand Challenge: Answering Visual Questions From Blind People. In CVPR . Computer Vision Foundation / IEEE Computer Society, 3608–3617
Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018 · 2018
Cited alongside, same era.
Women Also Snowboard: Overcoming Bias in Captioning Models. In ECCV (3) (Lecture Notes in Computer Science, Vol. 11207) . Springer, 793–811
Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach. 2018 · 2018
Cited alongside, same era.
OSCAR: Object-Semantics Aligned Pre-training for Vision-Language Tasks. In ECCV
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Later among the works it cites.
Visual Commonsense R-CNN. In CVPR . Computer Vision Foundation / IEEE, 10757–10767
Tan Wang, Jianqiang Huang, Hanwang Zhang, and Qianru Sun. 2020 · 2020
Later among the works it cites.
BERT Representations for Video Question Answering. In WACV . IEEE, 1545–1554
Zekun Yang, Noa Garcia, Chenhui Chu, Mayu Otani, Yuta Nakashima, and Haruo Takemura. 2020 · 2020
Later among the works it cites.
1st Workshop on Language for 3D Scenes
Panos Achlioptas, Zhenyu Chen, Mohamed Elhoseiny, Angel X Chang, Matthias Niessner, and Leonidas Guibas. 2021 · 2021
Later among the works it cites.
Workshop on Multilingual Multimodal Learning
Emanuele Bugliarello, Kai-Wei Chang, Desmond Elliott, Spandana Gella, Aishwarya Kamath, Liunian Harold Li, Fangyu Liu, Jonas Pfeiffer, Edoardo M. Ponti, Krishna Srinivasan, Ivan Vulić, Yinfei Yang, and Da Yin. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bilinear Attention Networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018 · 2018
Cited alongside, same era.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In ACL (1) . Association for Computational Linguistics, 2556–2565
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Cited alongside, same era.
RUBi: Reducing Unimodal Biases for Visual Question Answering
Rémi Cadène, Corentin Dancette, Hedi Ben-younes, Matthieu Cord, and Devi Parikh. 2019b · 2019
Cited alongside, same era.
Don’t Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering. In CVPR . Computer Vision Foundation / IEEE, 6700–6709
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
VisualBert: A Simple and Performant Baseline for Vision and Language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 2019
Cited alongside, same era.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Explicit Bias Discovery in Visual Question Answering Models. In CVPR . Computer Vision Foundation / IEEE, 9562–9571
Varun Manjunatha, Nirat Saini, and Larry S. Davis. 2019 · 2019
Cited alongside, same era.
4th Workshop on Closing the Loop Between Vision and Language
Mohamed Elhoseiny, Xin Eric Wang, Andrew Brown, Anna Rohrbach, and Marcus Rohrbach. 2021 · 2021
Later among the works it cites.
Visual Question Answering with Textual Representations for Images. In ICCVW . IEEE, 3147–3150
Yusuke Hirota, Noa Garcia, Mayu Otani, Chenhui Chu, Yuta Nakashima, Ittetsu Taniguchi, and Takao Onoye. 2021 · 2021
Later among the works it cites.
Seeing Out of the Box: End-to-End Pre-Training for Vision-Language Representation Learning. In CVPR . Computer Vision Foundation / IEEE, 12976–12985
Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu. 2021 · 2021
Later among the works it cites.
Roses Are Red, Violets Are Blue… but Should VQA Expect Them To?. In CVPR . Computer Vision Foundation / IEEE, 2776–2785
Corentin Kervadec, Grigory Antipov, Moez Baccouche, and Christian Wolf. 2021 · 2021
Later among the works it cites.
One Label, One Billion Faces: Usage and Consistency of Racial Categories in Computer Vision. In FAccT . ACM, 587–597
Zaid Khan and Yun Fu. 2021 · 2021
Later among the works it cites.
LANTERN - The Third Workshop Beyond Vision and Language: Integrating Real World Knowledge
Marius Mosbach, Sandro Pezzelle, Michael A. Hedderich, Dietrich Klakow, Marie-Francine Moens, and Zeynep Akata. 2021 · 2021
Later among the works it cites.
Counterfactual VQA: A Cause-Effect Look at Language Bias. In CVPR . Computer Vision Foundation / IEEE, 12700–12710
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen. 2021 · 2021
Later among the works it cites.
Visual Question Answering Workshop
Ayush Shrivastava, Yash Mukund Kant, Satwik Kottur, Dhruv Batra, Devi Parikh, and Aishwarya Agrawal. 2021 · 2021
Later among the works it cites.
Mitigating Gender Bias in Captioning Systems. In WWW . ACM / IW3C2, 633–645
Ruixiang Tang, Mengnan Du, Yuening Li, Zirui Liu, Na Zou, and Xia Hu. 2021 · 2021
Later among the works it cites.
Feature and Label Embedding Spaces Matter in Addressing Image Classifier Bias. In BMVC
William Thong and Cees GM Snoek. 2021 · 2021
Later among the works it cites.
From VQA to VLN: Recent Advances in Vision-and-Language Research
Qi Wu and Zhe Gan. 2021 · 2021
Later among the works it cites.
VinVL: Revisiting Visual Representations in Vision-Language Models. In CVPR . Computer Vision Foundation / IEEE, 5579–5588
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
Understanding and Evaluating Racial Biases in Image Captioning. In ICCV . IEEE, 14810–14820
Dora Zhao, Angelina Wang, and Olga Russakovsky. 2021 · 2021
Later among the works it cites.
Quantifying Societal Bias Amplification in Image Captioning. In CVPR
Yusuke Hirota, Yuta Nakashima, and Noa Garcia. 2022 · 2022
Closest in time.