Fetching the paper…
Reading the bibliography…
We present Answer-Me, a task-aware multi-task framework which unifies a variety of question answering tasks, such as, visual question answering, visual entailment, visual reasoning.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: CVPR (2009)
2009
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., Zitnick, C.L.: Microsoft COCO: Common objects in context. In: ECCV (2014)
2014
Earlier work this paper cites.
Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C.L., Batra, D., Parikh, D.: Vqa: Visual question answering. In: ICCV (2015)
2015
Earlier work this paper cites.
Bowman, S.R., Angeli, G., Potts, C., Manning, C.D.: A large annotated corpus for learning natural language inference. In: EMNLP (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. In: Advances in Neural Information Processing Systems (2015)
2015
Earlier work this paper cites.
Rohrbach, A., Rohrbach, M., Tandon, N., Schiele, B.: A dataset for movie description. In: CVPR (2015)
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A., Murphy, K.: Generation and comprehension of unambiguous object descriptions. In: CVPR (2016)
2016
Earlier work this paper cites.
Rohrbach, A., Rohrbach, M., Hu, R., Darrell, T., Schiele, B.: Grounding of textual phrases in images by reconstruction. In: ECCV (2016)
2016
Earlier work this paper cites.
Xu, J., Mei, T., Yao, T., Rui, Y.: Msr-vtt: A large video description dataset for bridging video and language. In: CVPR (2016)
2016
Earlier work this paper cites.
Yu, L., Poirson, P., Yang, S., Berg, A.C., Berg, T.L.: Modeling context in referring expressions. In: ECCV (2016)
2016
Earlier work this paper cites.
Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Moura, J.M.F., Parikh, D., Batra, D.: Visual dialog. In: CVPR (2017)
2017
Earlier work this paper cites.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In: CVPR (2017)
2017
Earlier work this paper cites.
Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., Lazebnik, S.: Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In: International Journal of Computer Vision (2017)
2017
Earlier work this paper cites.
Suhr, A., Lewis, M., Yeh, J., Artzi, Y.: A corpus of natural language for visual reasoning. In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: NeurIPS (2017)
2017
Earlier work this paper cites.
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering. In: CVPR (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Gurari, D., Li, Q., Stangl, A.J., Guo, A., Luo, C.L.K.G.J., Bigham, J.P.: VizWiz grand challenge: Answering visual questions from blind people. In: CVPR (2018)
2018
Earlier work this paper cites.
Kottur, S., Moura, J.M.F., Parikh, D., Batra, D., Rohrbach, M.: Visual coreference resolution in visual dialog using neural module networks. In: ECCV (2018)
2018
Earlier work this paper cites.
Margffoy-Tuay, E., Perez, J.C., Botero, E., Arbelaez, P.: Dynamic multimodal instance segmentation guided by natural language queries. In: ECCV (2018)
2018
Earlier work this paper cites.
Sharma, P., Ding, N., Goodman, S., Soricut, R.: Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: ACL (2018)
2018
Earlier work this paper cites.
Changpinyo, S., Pang, B., Sharma, P., Soricut, R.: Decoupled box proposal and featurization with ultrafine-grained semantic labels improve image captioning and visual question answering. In: Conference on Empirical Methods in Natural Language Processing (EMNLP) (2019)
2019
Earlier work this paper cites.
Hudson, D.A., Manning, C.D.: Gqa: a new dataset for compositional question answering over realworld images. In: CVPR (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Liu, D., Zhang, H., Wu, F., Zha, Z.J.: Learning to assemble neural module tree networks for visual grounding. In: ICCV (2019)
2019
Cited alongside, same era.
Liu, X., He, P., Chen, W., Gao, J.: Multi-task deep neural networks for natural language understanding. In: In Proceedings of ACL (2019)
2019
Cited alongside, same era.
Lu, J., Batra, D., Parikh, D., Lee, S.: Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In: CVPR (2019)
2019
Cited alongside, same era.
Nguyen, D.K., Okatani, T.: Multi-task learning of hierarchical vision-language representation. In: CVPR (2019)
2019
Cited alongside, same era.
Peng, G., Jiang, Z., You, H., Lu, P., Hoi, S., Wang, X., Li, H.: Dynamic fusion with intra- and inter- modality attention flow for visual question answering. In: CVPR (2019)
2019
Yang, Z., Chen, T., Wang, L., Luo, J.: Improving one-stage visual grounding by recursive sub-query construction. In: ECCV (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Changpinyo, S., Pont-Tuset, J., Ferrari, V., Soricut, R.: Telling the what while pointing to the where: Multimodal queries for image retrieval. In: Arxiv 2102.04980 (2021)
2021
Later among the works it cites.
Changpinyo, S., Sharma, P., Ding, N., Soricut, R.: Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In: CVPR (2021)
2021
Later among the works it cites.
Cho, J., Lei, J., Tan, H., Bansal, M.: Unifying vision-and-language tasks via text generation. In: Arxiv 2102.02779 (2021)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2019
Cited alongside, same era.
Singh, A., Natarjan, V., Shah, M., Jiang, Y., Chen, X., Parikh, D., Rohrbach, M.: Towards vqa models that can read. In: CVPR (2019)
2019
Cited alongside, same era.
Sun, C., Myers, A., Vondrick, C., Murphy, K., Schmid, C.: Videobert: A joint model for video and language representation learning. In: ICCV (2019)
2019
Cited alongside, same era.
Tan, H., Bansal, M.: Lxmert: Learning cross-modality encoder representations from transformers. In: EMNLP (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Yang, Z., Gong, B., Wang, L., Huang, W., Yu, D., Luo, J.: A fast and accurate one-stage approach to visual grounding. In: ICCV (2019)
2019
Cited alongside, same era.
Zellers, R., Bisk, Y., Farhadi, A., Choi, Y.: From recognition to cognition: Visual commonsense reasoning. In: CVPR (June 2019)
2019
Cited alongside, same era.
2021
Later among the works it cites.
Deng, C., Chen, S., Chen, D., He, Y., Wu, Q.: Sketch, ground, and refine: Top-down dense video captioning. In: CVPR (2021)
2021
Later among the works it cites.
Deng, J., Yang, Z., Chen, T., Zhou, W., Li, H.: Transvg: End-to-end visual grounding with transformers. ICCV (2021)
2021
Later among the works it cites.
Desai, K., Johnson, J.: Virtex: Learning visual representations from textual annotations. In: CVPR (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Huang, Z., Zeng, Z., Huang, Y., Liu, B., Fu, D., Fu, J.: Seeing out of the box:end-to-end pre-training for vision-language representation learning. In: CVPR (2021)
2021
Later among the works it cites.
Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q.V., Sung, Y., Li, Z., Duerig, T.: Scaling up visual and vision-language representation learning with noisy text supervision. In: ICML (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Kim, W., Son, B., Kim, I.: Vilt: Vision-and-language transformer without convolution or region supervision. In: ICML (2021)
2021
Later among the works it cites.
Lin, X., Bertasius, G., Wang, J., Chang, S.F., Parikh, D.: Vx2text: End-to-end learning of video-based text generation from multimodal inputs. In: CVPR (2021)
2021
Later among the works it cites.
Liu, Y., Huang, L., Song, L., Wang, B., Zhang, Y., Pan, P.: Enhancing textual cues in multi-modal transformers for VQA. In: VizWiz Challenge 2021 (2021)
2021
Later among the works it cites.
Liu, Z., Stent, S., Li, J., Gideon, J., Han, S.: Loctex: Learning data-efficient visual representations from localized textual supervision. In: ICCV (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: ICML (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Whitehead, S., Wu, H., Ji, H., Feris, R., Saenko, K.: Separating skills and concepts for novel visual question answering. In: CVPR (2021)
2021
Later among the works it cites.
Yin, D., Li, L.H., Hu, Z., Peng, N., Chang, K.W.: Broaden the vision: Geo-diverse visual commonsense reasoning. In: EMNLP (2021)
2021
Later among the works it cites.
Yu, F., Tang, J., Yin, W., Sun, Y., Tian, H., Wu, H., Wang, H.: Ernie-vil: Knowledge enhanced vision-language representations through scene graph. In: AAAI (2021)
2021
Later among the works it cites.
Yu, L., Lin, Z., Shen, X., Yang, J., Lu, X., Bansal, M., Berg, T.L.: Mattnet: Modular attention network for referring expression comprehension. In: CVPR (2021)
2021
Later among the works it cites.
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., Gao, J.: Vinvl: Revisiting visual representations in vision-language models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 5579–5588 (2021)
2021
Later among the works it cites.