Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) entails answering questions about images.
Lin, C.Y.: Rouge: A package for automatic evaluation of summaries. In: Text summarization branches out. pp. 74–81 (2004)
2004
Earlier work this paper cites.
Ignatova, K., Toprak, C., Bernhard, D., Gurevych, I.: Annotating question types in social q&a sites. In: Tagungsband des GSCL Symposiums ‘Sprachtechnologie und eHumanities. pp. 44–49 (2009)
2009
Earlier work this paper cites.
Harper, F.M., Weinberg, J., Logie, J., Konstan, J.A.: Question types in social q&a sites. First Monday (2010)
2010
Earlier work this paper cites.
Nie, L., Wang, M., Zha, Z., Li, G., Chua, T.S.: Multimedia answering: enriching text qa with media information. In: Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval. pp. 695–704 (2011)
2011
Earlier work this paper cites.
Ordonez, V., Kulkarni, G., Berg, T.: Im2text: Describing images using 1 million captioned photographs. Advances in neural information processing systems 24
2011
Earlier work this paper cites.
Chen, L., Zhang, D., Mark, L.: Understanding user intent in community question answering. In: Proceedings of the 21st international conference on world wide web. pp. 823–828 (2012)
2012
Earlier work this paper cites.
Brady, E., Morris, M.R., Zhong, Y., White, S., Bigham, J.P.: Visual challenges in the everyday lives of blind people. In: Proceedings of the SIGCHI conference on human factors in computing systems. pp. 2117–2126 (2013)
2013
Earlier work this paper cites.
Elliott, D., Keller, F.: Image description using visual dependency representations. In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. pp. 1292–1302 (2013)
2013
Earlier work this paper cites.
Toba, H., Ming, Z.Y., Adriani, M., Chua, T.S.: Discovering high quality answers in community question answering archives using a hierarchy of classifiers. Information Sciences 261
2014
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., Parikh, D.: Vqa: Visual question answering. In: Proceedings of the IEEE international conference on computer vision. pp. 2425–2433 (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Liu, Z.: Understanding and modeling user behavior in social question and answering (2015)
2015
Earlier work this paper cites.
Fu, H., Fan, Y.: Music information seeking via social q&a: An analysis of questions in music stackexchange community. In: Proceedings of the 16th ACM/IEEE-CS on Joint Conference on Digital Libraries. pp. 139–142 (2016)
2016
Earlier work this paper cites.
Kofler, C., Larson, M., Hanjalic, A.: User intent in multimedia search: a survey of the state of the art and future challenges. ACM Computing Surveys (CSUR) 49
2016
Earlier work this paper cites.
Srba, I., Bielikova, M.: A comprehensive survey and classification of approaches for community question answering. ACM Transactions on the Web (TWEB) 10
2016
Earlier work this paper cites.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. In: Conference on Computer Vision and Pattern Recognition (CVPR) (2017)
2017
Earlier work this paper cites.
Gurari, D., Grauman, K.: Crowdverge: Predicting if people will agree on the answer to a visual question. In: Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. pp. 3511–3522 (2017)
2017
Earlier work this paper cites.
He, B., Xia, M., Yu, X., Jian, P., Meng, H., Chen, Z.: An educational robot system of visual question answering for preschoolers. In: 2017 2nd International Conference on Robotics and Automation Engineering (ICRAE). pp. 441–445. IEEE (2017)
2017
Earlier work this paper cites.
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., Girshick, R.: Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2901–2910 (2017)
2017
Earlier work this paper cites.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al.: Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision 123
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Wu, Q., Teney, D., Wang, P., Shen, C., Dick, A., van den Hengel, A.: Visual question answering: A survey of methods and datasets. Computer Vision and Image Understanding 163
2017
Earlier work this paper cites.
Gurari, D., Li, Q., Stangl, A.J., Guo, A., Lin, C., Grauman, K., Luo, J., Bigham, J.P.: Vizwiz grand challenge: Answering visual questions from blind people. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3608–3617 (2018)
2018
Earlier work this paper cites.
Lei, J., Yu, L., Bansal, M., Berg, T.: Tvqa: Localized, compositional video question answering. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. pp. 1369–1379 (2018)
2018
Earlier work this paper cites.
Sharma, P., Ding, N., Goodman, S., Soricut, R.: Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 2556–2565 (2018)
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Hudson, D.A., Manning, C.D.: Gqa: A new dataset for real-world visual reasoning and compositional question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6700–6709 (2019)
2019
Earlier work this paper cites.
Marino, K., Rastegari, M., Farhadi, A., Mottaghi, R.: Ok-vqa: A visual question answering benchmark requiring external knowledge. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2019)
2019
Earlier work this paper cites.
Mishra, A., Shekhar, S., Singh, A.K., Chakraborty, A.: Ocr-vqa: Visual question answering by reading text in images. In: 2019 international conference on document analysis and recognition (ICDAR). pp. 947–952. IEEE (2019)
2019
Cited alongside, same era.
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., Rohrbach, M.: Towards vqa models that can read. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8317–8326 (2019)
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al.: Solving quantitative reasoning problems with language models. Advances in Neural Information Processing Systems 35
2022
Later among the works it cites.
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al.: Competition-level code generation with alphacode. Science 378
2022
Later among the works it cites.
Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., Jawahar, C.: Infographicvqa. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). pp. 1697–1706 (January 2022)
2022
Later among the works it cites.
Saikh, T., Ghosal, T., Mittal, A., Ekbal, A., Bhattacharyya, P.: Scienceqa: a novel resource for question answering on scholarly articles. International Journal on Digital Libraries 23
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Khot, T., Clark, P., Guerquin, M., Jansen, P., Sabharwal, A.: Qasc: A dataset for question answering via sentence composition. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 8082–8090 (2020)
2020
Cited alongside, same era.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J., et al.: Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21
2020
Cited alongside, same era.
Yasunaga, M., Liang, P.: Graph-based, self-supervised program repair from diagnostic feedback. In: International Conference on Machine Learning. pp. 10799–10808. PMLR (2020)
2020
Cited alongside, same era.
Zeng, X., Wang, Y., Chiu, T.Y., Bhattacharya, N., Gurari, D.: Vision skills needed to answer visual questions. Proceedings of the ACM on Human-Computer Interaction 4
2020
Cited alongside, same era.
Aggarwal, S., Mandowara, D., Agrawal, V., Khandelwal, D., Singla, P., Garg, D.: Explanations for commonsenseqa: New dataset and models. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). pp. 3050–3065 (2021)
2021
Cited alongside, same era.
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al.: Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems 35
2022
Later among the works it cites.
2022
Later among the works it cites.
Bitton-Guetta, N., Bitton, Y., Hessel, J., Schmidt, L., Elovici, Y., Stanovsky, G., Schwartz, R.: Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional images. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 2616–2627 (2023)
2023
Closest in time.
2023
Closest in time.
Dai, W., Li, J., Li, D., Tiong, A.M.H., Zhao, J., Wang, W., Li, B., Fung, P., Hoi, S.: Instructblip: Towards general-purpose vision-language models with instruction tuning (2023)
2023
Closest in time.
Exchange, S.: Gpt-4v pricing (2023), https://platform.openai.com/docs/guides/vision , accessed: 2023-11-16
2023
Closest in time.
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., Cao, Y.: Eva: Exploring the limits of masked visual representation learning at scale. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 19358–19369 (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
OpenAI: Gpt-4 technical report. ArXiv abs/2303.08774
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Jiang, C., Li, C., Xu, Y., Chen, H., Tian, J., Qi, Q., Zhang, J., Huang, F.: mplug-owl: Modularization empowers large language models with multimodality (2023)
2023
Closest in time.
2023
Closest in time.
González-Chávez, O., Ruiz, G., Moctezuma, D., Ramirez-delReal, T.: Are metrics measuring what they should? an evaluation of image captioning task metrics. Signal Processing: Image Communication 120
2024
Closest in time.
Zheng, L., Chiang, W.L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al.: Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36
2024
Closest in time.