Fots: Fast oriented text spotting with a unified network, in: CVPR
Liu, X., Liang, D., Yan, S., Chen, D., Qiao, Y., Yan, J., 2018 · 2018
Later among the works it cites.
Yolov3: An incremental improvement
Original
Redmon, J., Farhadi, A., 2018 · 2018
Later among the works it cites.
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Yu, Z., Yu, J., Xiang, C., Fan, J., Tao, D., 2018 · 2018
Later among the works it cites.
What is wrong with scene text recognition model comparisons? dataset and model analysis, in: ICCV
Baek, J., Kim, G., Lee, J., Park, S., Han, D., Yun, S., Oh, S.J., Lee, H., 2019 · 2019
Later among the works it cites.
Icdar 2019 competition on scene text visual question answering, in: ICDAR
Biten, A.F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Mathew, M., Jawahar, C., Valveny, E., Karatzas, D., 2019a · 2019
Later among the works it cites.
Ocr-vqa: Visual question answering by reading text in images, in: ICDAR
Mishra, A., Shekhar, S., Singh, A.K., Chakraborty, A., 2019 · 2019
Later among the works it cites.
Pythia-a platform for vision & language research, in: SysML Workshop, NeurIPS 2019
Singh, A., Natarajan, V., Jiang, Y., Chen, X., Shah, M., Rohrbach, M., Batra, D., Parikh, D., 2018 · 2019
Later among the works it cites.
Towards vqa models that can read, in: CVPR
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., Rohrbach, M., 2019 · 2019
Later among the works it cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., Duerig, T., Ferrari, V., 2020 · 2020
Closest in time.