Fetching the paper…
Reading the bibliography…
Benefiting from large-scale pretrained vision language models (VLMs), the performance of visual question answering (VQA) has approached human oracles.
Cho J, Lei J, Tan H, Bansal M (2021) Unifying vision-and-language tasks via text generation. In: International Conference on Machine Learning, pp 1931–1942
1942
Earlier work this paper cites.
Tishby N, Pereira FC, Bialek W (2000) The information bottleneck method. arXiv preprint physics/0004057
2000
Earlier work this paper cites.
Barber D, Agakov F (2003) The im algorithm: a variational approach to information maximization. In: Neural Information Processing Systems, pp 201–208
2003
Earlier work this paper cites.
Nguyen X, Wainwright MJ, Jordan MI (2010) Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Transactions on Information Theory 56(11):5847–5861
2010
Earlier work this paper cites.
Ordonez V, Kulkarni G, Berg T (2011) Im2text: Describing images using 1 million captioned photographs. In: Neural Information Processing Systems, pp 1143–1151
2011
Earlier work this paper cites.
Antol S, Agrawal A, Lu J, Mitchell M, Batra D, Zitnick CL, Parikh D (2015) Vqa: Visual question answering. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 2425–2433
2015
Earlier work this paper cites.
Tishby N, Zaslavsky N (2015) Deep learning and the information bottleneck principle. In: IEEE Information Theory Workshop, pp 1–5
2015
Earlier work this paper cites.
Bennasar M, Hicks Y, Setchi R (2015) Feature selection using joint mutual information maximisation. Expert Systems with Applications 42(22):8520–8532
2015
Earlier work this paper cites.
Chen X, Fang H, Lin TY, Vedantam R, Gupta S, Dollár P, Zitnick CL (2015) Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:150400325
2015
Earlier work this paper cites.
Lu J, Lin X, Batra D, Parikh D (2015) Deeper lstm and normalized cnn visual question answering model
2015
Earlier work this paper cites.
Yu L, Poirson P, Yang S, Berg AC, Berg TL (2016) Modeling context in referring expressions. In: European Conference on Computer Vision, pp 69–85
2016
Earlier work this paper cites.
Zhu Y, Groth O, Bernstein M, Fei-Fei L (2016) Visual7w: Grounded question answering in images. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 4995–5004
2016
Earlier work this paper cites.
Yang Z, He X, Gao J, Deng L, Smola A (2016) Stacked attention networks for image question answering. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 21–29
2016
Earlier work this paper cites.
Goyal Y, Khot T, Summers-Stay D, Batra D, Parikh D (2017) Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6904–6913
2017
Earlier work this paper cites.
Shwartz-Ziv R, Tishby N (2017) Opening the black box of deep neural networks via information. arXiv preprint arXiv:170300810
2017
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I (2017) Attention is all you need. In: Neural Information Processing Systems, pp 5998–6008
2017
Earlier work this paper cites.
Krishna R, Zhu Y, Groth O, Johnson J, Hata K, Kravitz J, Chen S, Kalantidis Y, Li LJ, Shamma DA, et al. (2017) Visual Genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision 123(1):32–73
2017
Earlier work this paper cites.
Kazemi V, Elqursh A (2017) Show, ask, attend, and answer: A strong baseline for visual question answering. arXiv preprint arXiv:170403162
2017
Earlier work this paper cites.
Anderson P, He X, Buehler C, Teney D, Johnson M, Gould S, Zhang L (2018) Bottom-up and top-down attention for image captioning and visual question answering. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6077–6086
2018
Earlier work this paper cites.
Jiang Y, Natarajan V, Chen X, Rohrbach M, Batra D, Parikh D (2018) Pythia v0. 1: the winning entry to the vqa challenge 2018. arXiv preprint arXiv:180709956
2018
Earlier work this paper cites.
Kim JH, Jun J, Zhang BT (2018) Bilinear attention networks. In: Neural Information Processing Systems, pp 1564–1574
2018
Earlier work this paper cites.
Sharma P, Ding N, Goodman S, Soricut R (2018) Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Annual Meeting of the Association for Computational Linguistics, pp 2556–2565
2018
Earlier work this paper cites.
Hu R, Andreas J, Darrell T, Saenko K (2018) Explainable neural computation via stack neural module networks. In: European Conference on Computer Vision, pp 53–69
2018
Earlier work this paper cites.
Oord Avd, Li Y, Vinyals O (2018) Representation learning with contrastive predictive coding. arXiv preprint arXiv:180703748
2018
Earlier work this paper cites.
Belghazi MI, Baratin A, Rajeswar S, Ozair S, Bengio Y, Courville A, Hjelm RD (2018) Mutual information neural estimation. In: International Conference on Machine Learning, pp 530–539
2018
Earlier work this paper cites.
Shah M, Chen X, Rohrbach M, Parikh D (2019) Cycle-consistency for robust visual question answering. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6649–6658
2019
Earlier work this paper cites.
Li LH, Yatskar M, Yin D, Hsieh CJ, Chang KW (2019) Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:190803557
2019
Earlier work this paper cites.
Lu J, Batra D, Parikh D, Lee S (2019) ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In: Neural Information Processing Systems, pp 13–23
2019
Earlier work this paper cites.
Tan H, Bansal M (2019) LXMERT: Learning cross-modality encoder representations from transformers. In: Conference on Empirical Methods in Natural Language Processing, pp 5099–5110
2019
Earlier work this paper cites.
Clark C, Yatskar M, Zettlemoyer L (2019) Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases. In: Conference on Empirical Methods in Natural Language Processing, pp 4067–4080
2019
Earlier work this paper cites.
Poole B, Ozair S, Van Den Oord A, Alemi A, Tucker G (2019) On variational bounds of mutual information. In: International Conference on Machine Learning, pp 5171–5180
2019
Earlier work this paper cites.
Hudson DA, Manning CD (2019) Gqa: A new dataset for real-world visual reasoning and compositional question answering. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6700–6709
2019
Earlier work this paper cites.
Shi J, Zhang H, Li J (2019) Explainable and explicit visual reasoning over scene graphs. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 8376–8384
2019
Earlier work this paper cites.
Ben-Younes H, Cadene R, Thome N, Cord M (2019) Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection. In: Association for the Advancement of Artificial Intelligence, pp 8102–8109
2019
Earlier work this paper cites.
Cadene R, Dancette C, Cord M, Parikh D, et al. (2019) RUBi: Reducing unimodal biases for visual question answering. In: Neural Information Processing Systems, pp 841–852
2019
Cited alongside, same era.
Liu X, Li L, Wang S, Zha ZJ, Meng D, Huang Q (2019) Adaptive reconstruction network for weakly supervised referring expression grounding. In: IEEE International Conference on Computer Vision, pp 2611–2620
2019
Cited alongside, same era.
Agarwal V, Shetty R, Fritz M (2020) Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 9690–9698
2020
Cited alongside, same era.
Su W, Zhu X, Cao Y, Li B, Lu L, Wei F, Dai J (2020) VL-BERT: pre-training of generic visual-linguistic representations. In: International Conference on Learning Representations
2020
Cited alongside, same era.
Bao F (2021) Disentangled variational information bottleneck for multiview representation learning. In: International Conference on Artificial Intelligence, pp 91–102
2021
Later among the works it cites.
Mahabadi RK, Belinkov Y, Henderson J (2021) Variational information bottleneck for effective low-resource fine-tuning. In: International Conference on Learning Representations
2021
Later among the works it cites.
Wang B, Wang S, Cheng Y, Gan Z, Jia R, Li B, Liu J (2021) InfoBERT: Improving robustness of language models from an information theoretic perspective. In: International Conference on Learning Representations
2021
Later among the works it cites.
Dong X, Luu AT, Lin M, Yan S, Zhang H (2021) How should pre-trained language models be fine-tuned towards adversarial robustness? In: Neural Information Processing Systems, pp 4356–4369
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen YC, Li L, Yu L, El Kholy A, Ahmed F, Gan Z, Cheng Y, Liu J (2020) UNITER: Universal image-text representation learning. In: European Conference on Computer Vision, pp 104–120
2020
Cited alongside, same era.
Whitehead S, Wu H, Fung YR, Ji H, Feris R, Saenko K (2020) Learning from lexical perturbations for consistent visual question answering. arXiv preprint arXiv:201113406
2020
Cited alongside, same era.
Teney D, Kafle K, Shrestha R, Abbasnejad E, Kanan C, Hengel Avd (2020) On the value of out-of-distribution testing: An example of goodhart’s law. In: Neural Information Processing Systems, pp 407–417
2020
Cited alongside, same era.
Li L, Gan Z, Liu J (2020) A closer look at the robustness of vision-and-language pre-trained models. arXiv preprint arXiv:201208673
2020
Cited alongside, same era.
Du Y, Xu J, Xiong H, Qiu Q, Zhen X, Snoek CG, Shao L (2020) Learning to learn with variational information bottleneck for domain generalization. In: European Conference on Computer Vision, pp 200–216
2020
Cited alongside, same era.
Federici M, Dutta A, Forré P, Kushman N, Akata Z (2020) Learning robust representations via multi-view information bottleneck. In: International Conference on Learning Representations
2020
Cited alongside, same era.
Dubois Y, Kiela D, Schwab DJ, Vedantam R (2020) Learning optimal representations with the decodable information bottleneck. In: Neural Information Processing Systems, pp 18674–18690
2020
Cited alongside, same era.
Huang Z, Zeng Z, Liu B, Fu D, Fu J (2020) Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:200400849
2020
Cited alongside, same era.
Pan Z, Niu L, Zhang J, Zhang L (2021) Disentangled information bottleneck. In: Association for the Advancement of Artificial Intelligence, pp 9285–9293
2021
Later among the works it cites.
Jeon I, Lee W, Pyeon M, Kim G (2021) Ib-gan: Disengangled representation learning with information bottleneck generative adversarial networks. In: Association for the Advancement of Artificial Intelligence, pp 7926–7934
2021
Later among the works it cites.
Li C, Yan M, Xu H, Luo F, Wang W, Bi B, Huang S (2021) SemVLP: Vision-language pre-training by aligning semantics at multiple levels. arXiv preprint arXiv:210307829
2021
Later among the works it cites.
Kim W, Son B, Kim I (2021) Vilt: Vision-and-language transformer without convolution or region supervision. In: International Conference on Machine Learning, pp 5583–5594
2021
Later among the works it cites.
Sun S, Chen YC, Li L, Wang S, Fang Y, Liu J (2021) Lightningdot: Pre-training visual-semantic embeddings for real-time image-text retrieval. In: Annual Meeting of the Association for Computational Linguistics, pp 982–997
2021
Later among the works it cites.
Huang Z, Zeng Z, Huang Y, Liu B, Fu D, Fu J (2021) Seeing out of the box: End-to-end pre-training for vision-language representation learning. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 12976–12985
2021
Later among the works it cites.
Zhang P, Li X, Hu X, Yang J, Zhang L, Wang L, Choi Y, Gao J (2021) Vinvl: Revisiting visual representations in vision-language models. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 5579–5588
2021
Later among the works it cites.
Yu F, Tang J, Yin W, Sun Y, Tian H, Wu H, Wang H (2021) Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In: Association for the Advancement of Artificial Intelligence, pp 3208–3216
2021
Later among the works it cites.
Li Y, Pan Y, Yao T, Chen J, Mei T (2021) Scheduled sampling in vision-language pretraining with decoupled encoder-decoder network. In: Association for the Advancement of Artificial Intelligence, pp 8518–8526
2021
Later among the works it cites.
Changpinyo S, Sharma P, Ding N, Soricut R (2021) Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 3558–3568
2021
Later among the works it cites.
Zeng Y, Zhang X, Li H, Wang J, Zhang J, Zhou W (2022) X
2022
Closest in time.
Wang P, Yang A, Men R, Lin J, Bai S, Li Z, Ma J, Zhou C, Zhou J, Yang H (2022) OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In: International Conference on Machine Learning, pp 23318–23340
2022
Closest in time.
Yu J, Wang Z, Vasudevan V, Yeung L, Seyedhosseini M, Wu Y (2022) Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:220501917
2022
Closest in time.
Wang Z, Yu J, Yu AW, Dai Z, Tsvetkov Y, Cao Y (2022) Simvlm: Simple visual language model pretraining with weak supervision. In: International Conference on Learning Representations
2022
Closest in time.
Li J, Li D, Xiong C, Hoi S (2022) Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. arXiv preprint arXiv:220112086
2022
Closest in time.
Li C, Xu H, Tian J, Wang W, Yan M, Bi B, Ye J, Chen H, Xu G, Cao Z, et al. (2022) mplug: Effective and efficient vision-language learning by cross-modal skip-connections. In: Conference on Empirical Methods in Natural Language Processing, pp 7241–7259
2022
Closest in time.
Agrawal A, Kajić I, Bugliarello E, Davoodi E, Gergely A, Blunsom P, Nematzadeh A (2022) Rethinking evaluation practices in visual question answering: A case study on out-of-distribution generalization. arXiv preprint arXiv:220512191
2022
Closest in time.
Pan Y, Li Z, Zhang L, Tang J (2022) Causal inference with knowledge distilling and curriculum learning for unbiased vqa. ACM Transactions on Multimedia Computing, Communications, and Applications 18(3):1–23
2022
Closest in time.
Li B, Shen Y, Wang Y, Zhu W, Li D, Keutzer K, Zhao H (2022) Invariant information bottleneck for domain generalization. In: Association for the Advancement of Artificial Intelligence, pp 7399–7407
2022
Closest in time.
Wang H, Guo X, Deng ZH, Lu Y (2022) Rethinking minimal sufficient representation in contrastive learning. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 16041–16050
2022
Closest in time.
Zhou D, Yu Z, Xie E, Xiao C, Anandkumar A, Feng J, Alvarez JM (2022) Understanding the robustness in vision transformers. In: International Conference on Machine Learning, pp 27378–27394
2022
Closest in time.
Dou ZY, Xu Y, Gan Z, Wang J, Wang S, Wang L, Zhu C, Zhang P, Yuan L, Peng N, et al. (2022) An empirical study of training end-to-end vision-and-language transformers. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 18166–18176
2022
Closest in time.
Zhong Y, Yang J, Zhang P, Li C, Codella N, Li LH, Zhou L, Dai X, Yuan L, Li Y, et al. (2022) Regionclip: Region-based language-image pretraining. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 16793–16803
2022
Closest in time.
Alayrac JB, Donahue J, Luc P, Miech A, Barr I, Hasson Y, Lenc K, Mensch A, Millican K, Reynolds M, et al. (2022) Flamingo: a visual language model for few-shot learning. In: Neural Information Processing Systems, pp 23716–23736
2022
Closest in time.
Zeng Y, Zhang X, Li H (2022) Multi-grained vision language pre-training: Aligning texts with visual concepts. In: International Conference on Machine Learning, pp 25994–26009
2022
Closest in time.
Ban Y, Dong Y (2022) Pre-trained adversarial perturbations. In: Neural Information Processing Systems, pp 1196–1209
2022
Closest in time.
Wang W, Bao H, Dong L, Bjorck J, Peng Z, Liu Q, Aggarwal K, Mohammed OK, Singhal S, Som S, et al. (2023) Image as a foreign language: Beit pretraining for all vision and vision-language tasks. In: IEEE Conference on Computer Vision and Pattern Recognition
2023
Closest in time.
Xu H, Ye Q, Yan M, Shi Y, Ye J, Xu Y, Li C, Bi B, Qian Q, Wang W, et al. (2023) mplug-2: A modularized multi-modal foundation model across text, image and video. arXiv preprint arXiv:230200402
2023
Closest in time.
Li L, Lei J, Gan Z, Liu J (2021) Adversarial vqa: A new benchmark for evaluating the robustness of vqa models. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 2042–2051
2051
Closest in time.