Fetching the paper…
Reading the bibliography…
Large-scale pre-training has recently revolutionized vision-and-language (VL) research.
Visual entailment: A novel task for fine-grained image understanding
Xie, N.; Lai, F.; Doran, D.; and Kadav, A. 2019 · 1901
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T.; Elsen, E.; and Hooker, S. 2019 · 1902
Earlier work this paper cites.
The difficulty of training sparse neural networks
Evci, U.; Pedregosa, F.; Gomez, A.; and Elsen, E. 2019 · 1906
Earlier work this paper cites.
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
Li, G.; Duan, N.; Fang, Y.; Jiang, D.; and Zhou, M. 2019a · 1908
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019b · 1908
Earlier work this paper cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Su, W.; Zhu, X.; Cao, Y.; Li, B.; Lu, L.; Wei, F.; and Dai, J. 2019 · 1908
Earlier work this paper cites.
Unified vision-language pre-training for image captioning and vqa
Zhou, L.; Palangi, H.; Zhang, L.; Hu, H.; Corso, J. J.; and Gao, J. 2019 · 1909
Earlier work this paper cites.
12-in-1: Multi-Task Vision and Language Representation Learning
Lu, J.; Goswami, V.; Rohrbach, M.; Parikh, D.; and Lee, S. 2019b · 1912
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W.; Wei, F.; Dong, L.; Bao, H.; Yang, N.; and Zhou, M. 2020b · 2002
Earlier work this paper cites.
Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Huang, Z.; Zeng, Z.; Liu, B.; Fu, D.; and Fu, J. 2020 · 2004
Earlier work this paper cites.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Li, X.; Yin, X.; Li, C.; Hu, X.; Zhang, P.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. 2020 · 2004
Earlier work this paper cites.
When bert plays the lottery, all tickets are winning
Prasanna, S.; Rogers, A.; and Rumshisky, A. 2020 · 2005
Earlier work this paper cites.
Optimal lottery tickets via subsetsum: Logarithmic over-parameterization is sufficient
Pensia, A.; Rajput, S.; Nagle, A.; Vishwakarma, H.; and Papailiopoulos, D. 2020 · 2006
Earlier work this paper cites.
Winning Lottery Tickets in Deep Generative Models
Kalibhat, N. M.; Balaji, Y.; and Feizi, S. 2020 · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. 2011 · 2011
Earlier work this paper cites.
Chen, T.; Frankle, J.; Chang, S.; Liu, S.; Zhang, Y.; Carbin, M.; and Wang, Z. 2020a · 2012
Earlier work this paper cites.
The Lottery Ticket Hypothesis for Object Recognition
Girish, S.; Maiya, S. R.; Gupta, K.; Chen, H.; Davis, L.; and Shrivastava, A. 2020 · 2012
Earlier work this paper cites.
MiniVLM: A Smaller and Faster Vision-Language Model
Wang, J.; Hu, X.; Zhang, P.; Li, X.; Wang, L.; Zhang, L.; Gao, J.; and Liu, Z. 2020a · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Cited alongside, same era.
Han, S.; Mao, H.; and Dally, W. J. 2015 · 2015
Cited alongside, same era.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Cited alongside, same era.
Fast sparse convnets
Elsen, E.; Dukhan, M.; Gale, T.; and Simonyan, K. 2020 · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J.; Dziugaite, G. K.; Roy, D.; and Carbin, M. 2020 · 2020
Later among the works it cites.
The Early Phase of Neural Network Training
Frankle, J.; Schwab, D. J.; and Morcos, A. S. 2020 · 2020
Later among the works it cites.
Large-scale adversarial training for vision-and-language representation learning
Gan, Z.; Chen, Y.-C.; Li, L.; Zhu, C.; Cheng, Y.; and Liu, J. 2020 · 2020
Later among the works it cites.
Proving the lottery ticket hypothesis: Pruning is all you need
Malach, E.; Yehudai, G.; Shalev-Schwartz, S.; and Shamir, O. 2020 · 2020
Later among the works it cites.
Comparing rewinding and fine-tuning in neural network pruning
Renda, A.; Frankle, J.; and Carbin, M. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Madry, A.; Makelov, A.; Schmidt, L.; Tsipras, D.; and Vladu, A. 2017 · 2017
Cited alongside, same era.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Stacked cross attention for image-text matching
Lee, K.-H.; Chen, X.; Hua, G.; Hu, H.; and He, X. 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Cited alongside, same era.
A corpus for reasoning about natural language grounded in photographs
Suhr, A.; Zhou, S.; Zhang, A.; Zhang, I.; Bai, H.; and Artzi, Y. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Winning the Lottery with Continuous Sparsification
Savarese, P.; Silva, H.; and Maire, M. 2020 · 2020
Later among the works it cites.
Picking Winning Tickets Before Training by Preserving Gradient Flow
Wang, C.; Zhang, G.; and Grosse, R. 2020 · 2020
Later among the works it cites.
Drawing Early-Bird Tickets: Toward More Efficient Training of Deep Networks
You, H.; Li, C.; Xu, P.; Fu, Y.; Wang, Y.; Chen, X.; Baraniuk, R. G.; Wang, Z.; and Lin, Y. 2020 · 2020
Later among the works it cites.
Playing the lottery with rewards and multiple languages: lottery tickets in rl and nlp
Yu, H.; Edunov, S.; Tian, Y.; and Morcos, A. S. 2020 · 2020
Later among the works it cites.
An Empirical Study of Training End-to-End Vision-and-Language Transformers
Dou, Z.-Y.; Xu, Y.; Gan, Z.; Wang, J.; Wang, S.; Wang, L.; Zhu, C.; Liu, Z.; Zeng, M.; et al. 2021 · 2021
Closest in time.
Compressing Visual-linguistic Model via Knowledge Distillation
Fang, Z.; Wang, J.; Hu, X.; Wang, L.; Yang, Y.; and Liu, Z. 2021 · 2021
Closest in time.
Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
Huang, Z.; Zeng, Z.; Huang, Y.; Liu, B.; Fu, D.; and Fu, J. 2021 · 2021
Closest in time.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Closest in time.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J.; Selvaraju, R. R.; Gotmare, A. D.; Joty, S.; Xiong, C.; and Hoi, S. 2021 · 2021
Closest in time.
Good Students Play Big Lottery Better
Ma, H.; Chen, T.; Hu, T.-K.; You, C.; Xie, X.; and Wang, Z. 2021 · 2021
Closest in time.
Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
Xue, H.; Huang, Y.; Liu, B.; Peng, H.; Fu, J.; Li, H.; and Luo, J. 2021 · 2021
Closest in time.
VinVL: Making Visual Representations Matter in Vision-Language Models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Closest in time.