Fetching the paper…
Reading the bibliography…
Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., G. Kulkarni, T. L. Berg · 2011
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., M. Maire, S. J. Belongie, et al · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., V. Ordonez, M. Matten, et al · 2014
Earlier work this paper cites.
VQA: visual question answering
Antol, S., A. Agrawal, J. Lu, et al · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., O. Vinyals, J. Dean · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., L. Wang, C. M. Cervantes, et al · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A., F. Li · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., G. Angeli, C. Potts, et al · 2015
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., P. Poirson, S. Yang, et al · 2016
Earlier work this paper cites.
Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
Das, A., H. Agrawal, C. L. Zitnick, et al · 2016
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., M. Cogswell, A. Das, et al · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Zagoruyko, S., N. Komodakis · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Tarvainen, A., H. Valpola · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., N. Shazeer, N. Parmar, et al · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Y. Zhu, O. Groth, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I., F. Hutter · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., T. Khot, D. Summers-Stay, et al · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., N. Ding, S. Goodman, et al · 2018
Earlier work this paper cites.
VSE++: improving visual-semantic embeddings with hard negatives
Faghri, F., D. J. Fleet, J. R. Kiros, et al · 2018
Earlier work this paper cites.
Born-again neural networks
Furlanello, T., Z. C. Lipton, M. Tschannen, et al · 2018
Earlier work this paper cites.
Deep mutual learning
Zhang, Y., T. Xiang, T. M. Hospedales, et al · 2018
Earlier work this paper cites.
Large scale distributed neural network training through online distillation
Anil, R., G. Pereyra, A. Passos, et al · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Y. Li, O. Vinyals · 2018
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension
Yu, L., Z. Lin, X. Shen, et al · 2018
Cited alongside, same era.
Bilinear attention networks
Kim, J., J. Jun, B. Zhang · 2018
Cited alongside, same era.
LXMERT: learning cross-modality encoder representations from transformers
Tan, H., M. Bansal · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Yu, F., J. Tang, W. Yin, et al · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., H. Fan, Y. Wu, et al · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Chen, T., S. Kornblith, M. Norouzi, et al · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., M. Cord, M. Douze, et al · 2020
Later among the works it cites.
Dividemix: Learning with noisy labels as semi-supervised learning
Li, J., R. Socher, S. C. Hoi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lu, J., D. Batra, D. Parikh, et al · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., M. Yatskar, D. Yin, et al · 2019
Cited alongside, same era.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., S. Zhou, A. Zhang, et al · 2019
Cited alongside, same era.
Visual semantic reasoning for image-text matching
Li, K., Y. Zhang, K. Li, et al · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., L. Debut, J. Chaumond, et al · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., M. Chang, K. Lee, et al · 2019
Cited alongside, same era.
Visual entailment: A novel task for fine-grained image understanding
Xie, N., F. Lai, D. Doran, et al · 2019
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., B. Zoph, J. Shlens, et al · 2020
Later among the works it cites.
What makes for good views for contrastive learning?
Tian, Y., C. Sun, B. Poole, et al · 2020
Later among the works it cites.
A mutual information maximization perspective of language representation learning
Kong, L., C. de Masson d’Autume, L. Yu, et al · 2020
Later among the works it cites.
e-snli-ve: Corrected visual-textual entailment with natural language explanations
Do, V., O.-M. Camburu, Z. Akata, et al · 2020
Later among the works it cites.
Counterfactual contrastive learning fo weakly-supervised vision-language grounding
Zhang, Z., Z. Zhao, Z. Lin, et al · 2020
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., J. W. Kim, C. Hallacy, et al · 2021
Closest in time.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Y. Yang, Y. Xia, et al · 2021
Closest in time.
Vinvl: Making visual representations matter in vision-language models
Zhang, P., X. Li, X. Hu, et al · 2021
Closest in time.
Seeing out of the box: End-to-end pre-training for vision-language representation learning
Huang, Z., Z. Zeng, Y. Huang, et al · 2021
Closest in time.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W., B. Son, I. Kim · 2021
Closest in time.
Prototypical contrastive learning of unsupervised representations
Li, J., P. Zhou, C. Xiong, et al · 2021
Closest in time.
Mopro: Webly supervised learning with momentum prototypes
Li, J., C. Xiong, S. C. Hoi · 2021
Closest in time.
Data-efficient language-supervised zero-shot learning with self-distillation
Cheng, R., B. Wu, P. Zhang, et al · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., L. Beyer, A. Kolesnikov, et al · 2021
Closest in time.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., P. Sharma, N. Ding, et al · 2021
Closest in time.
Unifying vision-and-language tasks via text generation
Cho, J., J. Lei, H. Tan, et al · 2021
Closest in time.