Fetching the paper…
Reading the bibliography…
Despite exciting progress in pre-training for visual-linguistic (VL) representations, very few aspire to a small VL model.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Model compression
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Gpt2: Empirical slant delay model for radio space geodetic techniques
K. Lagler, M. Schindelegger, J. Böhm, H. Krásná, and T. Nilsson · 2013
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Y. Kim and A. M. Rush · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Z. Huang and N. Wang · 2017
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Z. Huang and N. Wang · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
S. Zagoruyko and N. Komodakis · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
S. Ahn, S. X. Hu, A. Damianou, N. D. Lawrence, and Z. Dai · 2019
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Cited alongside, same era.
What does bert learn about the structure of language?
G. Jawahar, B. Sagot, and D. Seddah · 2019
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Z. Huang, Z. Zeng, B. Liu, D. Fu, and J. Fu · 2020
Later among the works it cites.
In defense of grid features for visual question answering
H. Jiang, I. Misra, M. Rohrbach, E. Learned-Miller, and X. Chen · 2020
Later among the works it cites.
TinyBERT: Distilling BERT for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu · 2020
Later among the works it cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
G. Li, N. Duan, Y. Fang, M. Gong, and D. Jiang · 2020
Later among the works it cites.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu · 2019
Cited alongside, same era.
Lit: Learned intermediate representation training for model compression
A. Koratana, D. Kang, P. Bailis, and M. Zaharia · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Cited alongside, same era.
A closer look at the robustness of vision-and-language pre-trained models
L. Li, Z. Gan, and J. Liu · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et al · 2020
Later among the works it cites.
12-in-1: Multi-task vision and language representation learning
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai · 2020
Later among the works it cites.
Contrastive distillation on intermediate representations for language model compression
S. Sun, Z. Gan, Y. Cheng, Y. Fang, S. Wang, and J. Liu · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou · 2020
Later among the works it cites.
Efficientdet: Scalable and efficient object detection
M. Tan, R. Pang, and Q. V. Le · 2020
Later among the works it cites.
Minivlm: A smaller and faster vision-language model
J. Wang, X. Hu, P. Zhang, X. Li, L. Wang, L. Zhang, J. Gao, and Z. Liu · 2020
Later among the works it cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le · 2020
Later among the works it cites.
Bert-of-theseus: Compressing bert by progressive module replacing
C. Xu, W. Zhou, T. Ge, F. Wei, and M. Zhou · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao · 2020
Later among the works it cites.
Actbert: Learning global-local video-text representations
L. Zhu and Y. Yang · 2020
Later among the works it cites.
Seed: Self-supervised distillation for visual representation
Z. Fang, J. Wang, L. Wang, L. Zhang, Y. Yang, and Z. Liu · 2021
Closest in time.
Vivo: Surpassing human performance in novel object captioning with visual vocabulary pre-training
X. Hu, X. Yin, K. Lin, L. Wang, L. Zhang, J. Gao, and Z. Liu · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu · 2021
Closest in time.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Closest in time.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Closest in time.
Vinvl: Making visual representations matter in vision-language models
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao · 2021
Closest in time.