Fetching the paper…
Reading the bibliography…
Real-world data contains a vast amount of multimodal information, among which vision and language are the two most representative modalities.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Karpathy, A., Joulin, A., and Fei-Fei, L. F · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Kiros, R., Salakhutdinov, R., and Zemel, R. S · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Guiding long-short term memory for image caption generation, 2015
Jia, X., Gavves, E., Fernando, B., and Tuytelaars, T · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering
Yang, Z., He, X., Gao, J., Deng, L., and Smola, A · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Instance-aware image and sentence matching with selective multimodal lstm
Huang, Y., Wang, W., and Wang, L · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
On compressing deep models by low rank and sparse decomposition
Yu, X., Liu, T., Wang, X., and Tao, D · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A · 2017
Earlier work this paper cites.
Encoder-decoder with atrous separable convolution for semantic image segmentation
Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., and Artzi, Y · 2018
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Cited alongside, same era.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Rao, Y., Zhao, W., Liu, B., Lu, J., Zhou, J., and Hsieh, C.-J · 2021
Later among the works it cites.
Intern: A new learning paradigm towards general vision
Shao, J., Chen, S., Li, Y., Wang, K., Yin, Z., He, Y., Teng, J., Sun, Q., Gao, M., Liu, J., et al · 2021
Later among the works it cites.
Multi-encoder parse-decoder network for sequential medical image segmentation
Shi, D., Liu, R., Tao, L., He, Z., and Huo, L · 2021
Later among the works it cites.
Segmenter: Transformer for semantic segmentation
Strudel, R., Garcia, R., Laptev, I., and Schmid, C · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The lottery ticket hypothesis for pre-trained bert networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Cited alongside, same era.
Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark, 2020
Contributors, M · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Power-bert: Accelerating bert inference via progressive word-vector elimination
Goyal, S., Choudhury, A. R., Raje, S., Chakaravarthy, V., Sabharwal, Y., and Verma, A · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., et al · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Cited alongside, same era.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L · 2021
Later among the works it cites.
Segformer: Simple and efficient design for semantic segmentation with transformers
Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., and Luo, P · 2021
Later among the works it cites.
Nvit: Vision transformer compression and parameter redistribution
Yang, H., Yin, H., Molchanov, P., Li, H., and Kautz, J · 2021
Later among the works it cites.
Zhu, M., Tang, Y., and Han, K · 2021
Later among the works it cites.
Vision transformer slimming: Multi-dimension searching in continuous optimization space
Chavan, A., Shen, Z., Liu, Z., Liu, Z., Cheng, K.-T., and Xing, E. P · 2022
Later among the works it cites.
Litevl: Efficient video-language learning with enhanced spatial-temporal modeling
Chen, D., Tao, C., Hou, L., Shang, L., Jiang, X., and Liu, Q · 2022
Later among the works it cites.
Playing lottery tickets with vision and language
Gan, Z., Chen, Y.-C., Li, L., Chen, T., Cheng, Y., Wang, S., Liu, J., Wang, L., and Liu, Z · 2022
Later among the works it cites.
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Later among the works it cites.
Not all patches are what you need: Expediting vision transformers via token reorganizations
Liang, Y., Ge, C., Tong, Z., Song, Y., Wang, J., and Xie, P · 2022
Later among the works it cites.
Heuristic dropout: An efficient regularization method for medical image segmentation models
Shi, D., Liu, R., Tao, L., and Yuan, C · 2022
Later among the works it cites.
Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al · 2022
Later among the works it cites.
Vitas: vision transformer architecture search
Su, X., You, S., Xie, J., Zheng, M., Wang, F., Qian, C., Zhang, C., Wang, X., and Xu, C · 2022
Later among the works it cites.
Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks
Sung, Y.-L., Cho, J., and Bansal, M · 2022
Later among the works it cites.
Compression of generative pre-trained language models via quantization
Tao, C., Hou, L., Zhang, W., Shang, L., Jiang, X., Liu, Q., Luo, P., and Wong, N · 2022
Later among the works it cites.
Deit iii: Revenge of the vit
Touvron, H., Cord, M., and Jégou, H · 2022
Later among the works it cites.
Masked generative distillation
Yang, Z., Li, Z., Shao, M., Shi, D., Yuan, Z., and Yuan, C · 2022
Later among the works it cites.
A-vit: Adaptive tokens for efficient vision transformer
Yin, H., Vahdat, A., Alvarez, J. M., Mallya, A., Kautz, J., and Molchanov, P · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y · 2022
Later among the works it cites.