Fetching the paper…
Reading the bibliography…
Vision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large number of visual and language tokens.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Structured pruning of deep convolutional neural networks
Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Guiding the long-short term memory model for image caption generation
X. Jia, E. Gavves, B. Fernando, and T. Tuytelaars · 2015
Earlier work this paper cites.
Learning efficient sparse and low rank models
Pablo Sprechmann, Alexander M Bronstein, and Guillermo Sapiro · 2015
Earlier work this paper cites.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo · 2016
Earlier work this paper cites.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
On compressing deep models by low rank and sparse decomposition
Xiyu Yu, Tongliang Liu, Xinchao Wang, and Dacheng Tao · 2017
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
What is the state of neural network pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
A survey of quantization methods for efficient neural network inference
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Not all patches are what you need: Expediting vision transformers via token reorganizations
Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie · 2022
Later among the works it cites.
Adavit: Adaptive vision transformers for efficient image recognition
Lingchen Meng, Hengduo Li, Bor-Chun Chen, Shiyi Lan, Zuxuan Wu, Yu-Gang Jiang, and Ser-Nam Lim · 2022
Later among the works it cites.
Balanced multimodal learning via on-the-fly gradient modulation
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Compressing bert: Studying the effects of weight pruning on transfer learning
Mitchell A Gordon, Kevin Duh, and Nicholas Andrews · 2020
Cited alongside, same era.
Compressing visual-linguistic model via knowledge distillation
Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang, Yezhou Yang, and Zicheng Liu · 2021
Cited alongside, same era.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao · 2021
Cited alongside, same era.
Dynamic neural networks: A survey
Yizeng Han, Gao Huang, Shiji Song, Le Yang, Honghui Wang, and Yulin Wang · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, JongWook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2021
Cited alongside, same era.
Savit: Structure-aware vision transformer pruning via collaborative optimization
Chuanyang Zheng, Kai Zhang, Zhi Yang, Wenming Tan, Jun Xiao, Ye Ren, Shiliang Pu, et al · 2022
Later among the works it cites.
Muti-scale and token mergence: Make your vit more efficient
Zhe Bian, Zhe Wang, Wenqiang Han, and Kangping Wang · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, YinTat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, MarcoTulio Ribeiro, and Yi Zhang · 2023
Later among the works it cites.
Pmr: Prototypical modal rebalance for multimodal learning
Yunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang, and Song Guo · 2023
Later among the works it cites.
Elip: Efficient language-image pre-training with fewer vision tokens
Yangyang Guo, Haoyu Zhang, Liqiang Nie, Yongkang Wong, and Mohan Kankanhalli · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
Adaptive sparse vit: Towards learnable adaptive token pruning by fully exploiting self-attention
Xiangcheng Liu, Tianyi Wu, and Guodong Guo · 2023
Later among the works it cites.
You need multiple exiting: Dynamic early exiting for accelerating unified vision language model
Shengkun Tang, Yaqing Wang, Zhenglun Kong, Tianchi Zhang, Yao Li, Caiwen Ding, Yanzhi Wang, Yi Liang, and Dongkuan Xu · 2023
Later among the works it cites.
Global vision transformer pruning with hessian-aware saliency
Huanrui Yang, Hongxu Yin, Maying Shen, Pavlo Molchanov, Hai Li, and Jan Kautz · 2023
Later among the works it cites.
X-pruner: explainable pruning for vision transformers
Lu Yu and Wei Xiang · 2023
Later among the works it cites.