Fetching the paper…
Reading the bibliography…
Encoder-decoder transformer models have achieved great success on various vision-language (VL) tasks, but they suffer from high inference latency.
Elbayad, M.; Gu, J.; Grave, E.; and Auli, M. 2019 · 1910
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2019 · 1910
Earlier work this paper cites.
Compressing bert: Studying the effects of weight pruning on transfer learning
Gordon, M. A.; Duh, K.; and Andrews, N. 2020 · 2002
Earlier work this paper cites.
Glancing Transformer for Non-Autoregressive Neural Machine Translation
Qian, L.; Zhou, H.; Bao, Y.; Wang, M.; Qiu, L.; Zhang, W.; Yu, Y.; and Li, L. 2021 · 2003
Earlier work this paper cites.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Xin, J.; Tang, R.; Lee, J.; Yu, Y.; and Lin, J. 2020 · 2004
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; Dean, J.; et al. 2015 · 2015
Earlier work this paper cites.
Deeply-supervised nets
Lee, C.-Y.; Xie, S.; Gallagher, P.; Zhang, Z.; and Tu, Z. 2015 · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J.; Huang, J.; Toshev, A.; Camburu, O.; Yuille, A. L.; and Murphy, K. 2016 · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
Teerapittayanon, S.; McDanel, B.; and Kung, H.-T. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Automatic differentiation in PyTorch
Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017 · 2017
Earlier work this paper cites.
Rosetta: Large scale system for text detection and recognition in images
Borisyuk, F.; Gordo, A.; and Sivakumar, V. 2018 · 2018
Earlier work this paper cites.
Non-Autoregressive Neural Machine Translation
Gu, J.; Bradbury, J.; Xiong, C.; Li, V. O.; and Socher, R. 2018 · 2018
Earlier work this paper cites.
Scene text visual question answering
Biten, A. F.; Tito, R.; Mafla, A.; Gomez, L.; Rusinol, M.; Valveny, E.; Jawahar, C.; and Karatzas, D. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Kenton, J. D. M.-W. C.; and Toutanova, L. K. 2019 · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Michel, P.; Levy, O.; and Neubig, G. 2019 · 2019
Cited alongside, same era.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A.; Shekhar, S.; Singh, A. K.; and Chakraborty, A. 2019 · 2019
Cited alongside, same era.
Towards vqa models that can read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Cited alongside, same era.
Uniter: Universal image-text representation learning
Chen, Y.-C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Docvqa: A dataset for vqa on document images
Mathew, M.; Karatzas, D.; and Jawahar, C. 2021 · 2021
Later among the works it cites.
Going full-tilt boogie on document understanding with text-image-layout transformer
Powalski, R.; Borchmann, Ł.; Jurkiewicz, D.; Dwojak, T.; Pietruszka, M.; and Pałka, G. 2021 · 2021
Later among the works it cites.
BERxiT: Early exiting for BERT with better fine-tuning and extension to regression
Xin, J.; Tang, R.; Yu, Y.; and Lin, J. 2021 · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2020
Cited alongside, same era.
TinyBERT: Distilling BERT for Natural Language Understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2020 · 2020
Cited alongside, same era.
FastBERT: a Self-distilling BERT with Adaptive Inference Time
Liu, W.; Zhou, P.; Wang, Z.; Zhao, Z.; Deng, H.; and Ju, Q. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P. J.; et al. 2020 · 2020
Cited alongside, same era.
The Right Tool for the Job: Matching Model and Instance Complexities
Schwartz, R.; Stanovsky, G.; Swayamdipta, S.; Dodge, J.; and Smith, N. A. 2020 · 2020
Cited alongside, same era.
Docformer: End-to-end transformer for document understanding
Appalaraju, S.; Jasani, B.; Kota, B. U.; Xie, Y.; and Manmatha, R. 2021 · 2021
Cited alongside, same era.
Latr: Layout-aware transformer for scene-text vqa
Biten, A. F.; Litman, R.; Xie, Y.; Appalaraju, S.; and Manmatha, R. 2022 · 2022
Later among the works it cites.
Pali: A jointly-scaled multilingual language-image model
Chen, X.; Wang, X.; Changpinyo, S.; Piergiovanni, A.; Padlewski, P.; Salz, D.; Goodman, S.; Grycner, A.; Mustafa, B.; Beyer, L.; et al. 2022 · 2022
Later among the works it cites.
Knowledge Distillation via the Target-aware Transformer
Lin, S.; Xie, H.; Wang, B.; Yu, K.; Chang, X.; Liang, X.; and Wang, G. 2022 · 2022
Later among the works it cites.
Unified-io: A unified model for vision, language, and multi-modal tasks
Lu, J.; Clark, C.; Zellers, R.; Mottaghi, R.; and Kembhavi, A. 2022 · 2022
Later among the works it cites.
Confident adaptive language modeling
Schuster, T.; Fisch, A.; Gupta, J.; Dehghani, M.; Bahri, D.; Tran, V.; Tay, Y.; and Metzler, D. 2022 · 2022
Later among the works it cites.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Wang, P.; Yang, A.; Men, R.; Lin, J.; Bai, S.; Li, Z.; Ma, J.; Zhou, C.; Zhou, J.; and Yang, H. 2022 · 2022
Later among the works it cites.
PCEE-BERT: Accelerating BERT Inference via Patient and Confident Early Exiting
Zhang, Z.; Zhu, W.; Zhang, J.; Wang, P.; Jin, R.; and Chung, T.-S. 2022 · 2022
Later among the works it cites.
Bert loses patience: Fast and robust inference with early exit
Zhou, W.; Xu, C.; Ge, T.; McAuley, J.; Xu, K.; and Wei, F. 2020b · 2022
Later among the works it cites.
DocFormerv2: Local Features for Document Understanding
Appalaraju, S.; Tang, P.; Dong, Q.; Sankaran, N.; Zhou, Y.; and Manmatha, R. 2023 · 2023
Closest in time.
A global past-future early exit method for accelerating inference of pre-trained language models
Liao, K.; Zhang, Y.; Ren, X.; Su, Q.; Sun, X.; and He, B. 2021 · 2023
Closest in time.