Fetching the paper…
Reading the bibliography…
Transformer architectures have brought about fundamental changes to computational linguistic field, which had been dominated by recurrent neural networks for many years.
1904
Earlier work this paper cites.
Yang J, Ren Z, Xu M, Chen X, Crandall D, Parikh D, Batra D (2019a) Embodied visual recognition
1904
Earlier work this paper cites.
1907
Earlier work this paper cites.
Zhou L, Palangi H, Zhang L, Hu H, Corso JJ, Gao J (2019) Unified vision-language pre-training for image captioning and vqa
1909
Earlier work this paper cites.
Sanh V, Debut L, Chaumond J, Wolf T (2020) Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
1910
Earlier work this paper cites.
Tan M, Pang R, Le QV (2020) Efficientdet: Scalable and efficient object detection
1911
Earlier work this paper cites.
Kervadec C, Antipov G, Baccouche M, Wolf C (2019) Weak supervision helps emergence of word-object alignment and improves vision-language tasks
1912
Earlier work this paper cites.
Elman JL (1990) Finding structure in time. COGNITIVE SCIENCE 14(2):179–211
1990
Earlier work this paper cites.
Miller GA (1995) Wordnet: A lexical database for english. COMMUNICATIONS OF THE ACM 38:39–41
1995
Earlier work this paper cites.
Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Computation 9(8):1735–1780
1997
Earlier work this paper cites.
LeCun Y, Bottou L, Bengio Y, Haffner P (1998) Gradient-based learning applied to document recognition. In: Proceedings of the IEEE, vol 86, pp 2278–2324, URL
1998
Earlier work this paper cites.
Qi D, Su L, Song J, Cui E, Bharti T, Sacheti A (2020) Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
2001
Earlier work this paper cites.
Luo H, Ji L, Shi B, Huang H, Duan N, Li T, Li J, Bharti T, Zhou M (2020b) Univl: A unified video and language pre-training model for multimodal understanding and generation
2002
Earlier work this paper cites.
Papineni K, Roukos S, Ward T, Zhu WJ (2002) Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Philadelphia, Pennsylvania, USA, pp 311–318, DOI
2002
Earlier work this paper cites.
Lin J, Yang A, Zhang Y, Liu J, Zhou J, Yang H (2021) M6-v0: Vision-and-language interaction for multi-modal pretraining
2003
Earlier work this paper cites.
Huang Z, Zeng Z, Liu B, Fu D, Fu J (2020b) Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
2004
Earlier work this paper cites.
Li X, Yin X, Li C, Zhang P, Hu X, Zhang L, Wang L, Hu H, Dong L, Wei F, Choi Y, Gao J (2020c) Oscar: Object-semantics aligned pre-training for vision-language tasks
2004
Earlier work this paper cites.
Lin CY (2004) ROUGE: A package for automatic evaluation of summaries. In: Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, pp 74–81, URL
2004
Earlier work this paper cites.
Sharir O, Peleg B, Shoham Y (2020) The cost of training nlp models: A concise overview
2004
Earlier work this paper cites.
Banerjee S, Lavie A (2005) METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In: Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Association for Computational Linguistics, Ann Arbor, Michigan, pp 65–72, URL
2005
Earlier work this paper cites.
Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S (2020) End-to-end object detection with transformers
2005
Earlier work this paper cites.
Li L, Chen YC, Cheng Y, Gan Z, Yu L, Liu J (2020b) Hero: Hierarchical encoder for video+language omni-representation pre-training
2005
Earlier work this paper cites.
Korbar B, Petroni F, Girdhar R, Torresani L (2020) Video understanding as machine translation
2006
Earlier work this paper cites.
Yu F, Tang J, Yin W, Sun Y, Tian H, Wu H, Wang H (2020) Ernie-vil: Knowledge enhanced vision-language representations through scene graph
2006
Earlier work this paper cites.
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, Uszkoreit J, Houlsby N (2020) An image is worth 16x16 words: Transformers for image recognition at scale
2010
Earlier work this paper cites.
Farhadi A, Hejrati M, Sadeghi M, Young P, Rashtchian C, Hockenmaier J, Forsyth D (2010) Every picture tells a story: Generating sentences from images. In: Computer Vision, ECCV 2010 - 11th European Conference on Computer Vision, Proceedings, Springer-Verlag Berlin Heidelberg, no. PART 4 in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), pp 15–29, DOI
2010
Earlier work this paper cites.
Luo F, Yang P, Li S, Ren X, Sun X (2020a) Capt: Contrastive pre-training for learning denoised sequence representations
2010
Earlier work this paper cites.
Ging S, Zolfaghari M, Pirsiavash H, Brox T (2020) Coot: Cooperative hierarchical transformer for video-text representation learning
2011
Earlier work this paper cites.
Ordonez V, Kulkarni G, Berg T (2011) Im2text: Describing images using 1 million captioned photographs. In: Shawe-Taylor J, Zemel R, Bartlett P, Pereira F, Weinberger KQ (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 24, pp 1143–1151, URL
2011
Earlier work this paper cites.
Barbu A, Bridge A, Burchill Z, Coroian D, Dickinson S, Fidler S, Michaux A, Mussman S, Narayanaswamy S, Salvi D, Schmidt L, Shangguan J, Siskind JM, Waggoner J, Wang S, Wei J, Yin Y, Zhang Z (2012) Video in sentences out
2012
Earlier work this paper cites.
Chen H, Wang Y, Guo T, Xu C, Deng Y, Liu Z, Ma S, Xu C, Xu C, Gao W (2020a) Pre-trained image processing transformer
2012
Earlier work this paper cites.
Guo J, Zhu C, Zhao Y, Wang H, Hu Y, He X, Cai D (2020) Lamp: Label augmented multimodal pretraining
2012
Earlier work this paper cites.
Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, Tang Y, Xiao A, Xu C, Xu Y, Yang Z, Zhang Y, Tao D (2021) A survey on visual transformer
2012
Earlier work this paper cites.
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: Pereira F, Burges CJC, Bottou L, Weinberger KQ (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 25, pp 1097–1105, URL
2012
Earlier work this paper cites.
Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, Jégou H (2020) Training data-efficient image transformers & distillation through attention
2012
Earlier work this paper cites.
Ushiku Y, Harada T, Kuniyoshi Y (2012) Efficient image annotation for automatic sentence generation. In: ACM Multimedia
2012
Earlier work this paper cites.
Wang H, Zhu Y, Adam H, Yuille A, Chen LC (2020a) Max-deeplab: End-to-end panoptic segmentation with mask transformers
2012
Earlier work this paper cites.
Wang J, Hu X, Zhang P, Li X, Wang L, Zhang L, Gao J, Liu Z (2020b) Minivlm: A smaller and faster vision-language model
2012
Earlier work this paper cites.
Das P, Xu C, Doell RF, Corso JJ (2013) A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. 2013 IEEE Conference on Computer Vision and Pattern Recognition pp 2634–2641
2013
Earlier work this paper cites.
Elliott D, Keller F (2013) Image description using visual dependency representations. In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Seattle, Washington, USA, pp 1292–1302, URL
2013
Earlier work this paper cites.
Hodosh M, Young P, Hockenmaier J (2013) Framing image description as a ranking task: Data, models and evaluation metrics. J Artif Intell Res 47:853–899
2013
Earlier work this paper cites.
Mikolov T, Sutskever I, Chen K, Corrado GS, Dean J (2013) Distributed representations of words and phrases and their compositionality. In: Burges CJC, Bottou L, Welling M, Ghahramani Z, Weinberger KQ (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 26, pp 3111–3119, URL
2013
Earlier work this paper cites.
Cho K, van Merriënboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, Bengio Y (2014) Learning phrase representations using RNN encoder–decoder for statistical machine translation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Doha, Qatar, pp 1724–1734, DOI
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets. In: Ghahramani Z, Welling M, Cortes C, Lawrence N, Weinberger KQ (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 27, pp 2672–2680, URL
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Karpathy A, Toderici G, Shetty S, Leung T, Sukthankar R, Fei-Fei L (2014) Large-scale video classification with convolutional neural networks. In: CVPR
2014
Earlier work this paper cites.
Kazemzadeh S, Ordonez V, Matten M, Berg T (2014) ReferItGame: Referring to objects in photographs of natural scenes. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Doha, Qatar, pp 787–798, DOI
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Pennington J, Socher R, Manning C (2014) GloVe: Global vectors for word representation. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Doha, Qatar, pp 1532–1543, DOI
Peters M, Neumann M, Iyyer M, Gardner M, Clark C, Lee K, Zettlemoyer L (2018) Deep contextualized word representations. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), Association for Computational Linguistics, New Orleans, Louisiana, pp 2227–2237, DOI
2018
Later among the works it cites.
Radford A, Sutskever I (2018) Improving language understanding by generative pre-training. In: arxiv
2018
Later among the works it cites.
Sharma P, Ding N, Goodman S, Soricut R (2018) Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, Melbourne, Australia, pp 2556–2565, DOI
2018
Later among the works it cites.
Vondrick C, Shrivastava A, Fathi A, Guadarrama S, Murphy K (2018) Tracking emerges by colorizing videos. In: Proceedings of the European Conference on Computer Vision (ECCV)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
Rezende DJ, Mohamed S, Wierstra D (2014) Stochastic backpropagation and approximate inference in deep generative models. In: Xing EP, Jebara T (eds) Proceedings of the 31st International Conference on Machine Learning, PMLR, Bejing, China, Proceedings of Machine Learning Research, vol 32, pp 1278–1286, URL
2014
Cited alongside, same era.
Young P, Lai A, Hodosh M, Hockenmaier J (2014) From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics 2:67–78, DOI
2014
Cited alongside, same era.
Agrawal P, Carreira J, Malik J (2015) Learning to see by moving. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
Cited alongside, same era.
Antol S, Agrawal A, Lu J, Mitchell M, Batra D, Zitnick CL, Parikh D (2015) Vqa: Visual question answering. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
Cited alongside, same era.
Girshick R (2015) Fast r-cnn. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
Cited alongside, same era.
Heilbron FC, Escorcia V, Ghanem B, Niebles JC (2015) Activitynet: A large-scale video benchmark for human activity understanding. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 961–970, DOI
2015
Cited alongside, same era.
Kiros R, Zhu Y, Salakhutdinov R, Zemel RS, Torralba A, Urtasun R, Fidler S (2015) Skip-thought vectors
2015
Cited alongside, same era.
Ren S, He K, Girshick R, Sun J (2015) Faster r-cnn: Towards real-time object detection with region proposal networks. In: Cortes C, Lawrence N, Lee D, Sugiyama M, Garnett R (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 28, pp 91–99, URL
2015
Cited alongside, same era.
2018
Later among the works it cites.
Xie S, Sun C, Huang J, Tu Z, Murphy K (2018) Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
2018
Later among the works it cites.
Zhou L, Zhou Y, Corso JJ, Socher R, Xiong C (2018) End-to-end dense video captioning with masked transformer. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
Later among the works it cites.
Alberti C, Ling J, Collins M, Reitter D (2019) Fusion of detected objects in text for visual question answering. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Association for Computational Linguistics, Hong Kong, China, pp 2131–2140, DOI
2019
Later among the works it cites.
Dai Z, Yang Z, Yang Y, Carbonell J, Le Q, Salakhutdinov R (2019) Transformer-XL: Attentive language models beyond a fixed-length context. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Florence, Italy, pp 2978–2988, DOI
2019
Later among the works it cites.
Devlin J, Chang MW, Lee K, Toutanova K (2019) BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Association for Computational Linguistics, Minneapolis, Minnesota, pp 4171–4186, DOI
2019
Later among the works it cites.
Hudson DA, Manning CD (2019) Gqa: A new dataset for real-world visual reasoning and compositional question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2019
Later among the works it cites.
Karras T, Laine S, Aila T (2019) A style-based generator architecture for generative adversarial networks. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 4396–4405, DOI
2019
Later among the works it cites.
Li LH, Yatskar M, Yin D, Hsieh CJ, Chang KW (2019) Visualbert: A simple and performant baseline for vision and language. In: Arxiv
2019
Later among the works it cites.
Lu J, Batra D, Parikh D, Lee S (2019) Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In: Wallach H, Larochelle H, Beygelzimer A, d'Alché-Buc F, Fox E, Garnett R (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 32, pp 13–23, URL
2019
Later among the works it cites.
Miech A, Zhukov D, Alayrac JB, Tapaswi M, Laptev I, Sivic J (2019) HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. In: ICCV
2019
Later among the works it cites.
Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I (2019) Language Models are Unsupervised Multitask Learners URL
2019
Later among the works it cites.
Shao S, Li Z, Zhang T, Peng C, Yu G, Zhang X, Li J, Sun J (2019) Objects365: A large-scale, high-quality dataset for object detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2019
Later among the works it cites.
Suhr A, Zhou S, Zhang A, Zhang I, Bai H, Artzi Y (2019) A corpus for reasoning about natural language grounded in photographs. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Florence, Italy, pp 6418–6428, DOI
2019
Later among the works it cites.
Sun C, Myers A, Vondrick C, Murphy K, Schmid C (2019) Videobert: A joint model for video and language representation learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2019
Later among the works it cites.
Tan H, Bansal M (2019) LXMERT: Learning cross-modality encoder representations from transformers. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Association for Computational Linguistics, Hong Kong, China, pp 5100–5111, DOI
2019
Later among the works it cites.
Tan M, Le Q (2019) EfficientNet: Rethinking model scaling for convolutional neural networks. In: Chaudhuri K, Salakhutdinov R (eds) Proceedings of the 36th International Conference on Machine Learning, PMLR, Proceedings of Machine Learning Research, vol 97, pp 6105–6114, URL
2019
Later among the works it cites.
Yang Z, Dai Z, Yang Y, Carbonell J, Salakhutdinov RR, Le QV (2019b) Xlnet: Generalized autoregressive pretraining for language understanding. In: Wallach H, Larochelle H, Beygelzimer A, d'Alché-Buc F, Fox E, Garnett R (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 32, pp 5753–5763, URL
2019
Later among the works it cites.
Zellers R, Bisk Y, Farhadi A, Choi Y (2019) From recognition to cognition: Visual commonsense reasoning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2019
Later among the works it cites.
2019
Later among the works it cites.
Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, Agarwal S, Herbert-Voss A, Krueger G, Henighan T, Child R, Ramesh A, Ziegler D, Wu J, Winter C, Hesse C, Chen M, Sigler E, Litwin M, Gray S, Chess B, Clark J, Berner C, McCandlish S, Radford A, Sutskever I, Amodei D (2020) Language models are few-shot learners. In: Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 33, pp 1877–1901, URL
2020
Later among the works it cites.
Chang WC, Yu FX, Chang YW, Yang Y, Kumar S (2020) Pre-training tasks for embedding-based large-scale retrieval. In: International Conference on Learning Representations, URL
2020
Later among the works it cites.
Gabeur V, Sun C, Alahari K, Schmid C (2020) Multi-modal Transformer for Video Retrieval. In: European Conference on Computer Vision (ECCV)
2020
Later among the works it cites.
Huang G, Pang B, Zhu Z, Rivera C, Soricut R (2020a) Multimodal pretraining for dense video captioning. In: Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Association for Computational Linguistics, Suzhou, China, pp 470–490, URL
2020
Later among the works it cites.
Jiao X, Yin Y, Shang L, Jiang X, Chen X, Li L, Wang F, Liu Q (2020) Tiny{bert}: Distilling {bert} for natural language understanding. URL
2020
Later among the works it cites.
Karras T, Laine S, Aittala M, Hellsten J, Lehtinen J, Aila T (2020) Analyzing and improving the image quality of stylegan. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2020
Later among the works it cites.
Kitaev N, Kaiser L, Levskaya A (2020) Reformer: The efficient transformer. In: International Conference on Learning Representations, URL
2020
Later among the works it cites.
Su W, Zhu X, Cao Y, Li B, Lu L, Wei F, Dai J (2020) Vl-bert: Pre-training of generic visual-linguistic representations. In: International Conference on Learning Representations, URL
2020
Later among the works it cites.
Sun C, Baradel F, Murphy K, Schmid C (2020) Learning video representations using contrastive bidirectional transformer. URL
2020
Later among the works it cites.
Wang Y, Mohamed A, Le D, Liu C, Xiao A, Mahadeokar J, Huang H, Tjandra A, Zhang X, Zhang F, et al (2020c) Transformer-based acoustic modeling for hybrid speech recognition. ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) DOI
2020
Later among the works it cites.
Zhang S, Jiang T, Wang T, Kuang K, Zhao Z, Zhu J, Yu J, Yang H, Wu F (2020) Devlbert. Proceedings of the 28th ACM International Conference on Multimedia DOI
2020
Later among the works it cites.
Zhu L, Yang Y (2020) Actbert: Learning global-local video-text representations. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2020
Later among the works it cites.
2021
Closest in time.
2021
Closest in time.
Fedus W, Zoph B, Shazeer N (2021) Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
2021
Closest in time.
2021
Closest in time.
Khan S, Naseer M, Hayat M, Zamir SW, Khan FS, Shah M (2021) Transformers in vision: A survey
2021
Closest in time.
Kim W, Son B, Kim I (2021) Vilt: Vision-and-language transformer without convolution or region supervision
2021
Closest in time.
Lei J, Li L, Zhou L, Gan Z, Berg TL, Bansal M, Liu J (2021) Less is more: Clipbert for video-and-language learning via sparse sampling
2021
Closest in time.
2021
Closest in time.
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I (2021) Learning transferable visual models from natural language supervision
2021
Closest in time.
Ramesh A, Pavlov M, Goh G, Gray S, Voss C, Radford A, Chen M, Sutskever I (2021) Zero-shot text-to-image generation
2021
Closest in time.
Xu K, Ba J, Kiros R, Cho K, Courville A, Salakhudinov R, Zemel R, Bengio Y (2015) Show, attend and tell: Neural image caption generation with visual attention. In: Bach F, Blei D (eds) Proceedings of the 32nd International Conference on Machine Learning, PMLR, Lille, France, Proceedings of Machine Learning Research, vol 37, pp 2048–2057, URL
2057
Closest in time.