Probing the Need for Visual Context in Multimodal Machine Translation
Caglayan, O., Madhyastha, P., Specia, L., and Barrault, L · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Original
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y., Khot, T., Agrawal, A., Summers-Stay, D., Batra, D., and Parikh, D · 2019
Later among the works it cites.
Attention on attention for image captioning
Huang, L., Wang, W., Chen, J., and Wei, X. Y · 2019
Later among the works it cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Original
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
Decoupled Weight Decay Regularization
Loshchilov, I. and Hutter, F · 2019
Later among the works it cites.
A Corpus for Reasoning About Natural Language Grounded in Photographs
Original
Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., and Artzi, Y · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Later among the works it cites.
From Recognition to Cognition: Visual Commonsense Reasoning
Original
Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Later among the works it cites.
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers
Cho, J., Lu, J., Schwenk, D., Hajishirzi, H., and Kembhavi, A · 2020
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D · 2020
Later among the works it cites.
Unified QA : Crossing Format Boundaries with a Single QA System
Khashabi, D., Min, S., Khot, T., Sabharwal, A., Tafjord, O., Clark, P., and Hajishirzi, H · 2020
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Later among the works it cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., Zettlemoyer, L., and Bart, P.-t · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
Miech, A., Alayrac, J.-B., Smaira, L., Laptev, I., Sivic, J., and Zisserman, A · 2020
Later among the works it cites.
Document Ranking with a Pretrained Sequence-to-Sequence Model
Nogueira, R., Jiang, Z., Lin, J., Mar, I. R., Pradeep, R., and Lin, J · 2020
Later among the works it cites.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., IV, R. L. L., Wallace, E., and Singh, S · 2020
Later among the works it cites.
Multimodal Transformer for Multimodal Machine Translation
Yao, S. and Wan, X · 2020
Later among the works it cites.
Actbert: Learning global-local video-text representations
Zhu, L. and Yang, Y · 2020
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Closest in time.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A., Wook, J., Chris, K., Aditya, H., Gabriel, R., Sandhini, G., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Closest in time.
VinVL: Making Visual Representations Matter in Vision-Language Models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, I · 2021
Closest in time.