Fetching the paper…
Reading the bibliography…
The multimedia community has shown a significant interest in perceiving and representing the physical world with multimodal pretrained neural network models, and among them, the visual-language pertaining (VLP) is, currently, the most captivating topic.
R. S. Jackendoff, Semantic structures . MIT press, 1992, vol. 18
1992
Earlier work this paper cites.
H. Zeijlstra, “Negation in natural language: On the form and meaning of negative elements,” Language and Linguistics Compass , vol. 1, no. 5, pp. 498–518, 2007
2007
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in Indian Conference on Computer Vision, Graphics and Image Processing , Dec 2008
2008
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in International Conference on Machine Learning , 2009
2009
Earlier work this paper cites.
M. J. Choi, A. Torralba, and A. S. Willsky, “Context models and out-of-context objects,” Pattern Recognition Letters , 2012
2012
Earlier work this paper cites.
I. Orenes, D. Beltrán, and C. Santamaría, “How negation is understood: Evidence from the visual world paradigm,” Journal of memory and language , vol. 74, pp. 36–45, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference of Computer Vision . Springer, 2014
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Transactions of the Association for Computational Linguistics , 2014
2014
Earlier work this paper cites.
M. J. Cresswell, Logics and languages . Routledge, 2016
2016
Earlier work this paper cites.
R. Shekhar, S. Pezzelle, Y. Klimovich, A. Herbelot, M. Nabi, E. Sangineto, and R. Bernardi, “Foil it! find one mismatch between image and language caption,” in ACL 2017 The 55th Annual Meeting of the Association for Computational Linguistics: Proceedings of the Conference, Vol. 1 (Long Papers) . Association for Computational Linguistics, 2017, pp. 255–265
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International journal of computer vision , 2017
2017
Earlier work this paper cites.
M. Honnibal and I. Montani, “spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing,” To appear , vol. 7, no. 1, pp. 411–420, 2017
2017
Earlier work this paper cites.
A. Conneau, G. Kruszewski, G. Lample, L. Barrault, and M. Baroni, “What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Melbourne, Australia: Association for Computational Linguistics, Jul. 2018, pp. 2126–2136
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Wang, A. Singh et al. , “Glue: A multi-task benchmark and analysis platform for natural language understanding,” in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP , 2018
2018
Earlier work this paper cites.
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang, “Visualbert: A simple and performant baseline for vision and language,” arXiv preprint , 2019
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” Advances in neural information processing systems , 2019
2019
Earlier work this paper cites.
O. Caglayan, P. S. Madhyastha, L. Specia, and L. Barrault, “Probing the need for visual context in multimodal machine translation,” in North American Chapter of the Association for Computational Linguistics , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6700–6709
2019
Earlier work this paper cites.
A. Fyshe, G. Sudre, L. Wehbe, N. Rafidi, and T. M. Mitchell, “The lexical semantics of adjective–noun phrases in the human brain,” Human brain mapping , vol. 40, no. 15, pp. 4457–4469, 2019
2019
Earlier work this paper cites.
J. Cao, Z. Gan, Y. Cheng, L. Yu, Y.-C. Chen, and J. Liu, “Behind the scene: Revealing the secrets of pre-trained vision-and-language models,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16 , 2020
2020
Earlier work this paper cites.
S. Pratt, M. Yatskar, L. Weihs, A. Farhadi, and A. Kembhavi, “Grounded situation recognition,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16 . Springer, 2020, pp. 314–332
2020
Earlier work this paper cites.
L. Ding, L. Wang, D. Wu, D. Tao, and Z. Tu, “Context-aware cross-attention for non-autoregressive translation,” in Proceedings of the 28th International Conference on Computational Linguistics , 2020, pp. 4396–4402
2020
Earlier work this paper cites.
A. Ettinger, “What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models,” Transactions of the Association for Computational Linguistics , 2020
2020
Earlier work this paper cites.
A. Warstadt, A. Parrish, H. Liu, A. Mohananey, W. Peng, S.-F. Wang, and S. R. Bowman, “Blimp: The benchmark of linguistic minimal pairs for english,” Transactions of the Association for Computational Linguistics , 2020
2020
Earlier work this paper cites.
L. Ding, L. Wang, and D. Tao, “Self-attention with cross-lingual position representation,” in Annual Meeting of the Association for Computational Linguistics , 2020
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning , 2021
2021
Cited alongside, same era.
E. Bugliarello, R. Cotterell, N. Okazaki, and D. Elliott, “Multimodal pretraining unmasked: A meta-analysis and a unified framework of vision-and-language berts,” Transactions of the Association for Computational Linguistics , 2021
2021
Cited alongside, same era.
T. Iki and A. Aizawa, “Effect of visual extensions on natural language understanding in vision-and-language models,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021
2021
Cited alongside, same era.
P. J. Rösch and J. Libovickỳ, “Probing the role of positional information in vision-language models,” in Findings of the Association for Computational Linguistics: NAACL 2022 , 2022, pp. 1031–1041
2022
Later among the works it cites.
T. Thrush, R. Jiang et al. , “Winoground: Probing vision and language models for visio-linguistic compositionality,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022
2022
Later among the works it cites.
X. Liu, D. Yin, Y. Feng, and D. Zhao, “Things not written in text: Exploring spatial commonsense from visual signals,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 2365–2376
2022
Later among the works it cites.
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela, “Flava: A foundational language and vision alignment model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 638–15 650
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
L. A. Hendricks and A. Nematzadeh, “Probing image-language transformers for verb understanding,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 3635–3644
2021
Cited alongside, same era.
A. Radford, J. W. Kim et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
K. Pham, K. Kafle, Z. Lin, Z. Ding, S. Cohen, Q. Tran, and A. Shrivastava, “Learning to predict visual attributes in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 13 018–13 028
2021
Cited alongside, same era.
L. Ding, L. Wang, X. Liu, D. F. Wong, D. Tao, and Z. Tu, “Understanding and improving lexical choice in non-autoregressive translation,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
T. Pham, T. Bui, L. Mai, and A. Nguyen, “Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks?” in Findings of ACL , 2021
2021
Cited alongside, same era.
K. Sinha, R. Jia, D. Hupkes, J. Pineau, A. Williams, and D. Kiela, “Masked language modeling and the distributional hypothesis: Order word matters pre-training for little,” in Empirical Methods in Natural Language Processing , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou, “When and why vision-language models behave like bags-of-words, and what to do about it?” in The Eleventh International Conference on Learning Representations , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Rao, L. Ding, S. Qi, M. Fang, Y. Liu, L. Shen, and D. Tao, “Dynamic contrastive distillation for image-text retrieval,” IEEE Transactions on Multimedia , 2023
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” 2023
2023
Closest in time.
H. Touvron, T. Lavril, G. Izacard et al. , “Llama: Open and efficient foundation language models,” arXiv preprint , 2023
2023
Closest in time.
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, “A survey on multimodal large language models,” arXiv preprint , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML , 2023
2023
Closest in time.
2023
Closest in time.
W. X. Zhao, K. Zhou, J. Li, T. Tang et al. , “A survey of large language models,” arXiv preprint , 2023
2023
Closest in time.
2023
Closest in time.
Q. Zhong, L. Ding, J. Liu, B. Du, and D. Tao, “Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert,” arXiv preprint , 2023
2023
Closest in time.
Q. Lu, B. Qiu, L. Ding, L. Xie, and D. Tao, “Error analysis prompting enables human-like translation evaluation in large language models: A case study on chatgpt,” arXiv preprint , 2023
2023
Closest in time.
K. Peng, L. Ding, Q. Zhong, L. Shen, X. Liu, M. Zhang, Y. Ouyang, and D. Tao, “Towards making the most of chatgpt for machine translation,” arxiv preprint , 2023
2023
Closest in time.
H. Deng, L. Ding, X. Liu, M. Zhang, D. Tao, and M. Zhang, “Improving simultaneous machine translation with monolingual data,” in AAAI Conference on Artificial Intelligence , 2023
2023
Closest in time.
Z. Peng, W. Wang, L. Dong, Y. Hao, S. Huang, S. Ma, and F. Wei, “Kosmos-2: Grounding multimodal large language models to the world,” ArXiv , vol. abs/2306, 2023
2023
Closest in time.