Fetching the paper…
Reading the bibliography…
Previous research on multimodal entity linking (MEL) has primarily employed contrastive learning as the primary objective.
Supervised contrastive learning
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020 · 2004
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
S. Chopra, R. Hadsell, and Y. LeCun. 2005 · 2005
Earlier work this paper cites.
What makes for good views for contrastive learning?
Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. 2020 · 2005
Earlier work this paper cites.
Multimodal named entity disambiguation for noisy social media posts
Seungwhan Moon, Leonardo Neves, and Vitor Carvalho. 2018 · 2008
Earlier work this paper cites.
Conditional negative sampling for contrastive learning of visual representations
Mike Wu, Milan Mosse, Chengxu Zhuang, Daniel Yamins, and Noah Goodman. 2020a · 2010
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandecic and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Learning local feature descriptors with triplets and shallow convolutional neural networks
Vassileios Balntas, Edgar Riba, Daniel Ponsa, and Krystian Mikolajczyk. 2016 · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn. 2016 · 2016
Earlier work this paper cites.
A discriminative feature learning approach for deep face recognition
Yandong Wen, Kaipeng Zhang, Zhifeng Li, and Yu Qiao. 2016 · 2016
Earlier work this paper cites.
Sphereface: Deep hypersphere embedding for face recognition
Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song. 2017 · 2017
Earlier work this paper cites.
Noise mitigation for neural entity typing and relation extraction
Yadollah Yaghoobzadeh, Heike Adel, and Hinrich Schütze. 2017 · 2017
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
Jiankang Deng, Jia Guo, and Stefanos Zafeiriou. 2018 · 2018
Earlier work this paper cites.
Cosface: Large margin cosine loss for deep face recognition
Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Zhifeng Li, Dihong Gong, Jingchao Zhou, and Wei Liu. 2018 · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Contragan: Contrastive learning for conditional image generation
Minguk Kang and Jaesik Park. 2020 · 2020
Cited alongside, same era.
Richpedia: A large-scale, comprehensive multi-modal knowledge graph
Meng Wang, Haofen Wang, Guilin Qi, and Qiushuo Zheng. 2020 · 2020
Cited alongside, same era.
Short text topic modeling with topic distribution quantization and negative sampling decoder
Xiaobao Wu, Chunping Li, Yan Zhu, and Yishu Miao. 2020b · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Cited alongside, same era.
Multimodal entity linking: A new dataset and A baseline
Jingru Gan, Jinchang Luo, Haiwei Wang, Shuhui Wang, Wei He, and Qingming Huang. 2021 · 2021
Cited alongside, same era.
Conditional contrastive learning for improving fairness in self-supervised learning
Martin Q. Ma, Yao-Hung Hubert Tsai, Paul Pu Liang, Han Zhao, Kun Zhang, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022 · 2022
Later among the works it cites.
Adaptive contrastive learning on multimodal transformer for review helpfulness prediction
Thong Nguyen, Xiaobao Wu, Anh Tuan Luu, Zhen Hai, and Lidong Bing. 2022 · 2022
Later among the works it cites.
Conditional contrastive learning with kernel
Yao-Hung Hubert Tsai, Tianqin Li, Martin Q Ma, Han Zhao, Kun Zhang, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2022 · 2022
Later among the works it cites.
Mitigating data sparsity for short text topic modeling by topic-semantic contrastive learning
Xiaobao Wu, Anh Tuan Luu, and Xinshuai Dong. 2022 · 2022
Later among the works it cites.
A contrastive framework for learning sentence representations from pairwise and triple-wise perspective in angular space
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021 · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. 2021 · 2021
Cited alongside, same era.
Entity-based knowledge conflicts in question answering
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Contrastive learning for neural topic model
Thong Nguyen and Anh Tuan Luu. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
Integrating auxiliary information in self-supervised learning
Yao-Hung Hubert Tsai, Tianqin Li, Weixin Liu, Peiyuan Liao, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2021 · 2021
Cited alongside, same era.
Attention-based multimodal entity linking with high-quality images
Li Zhang, Zhixu Li, and Qiang Yang. 2021 · 2021
Cited alongside, same era.
Yuhao Zhang, Hongji Zhu, Yongliang Wang, Nan Xu, Xiaobo Li, and Binqiang Zhao. 2022 · 2022
Later among the works it cites.
Multi-grained multimodal interaction network for entity linking
Pengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu, Linli Xu, and Enhong Chen. 2023 · 2023
Later among the works it cites.
A dual-way enhanced framework from text matching point of view for multimodal entity linking
Shezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan, Shasha Li, Xiaoguang Mao, and Meng Wang. 2023 · 2023
Later among the works it cites.
MMEL: A joint learning framework for multi-mention entity linking
Chengmei Yang, Bowei He, Yimeng Wu, Chao Xing, Lianghua He, and Chen Ma. 2023 · 2023
Later among the works it cites.
Visual entity linking via multi-modal learning
Qiushuo Zheng, Hao Wen, Meng Wang, and Guilin Qi. 2022 · 2023
Later among the works it cites.
Generative multimodal entity linking
Senbao Shi, Zhenran Xu, Baotian Hu, and Min Zhang. 2024 · 2024
Later among the works it cites.
Learning facial expression and body gesture visual information for video emotion recognition
Jie Wei, Guanyu Hu, Xinyu Yang, Anh Tuan Luu, and Yizhuo Dong. 2024 · 2024
Later among the works it cites.
Xiaobao Wu, Xinshuai Dong, Liangming Pan, Thong Nguyen, and Anh Tuan Luu. 2024a · 2024
Later among the works it cites.
Meta-optimized angular margin contrastive framework for video-language representation learning
Thong Nguyen, Yi Bin, Xiaobao Wu, Xinshuai Dong, Zhiyuan Hu, Khoi Le, Cong-Duy Nguyen, See-Kiong Ng, and Luu Anh Tuan. 2025 · 2025
Closest in time.