Fetching the paper…
Reading the bibliography…
Visual Entity Linking (VEL) is a task to link regions of images with their corresponding entities in Knowledge Bases (KBs), which is beneficial for many computer vision tasks such as image retrieval, image caption, and visual question answering.
Zero-shot entity linking by reading entity descriptions
Lajanugen Logeswaran, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, Jacob Devlin, and Honglak Lee. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Global entity disambiguation with pretrained contextualized embeddings of words and entities
Ikuya Yamada, Koki Washio, Hiroyuki Shindo, and Yuji Matsumoto. 2019 · 1909
Earlier work this paper cites.
Scalable zero-shot entity linking with dense entity retrieval
Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel, and Luke Zettlemoyer. 2019 · 1911
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary G. Ives. 2007 · 2007
Earlier work this paper cites.
Yago: A core of semantic knowledge unifying wordnet and wikipedia
M Fabian, Kasneci Gjergji, WEIKUM Gerhard, et al. 2007 · 2007
Earlier work this paper cites.
Image retrieval: Ideas, influences, and trends of the new age
Ritendra Datta, Dhiraj Joshi, Jia Li, and James Z Wang. 2008 · 2008
Earlier work this paper cites.
Autoregressive entity retrieval
Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. 2020 · 2010
Earlier work this paper cites.
Robust disambiguation of named entities in text
Johannes Hoffart, Mohamed Amir Yosef, Ilaria Bordino, Hagen Fürstenau, Manfred Pinkal, Marc Spaniol, Bilyana Taneva, Stefan Thater, and Gerhard Weikum. 2011 · 2011
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Visual entity linking: A preliminary study
Rebecka Weegar, Linus Hammarlund, Agnes Tegen, Magnus Oskarsson, Kalle Åström, and Pierre Nugues. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Cited alongside, same era.
Mediaeval 2016: A multimodal system for the verifying multimedia use task
Cédric Maigrot, Vincent Claveau, Ewa Kijak, and Ronan Sicre. 2016 · 2016
Cited alongside, same era.
Joint face detection and alignment using multitask cascaded convolutional networks
Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li, and Yu Qiao. 2016 · 2016
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019 · 2019
Later among the works it cites.
Multimodal entity linking for tweets
Omar Adjali, Romaric Besançon, Olivier Ferret, Herve Le Borgne, and Brigitte Grau. 2020 · 2020
Later among the works it cites.
Autoregressive entity retrieval
N. D. Cao, G. Izacard, S. Riedel, and F. Petroni. 2020 · 2020
Later among the works it cites.
Vt-linker: Visual-textual-knowledge entity linker
Shahi Dost, Luciano Serafini, Marco Rospocher, Lamberto Ballan, and Alessandro Sperduti. 2020 · 2020
Later among the works it cites.
Visualnews : Benchmark and challenges in entity-aware image captioning
F. Liu, Y. Wang, T. Wang, and V. Ordonez. 2020 · 2020
Later among the works it cites.
Evaluating the impact of knowledge graph context on entity disambiguation models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Breakingnews: Article annotation by image and text processing
A. Ramisa, Fei Yan, Francesc Moreno-Noguer, and K. Mikolajczyk. 2017 · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. 2017 · 2017
Cited alongside, same era.
A context-driven extractive framework for generating realistic image descriptions
Amara Tariq and Hassan Foroosh. 2017 · 2017
Cited alongside, same era.
Visual entity linking
Neha Tilak, Sunil Gandhi, and Tim Oates. 2017 · 2017
Cited alongside, same era.
Vggface2: A dataset for recognising faces across pose and age
Qiong Cao, Li Shen, Weidi Xie, Omkar M Parkhi, and Andrew Zisserman. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Multimodal named entity disambiguation for noisy social media posts
S. Moon, L. Neves, and V. Carvalho. 2018 · 2018
Cited alongside, same era.
Isaiah Onando Mulang’, Kuldeep Singh, Chaitali Prabhu, Abhishek Nadgeri, Johannes Hoffart, and Jens Lehmann. 2020 · 2020
Later among the works it cites.
Transform and tell: Entity-aware news image captioning
Alasdair Tran, Alexander Mathews, and Lexing Xie. 2020 · 2020
Later among the works it cites.
Multimodal entity linking: a new dataset and a baseline
Jingru Gan, Jinchang Luo, Haiwei Wang, Shuhui Wang, Wei He, and Qingming Huang. 2021 · 2021
Later among the works it cites.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. 2021 · 2021
Later among the works it cites.
Multimodal news analytics using measures of cross-modal entity and context consistency
Eric Müller-Budack, Jonas Theiner, Sebastian Diering, Maximilian Idahl, Sherzod Hakimov, and Ralph Ewerth. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, and J. Clark. 2021 · 2021
Later among the works it cites.
Neural entity linking: A survey of models based on deep learning
Özge Sevgili, Artem Shelmanov, Mikhail Y. Arkhipov, Alexander Panchenko, and Chris Biemann. 2022 · 2022
Closest in time.
Wikidiverse: A multimodal entity linking dataset with diversified contextual topics and entity types
X. Wang, J. Tian, M. Gui, Z. Li, R. Wang, M. Yan, L. Chen, and Y. Xiao. 2022 · 2022
Closest in time.
Visual entity linking via multi-modal learning
Qiushuo Zheng, Hao Wen, Meng Wang, and Guilin Qi. 2022 · 2022
Closest in time.