Fetching the paper…
Reading the bibliography…
Recently, Multi-modal Named Entity Recognition (MNER) has attracted a lot of attention.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
Named entity task definition, version 2.1
Beth M. Sundheim. 1995 · 1995
Earlier work this paper cites.
Visual attention model for name tagging in multimodal social media
Di Lu, Leonardo Neves, Vitor Carvalho, Ning Zhang, and Heng Ji. 2018 · 1999
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001 · 2001
Earlier work this paper cites.
Introduction to the CoNLL-2002 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang. 2002 · 2002
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
An overview of the tesseract ocr engine
Ray Smith. 2007 · 2007
Earlier work this paper cites.
Flert: Document-level features for named entity recognition
Stefan Schweter and Alan Akbik. 2020 · 2011
Earlier work this paper cites.
Ncbi disease corpus: a resource for disease name recognition and concept normalization
Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, M. Hodosh, and J. Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. L. Zitnick, Devi Parikh, and Dhruv Batra. 2015 · 2015
Earlier work this paper cites.
Bidirectional lstm-crf models for sequence tagging
Zhiheng Huang, W. Xu, and Kailiang Yu. 2015 · 2015
Earlier work this paper cites.
Context-aware image tweet modelling and recommendation
Tao Chen, Xiangnan He, and Min-Yen Kan. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Biocreative v cdr task corpus: a resource for chemical disease relation extraction
Jiao Li, Yueping Sun, Robin J Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J Mattingly, Thomas C Wiegers, and Zhiyong Lu. 2016 · 2016
Cited alongside, same era.
Results of the WNUT16 named entity recognition shared task
Benjamin Strauss, Bethany Toma, Alan Ritter, Marie-Catherine de Marneffe, and Wei Xu. 2016 · 2016
Cited alongside, same era.
What value do explicit high level concepts have in vision to language problems?
Qi Wu, Chunhua Shen, Lingqiao Liu, Anthony Dick, and Anton Van Den Hengel. 2016 · 2016
Cited alongside, same era.
Results of the WNUT2017 shared task on novel and emerging entity recognition
Leon Derczynski, Eric Nichols, Marieke van Erp, and Nut Limsopatham. 2017 · 2017
Cited alongside, same era.
Self-critical sequence training for image captioning
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017 · 2017
Cited alongside, same era.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Hao Tan and M. Bansal. 2019 · 2019
Later among the works it cites.
Categorizing and inferring the relationship between the text and image of Twitter posts
Alakananda Vempala and Daniel Preoţiuc-Pietro. 2019 · 2019
Later among the works it cites.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Contextual string embeddings for sequence labeling
Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018 · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Cited alongside, same era.
Discriminability objective for training descriptive captions
R. Luo, Brian L. Price, Scott D. Cohen, and Gregory Shakhnarovich. 2018 · 2018
Cited alongside, same era.
Multimodal named entity recognition for short social media posts
Seungwhan Moon, Leonardo Neves, and Vitor Carvalho. 2018 · 2018
Cited alongside, same era.
Adaptive co-attention network for named entity recognition in tweets
Qi Zhang, Jinlan Fu, Xiaoyu Liu, and Xuanjing Huang. 2018 · 2018
Cited alongside, same era.
Pooled contextualized embeddings for named entity recognition
Alan Akbik, Tanja Bergmann, and Roland Vollgraf. 2019 · 2019
Cited alongside, same era.
RIVA: A pre-trained tweet multimodal model based on text-image relation for multimodal NER
Lin Sun, Jiquan Wang, Yindu Su, Fangsheng Weng, Yuxuan Sun, Zengwei Zheng, and Yuanyi Chen. 2020 · 2020
Later among the works it cites.
Cross-media keyphrase prediction: A unified framework with multi-modality multi-head attention and image wordings
Yue Wang, Jing Li, Michael Lyu, and Irwin King. 2020 · 2020
Later among the works it cites.
Multimodal representation with embedded visual guiding objects for named entity recognition in social media posts
Zhiwei Wu, Changmeng Zheng, Yi Cai, Junying Chen, Ho-fung Leung, and Qing Li. 2020 · 2020
Later among the works it cites.
LUKE: Deep contextualized entity representations with entity-aware self-attention
Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda, and Yuji Matsumoto. 2020 · 2020
Later among the works it cites.
Improving multimodal named entity recognition via entity span detection with unified multimodal transformer
Jianfei Yu, Jing Jiang, Li Yang, and Rui Xia. 2020 · 2020
Later among the works it cites.
Gazetteer enhanced named entity recognition for code-mixed web queries
Besnik Fetahu, Anjie Fang, Oleg Rokhlenko, and Shervin Malmasi. 2021 · 2021
Closest in time.
Rpbert: A text-image relation propagation-based bert model for multimodal ner
Lin Sun, Jiquan Wang, Kai Zhang, Yindu Su, and Fangsheng Weng. 2021 · 2021
Closest in time.
Improving named entity recognition by external context retrieving and cooperative learning
Xinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, and Kewei Tu. 2021 · 2021
Closest in time.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021 · 2021
Closest in time.