Fetching the paper…
Reading the bibliography…
Multimodal named entity recognition and relation extraction (MNER and MRE) is a fundamental and crucial branch in information extraction.
Aiding intra-text representations with visual context for multimodal named entity recognition
Omer Arshad, Ignazio Gallo, Shah Nawaz, and Alessandro Calefati. 2019 · 1904
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
Learning rich features at high-speed for single-shot object detection
Tiancai Wang, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, and Ling Shao. 2019 · 1980
Earlier work this paper cites.
Visual attention model for name tagging in multimodal social media
Di Lu, Leonardo Neves, Vitor Carvalho, Ning Zhang, and Heng Ji. 2018 · 1999
Earlier work this paper cites.
Shuguang Chen, Gustavo Aguilar, Leonardo Neves, and Thamar Solorio. 2020a · 2010
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Distant supervision for relation extraction via piecewise convolutional neural networks
Daojian Zeng, Kang Liu, Yubo Chen, and Jun Zhao. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Neural architectures for named entity recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016a · 2016
Earlier work this paper cites.
Neural architectures for named entity recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016b · 2016
Earlier work this paper cites.
End-to-end sequence labeling via bi-directional lstm-cnns-crf
Xuezhe Ma and Eduard H. Hovy. 2016 · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. 2017 · 2017
Earlier work this paper cites.
Parallel feature pyramid network for object detection
Seung-Wook Kim, Hyong-Keun Kook, Jee-Young Sun, Mun-Cheon Kang, and Sung-Jea Ko. 2018 · 2018
Earlier work this paper cites.
Path aggregation network for instance segmentation
Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. 2018 · 2018
Earlier work this paper cites.
Multimodal named entity recognition for short social media posts
Seungwhan Moon, Leonardo Neves, and Vitor Carvalho. 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Cited alongside, same era.
Adaptive co-attention network for named entity recognition in tweets
Qi Zhang, Jinlan Fu, Xiaoyu Liu, and Xuanjing Huang. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Implicit entity recognition, classification and linking in tweets
Hawre Hosseini. 2019 · 2019
Cited alongside, same era.
What does BERT learn about the structure of language?
Unbiased scene graph generation from biased training
Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang. 2020 · 2020
Later among the works it cites.
Multimodal representation with embedded visual guiding objects for named entity recognition in social media posts
Zhiwei Wu, Changmeng Zheng, Yi Cai, Junying Chen, Ho-Fung Leung, and Qing Li 0001. 2020 · 2020
Later among the works it cites.
Improving multimodal named entity recognition via entity span detection with unified multimodal transformer
Jianfei Yu, Jing Jiang, Li Yang, and Rui Xia. 2020 · 2020
Later among the works it cites.
Openue: An open toolkit of universal extraction from text
Ningyu Zhang, Shumin Deng, Zhen Bi, Haiyang Yu, Jiacheng Yang, Mosha Chen, Fei Huang, Wei Zhang, and Huajun Chen. 2020 · 2020
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Cited alongside, same era.
GCDT: A global context enhanced deep transition architecture for sequence labeling
Yijin Liu, Fandong Meng, Jinchao Zhang, Jinan Xu, Yufeng Chen, and Jie Zhou. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Matching the blanks: Distributional similarity for relation learning
Livio Baldini Soares, Nicholas FitzGerald, Jeffrey Ling, and Tom Kwiatkowski. 2019 · 2019
Cited alongside, same era.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
A fast and accurate one-stage approach to visual grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo. 2019 · 2019
Cited alongside, same era.
UNITER: universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020b · 2020
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Later among the works it cites.
Noisy-labeled NER with confidence estimation
Kun Liu, Yao Fu, Chuanqi Tan, Mosha Chen, Ningyu Zhang, Songfang Huang, and Sheng Gao. 2021 · 2021
Later among the works it cites.
ERICA: improving entity and relation understanding for pre-trained language models via contrastive learning
Yujia Qin, Yankai Lin, Ryuichi Takanobu, Zhiyuan Liu, Peng Li, Heng Ji, Minlie Huang, Maosong Sun, and Jie Zhou. 2021 · 2021
Later among the works it cites.
Rpbert: A text-image relation propagation-based BERT model for multimodal NER
Lin Sun, Jiquan Wang, Kai Zhang, Yindu Su, and Fangsheng Weng. 2021 · 2021
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2021 · 2021
Later among the works it cites.
Multi-modal graph fusion for named entity recognition with targeted visual guidance
Dong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu, Qiaoming Zhu, and Guodong Zhou. 2021a · 2021
Later among the works it cites.
Document-level relation extraction as semantic segmentation
Ningyu Zhang, Xiang Chen, Xin Xie, Shumin Deng, Chuanqi Tan, Mosha Chen, Fei Huang, Luo Si, and Huajun Chen. 2021b · 2021
Later among the works it cites.
Multimodal relation extraction with efficient graph alignment
Changmeng Zheng, Junhao Feng, Ze Fu, Yi Cai, Qing Li, and Tao Wang. 2021 · 2021
Later among the works it cites.
Contrastive demonstration tuning for pre-trained language models
Xiaozhuan Liang, Ningyu Zhang, Siyuan Cheng, Zhen Bi, Zhenru Zhang, Chuanqi Tan, Songfang Huang, Fei Huang, and Huajun Chen. 2022 · 2022
Closest in time.