Fetching the paper…
Reading the bibliography…
We introduce a new task, MultiMedia Event Extraction (M2E2), which aims to extract events and their arguments from multimedia documents.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid. 2019a · 1906
Earlier work this paper cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou. 2019a · 1908
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019b · 1908
Earlier work this paper cites.
M-bert: Injecting multimodal information in the bert structure
Wasifur Rahman, Md Kamrul Hasan, Amir Zadeh, Louis-Philippe Morency, and Mohammed Ehsan Hoque. 2019 · 1908
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019b · 1908
Earlier work this paper cites.
Open event extraction from online text using a generative adversarial network
Rui Wang, Deyu Zhou, and Yulan He. 2019 · 1908
Earlier work this paper cites.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 1909
Earlier work this paper cites.
Entity, relation, and event extraction with contextualized span representations
David Wadden, Ulme Wennberg, Yi Luan, and Hannaneh Hajishirzi. 2019 · 1909
Earlier work this paper cites.
Wordnet: a lexical database for english
George A Miller. 1995 · 1995
Earlier work this paper cites.
The Rise of the Image, The Fall of the Word
Mitchell Stephens. 1998 · 1998
Earlier work this paper cites.
Background to framenet
Charles J Fillmore, Christopher R Johnson, and Miriam RL Petruck. 2003 · 2003
Earlier work this paper cites.
Ace 2005 multilingual training corpus
Christopher Walker, Stephanie Strassel, Julie Medero, and Kazuaki Maeda. 2006 · 2005
Earlier work this paper cites.
Viterbi based alignment between text images and their transcripts
Alejandro H Toselli, Verónica Romero, and Enrique Vidal. 2007 · 2007
Earlier work this paper cites.
Semantic event extraction from basketball games using multi-modal analysis
Yifan Zhang, Changsheng Xu, Yong Rui, Jinqiao Wang, and Hanqing Lu. 2007 · 2007
Earlier work this paper cites.
Refining event extraction through cross-document inference
Heng Ji and Ralph Grishman. 2008 · 2008
Earlier work this paper cites.
Acquiring topic features to improve event extraction: in pre-selected and balanced collections
Shasha Liao and Ralph Grishman. 2011 · 2011
Earlier work this paper cites.
Bootstrapped training of event extraction classifiers
Ruihong Huang and Ellen Riloff. 2012 · 2012
Earlier work this paper cites.
Trecvid 2012 genie: Multimedia event detection and recounting
AG Amitha Perera, Sangmin Oh, P Megha, Tianyang Ma, Anthony Hoogs, Arash Vahdat, Kevin Cannons, Greg Mori, Scott Mccloskey, Ben Miller, et al. 2012 · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012 · 2012
Earlier work this paper cites.
Abstract meaning representation for sembanking
Laura Banarescu, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider. 2013 · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. 2013 · 2013
Cited alongside, same era.
Joint event extraction via structured prediction with global features
Qi Li, Heng Ji, and Liang Huang. 2013 · 2013
Cited alongside, same era.
The Stanford CoreNLP natural language processing toolkit
Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
Improving event extraction via multimodal integration
Tongtao Zhang, Spencer Whitehead, Hanwang Zhang, Hongzhi Li, Joseph Ellis, Lifu Huang, Wei Liu, Heng Ji, and Shih-Fu Chang. 2017 · 2017
Later among the works it cites.
Collective event detection via a hierarchical and bias tagging networks with gated multi-level attention mechanisms
Yubo Chen, Hang Yang, Kang Liu, Jun Zhao, and Yantao Jia. 2018 · 2018
Later among the works it cites.
Videocapsulenet: A simplified network for action detection
Kevin Duarte, Yogesh Rawat, and Mubarak Shah. 2018 · 2018
Later among the works it cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018 · 2018
Later among the works it cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015 · 2015
Cited alongside, same era.
Event extraction via dynamic multi-pooling convolutional neural networks
Yubo Chen, Liheng Xu, Kang Liu, Daojian Zeng, and Jun Zhao. 2015 · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
A transition-based algorithm for amr parsing
Chuan Wang, Nianwen Xue, and Sameer Pradhan. 2015b · 2015
Cited alongside, same era.
Linking entities across images and text
Rebecka Weegar, Kalle Åström, and Pierre Nugues. 2015 · 2015
Cited alongside, same era.
Eventnet: A large scale structured concept library for complex event detection in video
Guangnan Ye, Yitong Li, Hongliang Xu, Dong Liu, and Shih-Fu Chang. 2015 · 2015
Cited alongside, same era.
Learning translations via images with a massively multilingual image dataset
John Hewitt, Daphne Ippolito, Brendan Callahan, Reno Kriz, Derry Tanti Wijaya, and Chris Callison-Burch. 2018 · 2018
Later among the works it cites.
Self-regulation: Employing a generative adversarial network to improve event detection
Yu Hong, Wenxuan Zhou, jingli zhang jingli, Guodong Zhou, and Qiaoming Zhu. 2018 · 2018
Later among the works it cites.
Compositional learning for human object interaction
Keizo Kato, Yin Li, and Abhinav Gupta. 2018 · 2018
Later among the works it cites.
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, et al. 2018 · 2018
Later among the works it cites.
Recurrent tubelet proposal and recognition networks for action detection
Dong Li, Zhaofan Qiu, Qi Dai, Ting Yao, and Tao Mei. 2018 · 2018
Later among the works it cites.
Jointly multiple events extraction via attention-based graph information aggregation
Xiao Liu, Zhunchen Luo, and Heyan Huang. 2018a · 2018
Later among the works it cites.
Learning visually-grounded semantics from contrastive adversarial samples
Haoyue Shi, Jiayuan Mao, Tete Xiao, Yuning Jiang, and Jian Sun. 2018 · 2018
Later among the works it cites.
Grounding semantic roles in images
Carina Silberer and Manfred Pinkal. 2018 · 2018
Later among the works it cites.
Reliability-aware dynamic feature composition for name tagging
Ying Lin, Liyuan Liu, Heng Ji, Dong Yu, and Jiawei Han. 2019 · 2019
Later among the works it cites.
Focus your attention: A bidirectional focal attention network for image-text matching
Chunxiao Liu, Zhendong Mao, An-An Liu, Tianzhu Zhang, Bin Wang, and Yongdong Zhang. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Adversarial representation learning for text-to-image matching
Nikolaos Sarafianos, Xiang Xu, and Ioannis A. Kakadiaris. 2019 · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Exploring pre-trained language models for event extraction and generation
Sen Yang, Dawei Feng, Linbo Qiao, Zhigang Kan, and Dongsheng Li. 2019 · 2019
Later among the works it cites.
Joint entity and event extraction with generative adversarial imitation learning
Tongtao Zhang, Heng Ji, and Avirup Sil. 2019 · 2019
Later among the works it cites.