Fetching the paper…
Reading the bibliography…
Visual and textual modalities contribute complementary information about events described in multimedia documents.
The Berkeley FrameNet project
Collin F. Baker, Charles J. Fillmore, and John B. Lowe. 1998 · 1998
Earlier work this paper cites.
The stages of event extraction
David Ahn. 2006 · 2006
Earlier work this paper cites.
Refining event extraction through cross-document inference
Heng Ji and Ralph Grishman. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li. 2009 · 2009
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Fei-Fei Li. 2014 · 2014
Earlier work this paper cites.
Event extraction via dynamic multi-pooling convolutional neural networks
Yubo Chen, Liheng Xu, Kang Liu, Daojian Zeng, and Jun Zhao. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Earlier work this paper cites.
Joint event extraction via recurrent neural networks
Thien Huu Nguyen, Kyunghyun Cho, and Ralph Grishman. 2016 · 2016
Earlier work this paper cites.
Situation recognition: Visual semantic role labeling for image understanding
Mark Yatskar, Luke S. Zettlemoyer, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
João Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Automatically labeled data generation for large scale event extraction
Yubo Chen, Shulin Liu, Xiang Zhang, Kang Liu, and Jun Zhao. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017 · 2017
Earlier work this paper cites.
Situation recognition with graph neural networks
Ruiyu Li, Makarand Tapaswi, Renjie Liao, Jiaya Jia, Raquel Urtasun, and Sanja Fidler. 2017 · 2017
Earlier work this paper cites.
Recurrent models for situation recognition
Arun Mallya and Svetlana Lazebnik. 2017 · 2017
Cited alongside, same era.
Improving event extraction via multimodal integration
Tongtao Zhang, Spencer Whitehead, Hanwang Zhang, Hongzhi Li, Joseph G. Ellis, Lifu Huang, Wei Liu, Heng Ji, and Shih-Fu Chang. 2017 · 2017
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018 · 2018
Cited alongside, same era.
Zero-shot transfer learning for event extraction
Lifu Huang, Heng Ji, Kyunghyun Cho, Ido Dagan, Sebastian Riedel, and Clare Voss. 2018 · 2018
Cited alongside, same era.
Bsn: Boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang. 2018 · 2018
Cited alongside, same era.
Jointly extracting event triggers and arguments by dependency-bridge RNN and tensor-based argument interaction
Improving action segmentation via graph-based temporal reasoning
Yifei Huang, Yusuke Sugano, and Yoichi Sato. 2020 · 2020
Later among the works it cites.
Cross-media structured common space for multimedia event extraction
Manling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead, Di Lu, Heng Ji, and Shih-Fu Chang. 2020 · 2020
Later among the works it cites.
Event extraction as machine reading comprehension
Jian Liu, Yubo Chen, Kang Liu, Wei Bi, and Xiaojiang Liu. 2020 · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020 · 2020
Later among the works it cites.
Grounded situation recognition
Sarah Pratt, Mark Yatskar, Luca Weihs, Ali Farhadi, and Aniruddha Kembhavi. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lei Sha, Feng Qian, Baobao Chang, and Zhifang Sui. 2018 · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. 2018 · 2018
Cited alongside, same era.
DCFEE: A document-level Chinese financial event extraction system based on automatically labeled training data
Hang Yang, Yubo Chen, Kang Liu, Yang Xiao, and Jun Zhao. 2018 · 2018
Cited alongside, same era.
Scale up event extraction learning via automatic training data generation
Ying Zeng, Yansong Feng, Rong Ma, Zheng Wang, Rui Yan, Chongde Shi, and Dongyan Zhao. 2018 · 2018
Cited alongside, same era.
Multi-level multimodal common semantic space for image-phrase grounding
Hassan Akbari, Svebor Karaman, Surabhi Bhargava, Brian Chen, Carl Vondrick, and Shih-Fu Chang. 2019 · 2019
Cited alongside, same era.
BMN: boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen. 2019 · 2019
Cited alongside, same era.
Neural cross-lingual event detection with minimal parallel resources
Jian Liu, Yubo Chen, Kang Liu, and Jun Zhao. 2019 · 2019
Cited alongside, same era.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Image enhanced event detection in news articles
Meihan Tong, Shuai Wang, Yixin Cao, Bin Xu, Juanzi Li, Lei Hou, and Tat-Seng Chua. 2020 · 2020
Later among the works it cites.
Neural Gibbs Sampling for Joint Event Argument Extraction
Xiaozhi Wang, Shengyu Jia, Xu Han, Zhiyuan Liu, Juanzi Li, Peng Li, and Jie Zhou. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
Counterfactual contrastive learning for weakly-supervised vision-language grounding
Zhu Zhang, Zhou Zhao, Zhijie Lin, Xiuqiang He, et al. 2020 · 2020
Later among the works it cites.
Multimodal clustering networks for self-supervised learning from unlabeled videos
Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie Boggust, Rameswar Panda, Brian Kingsbury, Rogerio Feris, David Harwath, et al. 2021 · 2021
Closest in time.
Document-level event argument extraction by conditional generation
Sha Li, Heng Ji, and Jiawei Han. 2021 · 2021
Closest in time.
Vx2text: End-to-end learning of video-based text generation from multimodal inputs
Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, and Lorenzo Torresani. 2021 · 2021
Closest in time.
Refineloc: Iterative refinement for weakly-supervised action localization
Alejandro Pardo, Humam Alwassel, Fabian Caba, Ali Thabet, and Bernard Ghanem. 2021 · 2021
Closest in time.
Visual semantic role labeling for video understanding
Arka Sadhu, Tanmay Gupta, Mark Yatskar, Ram Nevatia, and Aniruddha Kembhavi. 2021 · 2021
Closest in time.
Bsn++: Complementary boundary regressor with scale-balanced relation modeling for temporal action proposal generation
Haisheng Su, Weihao Gan, Wei Wu, Yu Qiao, and Junjie Yan. 2021 · 2021
Closest in time.
Abstract Meaning Representation guided graph encoding and decoding for joint information extraction
Zixuan Zhang and Heng Ji. 2021 · 2021
Closest in time.