Fetching the paper…
Reading the bibliography…
In the past few years, there has been a surge of interest in multi-modal problems, from image captioning to visual question answering and beyond.
The use of the area under the roc curve in the evaluation of machine learning algorithms
Bradley, A. P · 1997
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on Twitter
Waseem, Z. and Hovy, D · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
Devlin, J., Gupta, S., Girshick, R., Mitchell, M., and Zitnick, C. L · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y · 2015
Earlier work this paper cites.
Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter
Waseem, Z · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M. W., and Weber, I · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Detecting hate speech in social media
Malmasi, S. and Zampieri, M · 2017
Cited alongside, same era.
A survey of multimodal sentiment analysis
Soleymani, M., Garcia, D., Jou, B., Schuller, B., Chang, S.-F., and Pantic, M · 2017
Cited alongside, same era.
Cross-media learning for image sentiment analysis in the wild
Vadicamo, L., Carrara, F., Cimino, A., Cresci, S., Dell’Orletta, F., Falchi, F., and Tesconi, M · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Contextual inter-modal attention for multi-modal sentiment analysis
Exploring deep multimodal fusion of text and photo for hate speech classification
Yang, F., Peng, X., Ghosh, G., Shilon, R., Ma, H., Moore, E., and Predovic, G · 2019
Later among the works it cites.
Uniter: Universal image-text representation learning, 2020
Chen, Y.-C., Li, L., Yu, L., Kholy, A. E., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J · 2020
Closest in time.
Exploring hate speech detection in multimodal publications
Gomez, R., Gibert, J., Gomez, L., and Karatzas, D · 2020
Closest in time.
Captioning images taken by people who are blind, 2020
Gurari, D., Zhao, Y., Zhang, M., and Bhattacharya, N · 2020
Closest in time.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D., Firooz, H., Mohan, A., Goswami, V., Singh, A., Ringshia, P., and Testuggine, D · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ghosal, D., Akhtar, M. S., Chauhan, D., Poria, S., Ekbal, A., and Bhattacharyya, P · 2018
Cited alongside, same era.
Benchmarking aggression identification in social media
Kumar, R., Ojha, A. K., Malmasi, S., and Zampieri, M · 2018
Cited alongside, same era.
Multimodal sentiment analysis using hierarchical fusion with context modeling
Majumder, N., Hazarika, D., Gelbukh, A., Cambria, E., and Poria, S · 2018
Cited alongside, same era.
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
Zadeh, A. B., Liang, P. P., Poria, S., Cambria, E., and Morency, L.-P · 2018
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Kumar, A. and Vepa, J · 2020
Closest in time.
Oscar: Object-semantics aligned pre-training for vision-language tasks, 2020
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., and Gao, J · 2020
Closest in time.
Shenoy, A. and Sardana, A · 2020
Closest in time.
Textcaps: a dataset for image captioning with reading comprehension, 2020
Sidorov, O., Hu, R., Rohrbach, M., and Singh, A · 2020
Closest in time.
Mmf: A multimodal framework for vision and language research
Singh, A., Goswami, V., Natarajan, V., Jiang, Y., Chen, X., Shah, M., Rohrbach, M., Batra, D., and Parikh, D · 2020
Closest in time.