Fetching the paper…
Reading the bibliography…
Multimodal deep learning, especially vision-language models, have gained significant traction in recent years, greatly improving performance on many downstream tasks, including content moderation and violence detection.
“Cognitron: A self-organizing multilayered neural network,”
Kunihiko Fukushima, · 1975
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Crisishatemm: Multimodal analysis of directed and undirected hate speech in text-embedded images from russia-ukraine conflict,”
Aashish Bhandari, Siddhant B Shah, Surendrabikram Thapa, Usman Naseem, and Mehwish Nasim, · 2002
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“A multimodal approach to violence detection in video sharing sites,”
Theodoros Giannakopoulos, Aggelos Pikrakis, and Sergios Theodoridis, · 2010
Earlier work this paper cites.
“Multimodal deep learning,”
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng, · 2011
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Multimodal machine learning: A survey and taxonomy,”
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency, · 2018
Earlier work this paper cites.
“Multimodal unsupervised image-to-image translation,”
Xun Huang, Ming-Yu Liu, Serge Belongie, and Jan Kautz, · 2018
Earlier work this paper cites.
“Empowering first responders through automated multimodal content moderation,”
Divam Gupta, Indira Sen, Niharika Sachdeva, Ponnurangam Kumaraguru, and Arun Balaji Buduru, · 2018
Earlier work this paper cites.
“Learning factorized multimodal representations,”
Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov, · 2018
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
“Multimodal transformer with multi-view visual representation for image captioning,”
Jun Yu, Jing Li, Zhou Yu, and Qingming Huang, · 2019
Cited alongside, same era.
“A multimodal fusion approach for image captioning,”
Dexin Zhao, Zhi Chang, and Shutao Guo, · 2019
Cited alongside, same era.
“Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,”
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf, · 2019
“Multimodal sentiment analysis with image-text interaction network,”
Tong Zhu, Leida Li, Jufeng Yang, Sicheng Zhao, Hantao Liu, and Jiansheng Qian, · 2022
Later among the works it cites.
Cem Akkus, Luyang Chu, Vladana Djakovic, Steffen Jauch-Walser, Philipp Koch, Giacomo Loss, Christopher Marquardt, Marco Moldovan, Nadja Sauter, Maximilian Schneider, et al., · 2023
Closest in time.
“Multi-modal medical image classification using deep residual network and genetic algorithm,”
Muhammad Haris Abid, Rehan Ashraf, Toqeer Mahmood, and CM Nadeem Faisal, · 2023
Closest in time.
“Evaluating machine learning models with nero: Non-equivariance revealed on orbits,”
Zhuokai Zhao, Takumi Matsuzawa, William Irvine, Michael Maire, and Gordon L Kindlmann, · 2023
Closest in time.
“Towards good practices for missing modality robust action recognition,”
Sangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho, and Changick Kim, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Multi-modal classification using images and text,”
Stuart J Miller, Justin Howard, Paul Adams, Mel Schwan, and Robert Slater, · 2020
Cited alongside, same era.
Video understanding using multimodal deep learning
Arsha Nagrani, · 2020
Cited alongside, same era.
“Not only look, but also listen: Learning multimodal violence detection under weak supervision,”
Peng Wu, Jing Liu, Yujia Shi, Yujia Sun, Fangtao Shao, Zhaoyang Wu, and Zhiwei Yang, · 2020
Cited alongside, same era.
“Smil: Multimodal learning with severely missing modality,”
Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng, · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Cited alongside, same era.
Multimodal Learning from Videos: Exploring Models and Task Complexities
Shruti Palaskar, · 2022
Cited alongside, same era.
“Are multimodal transformers robust to missing modality?,”
Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine, and Xi Peng, · 2022
Cited alongside, same era.
Closest in time.
“A survey of video violence detection,”
Huiling Yao and Xing Hu, · 2023
Closest in time.
“Mobileone: An improved one millisecond mobile backbone,”
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, and Anurag Ranjan, · 2023
Closest in time.
“Direct acquisition optimization for low-budget active learning,”
Zhuokai Zhao, Yibo Jiang, and Yuxin Chen, · 2024
Closest in time.
“Rankclip: Ranking-consistent language-image pretraining,”
Yiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng, Zenghui Ding, and Yining Sun, · 2024
Closest in time.
“Halc: Object hallucination reduction via adaptive focal-contrast decoding,”
Zhaorun Chen, Zhuokai Zhao, Hongyin Luo, Huaxiu Yao, Bo Li, and Jiawei Zhou, · 2024
Closest in time.
Zhaorun Chen, Zhuokai Zhao, Zhihong Zhu, Ruiqi Zhang, Xiang Li, Bhiksha Raj, and Huaxiu Yao, · 2024
Closest in time.