Fetching the paper…
Reading the bibliography…
Deep Learning and its applications have cascaded impactful research and development with a diverse range of modalities present in the real-world data.
Generating multiple objects at spatially distinct locations
Hinz, T., Heinrich, S., Wermter, S., 2019 · 1901
Earlier work this paper cites.
Self-monitoring navigation agent via auxiliary progress estimation
Ma, C.Y., Lu, J., Wu, Z., Al-Regib, G., Kira, Z., Socher, R., Xiong, C., 2019a · 1901
Earlier work this paper cites.
Embodied multimodal multitask learning
Chaplot, D.S., Lee, L., Salakhutdinov, R., Parikh, D., Batra, D., 2019 · 1902
Earlier work this paper cites.
Gqa: a new dataset for compositional question answering over real-world images
Hudson, D.A., Manning, C.D., 2019 · 1902
Earlier work this paper cites.
Dual attention networks for visual reference resolution in visual dialog
Kang, G.C., Lim, J., Zhang, B.T., 2019 · 1902
Earlier work this paper cites.
Large-scale answerer in questioner’s mind for visual dialog question generation
Lee, S.W., Gao, T., Yang, S., Yoo, J., Ha, J.W., 2019 · 1902
Earlier work this paper cites.
Multimodal machine translation with embedding prediction
Hirasawa, T., Yamagishi, H., Matsumura, Y., Komachi, M., 2019 · 1904
Earlier work this paper cites.
Learning to navigate unseen environments: Back translation with environmental dropout
Tan, H., Yu, L., Bansal, M., 2019b · 1904
Earlier work this paper cites.
Multi-task learning for multi-modal emotion recognition and sentiment analysis
Akhtar, M.S., Chauhan, D.S., Ghosal, D., Poria, S., Ekbal, A., Bhattacharyya, P., 2019 · 1905
Earlier work this paper cites.
Debiasing word embeddings improves multimodal machine translation
Hirasawa, T., Komachi, M., 2019 · 1905
Earlier work this paper cites.
Leveraging medical visual question answering with supporting facts
Kornuta, T., Rajan, D., Shivade, C., Asseman, A., Ozcan, A.S., 2019 · 1905
Earlier work this paper cites.
What makes training multi-modal networks hard?
Wang, W., Tran, D., Feiszli, M., 2019b · 1905
Earlier work this paper cites.
Self-critical reasoning for robust visual question answering
Wu, J., Mooney, R.J., 2019b · 1905
Earlier work this paper cites.
Distilling translations with visual awareness
Ive, J., Madhyastha, P., Specia, L., 2019 · 1906
Earlier work this paper cites.
Stand-alone self-attention in vision models
Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., Shlens, J., 2019 · 1906
Earlier work this paper cites.
Multi-scale guided attention for medical image segmentation
Sinha, A., Dolz, J.E., 2019 · 1906
Earlier work this paper cites.
Contrastive bidirectional transformer for temporal representation learning
Sun, C., Baradel, F., Murphy, K., Schmid, C., 2019a · 1906
Earlier work this paper cites.
Selfie: Self-supervised pretraining for image embedding
Trinh, T.H., Luong, M.T., Le, Q.V., 2019 · 1906
Earlier work this paper cites.
Informative image captioning with external sources of information
Zhao, S., Sharma, P., Levinboim, T., Soricut, R., 2019b · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V., 2019b · 1907
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L.H., Yatskar, M., Yin, D., Hsieh, C.J., Chang, K.W., 2019b · 1908
Earlier work this paper cites.
Focal visual-text attention for memex question answering
Liang, J., Jiang, L., Cao, L., Kalantidis, Y., Li, L.J., Hauptmann, A.G., 2019a · 1908
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., Lee, S., 2019a · 1908
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., Dai, J., 2020 · 1908
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H.H., Bansal, M., 2019 · 1908
Earlier work this paper cites.
Multimodal unified attention networks for vision-and-language interactions
Yu, Z., Cui, Y., Yu, J., Tao, D., Tian, Q., 2019b · 1908
Earlier work this paper cites.
Explainable high-order visual question reasoning: A new benchmark and knowledge-routed network
Cao, Q., Li, B., Liang, X., Lin, L., 2019a · 1909
Earlier work this paper cites.
Tiger: Text-to-image grounding for image caption evaluation
Jiang, M., Huang, Q., Zhang, L., Wang, X., Zhang, P., Gan, Z., Diesner, J., Gao, J., 2019 · 1909
Earlier work this paper cites.
Supervised multimodal bitransformers for classifying images and text
Kiela, D., Bhooshan, S., Firooz, H., Testuggine, D., 2019 · 1909
Earlier work this paper cites.
Probabilistic framework for solving visual dialog
Patro, B.N., Anupriy, Namboodiri, V.P., 2019 · 1909
Earlier work this paper cites.
Hierarchical attention networks for medical image segmentation
Ding, F., Yang, G., Liu, J., Wu, J., Ding, D., Xu, J., Cheng, G., Li, X., 2019 · 1911
Earlier work this paper cites.
Counterfactual vision-and-language navigation via adversarial path sampling
Fu, T.J., Wang, X., Peterson, M., Grafton, S.T., Eckstein, M., Wang, W.Y., 2019 · 1911
Earlier work this paper cites.
Perceive, transform, and act: Multi-modal attention networks for vision-and-language navigation
Landi, F., Baraldi, L., Cornia, M., Corsini, M., Cucchiara, R., 2019a · 1911
Earlier work this paper cites.
Question-conditioned counterfactual image generation for vqa
Pan, J., Goyal, Y., Lee, S., 2019 · 1911
Earlier work this paper cites.
Dynamic fusion for multimodal data
Sahu, G., Vechtomova, O., 2019 · 1911
Earlier work this paper cites.
Sun, Z., Sarma, P.K., Sethares, W.A., Liang, Y., 2019c · 1911
Earlier work this paper cites.
Open-ended visual question answering by multi-modal domain adaptation
Xu, Y., Chen, L., Cheng, Z., Duan, L., Luo, J., 2019a · 1911
Earlier work this paper cites.
Agarwal, V., Shetty, R., Fritz, M., 2019 · 1912
Earlier work this paper cites.
Exposing and correcting the gender bias in image captioning datasets and models
Bhargava, S., Forsyth, D., 2019 · 1912
Earlier work this paper cites.
M2: Meshed-memory transformer for image captioning
Cornia, M., Stefanini, M., Baraldi, L., Cucchiara, R., 2019b · 1912
Earlier work this paper cites.
Assessing the robustness of visual question answering
Huang, J., Alfadly, M., Ghanem, B., Worring, M., 2019b · 1912
Earlier work this paper cites.
12-in-1: Multi-task vision and language representation learning
Lu, J., Goswami, V., Rohrbach, M., Parikh, D., Lee, S., 2019b · 1912
Earlier work this paper cites.
Large-scale pretraining for visual dialog: A simple state-of-the-art baseline
Murahari, V.S., Batra, D., Parikh, D., Das, A., 2019 · 1912
Earlier work this paper cites.
Identity-aware textual-visual matching with latent co-attention
Li, S., Xiao, T., Li, H., Yang, W., Wang, X., 2017 · 1917
Earlier work this paper cites.
Variational structured semantic inference for diverse image captioning, in: NeurIPS, pp. 1931–1941
Chen, F., Ji, R., Ji, J., Sun, X., Zhang, B., Ge, X., Wu, Y., Huang, F., Wang, Y., 2019a · 1941
Earlier work this paper cites.
Stackgan++: Realistic image synthesis with stacked generative adversarial networks
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., Metaxas, D.N., 2019b · 1962
Earlier work this paper cites.
An analysis of visual question answering algorithms
Kafle, K., Kanan, C., 2017 · 1991
Earlier work this paper cites.
Multitask learning
Caruana, R., 1997 · 1997
Earlier work this paper cites.
Murel: Multimodal relational reasoning for visual question answering
Cadène, R., Ben-younes, H., Cord, M., Thome, N., 2019a · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., Haffner, P., 1998 · 1998
Earlier work this paper cites.
Delving deeper into the decoder for video captioning
Chen, H., Li, J., Hu, X., 2020a · 2001
Earlier work this paper cites.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Qi, D., Su, L., Song, J., Cui, E.D.B., Bharti, T., Sacheti, A., 2020a · 2001
Earlier work this paper cites.
Show, recall, and tell: Image captioning with recall mechanism
Wang, L., Bai, Z., Zhang, Y., Lu, H., 2020a · 2001
Earlier work this paper cites.
Dual multi-head co-attention for multi-choice reading comprehension
Zhu, P.F., Zhao, H., Li, X., 2020b · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation, in: ACL
Papineni, K., Roukos, S., Ward, T., Zhu, W.J., 2002 · 2002
Earlier work this paper cites.
On the general value of evidence, and bilingual scene-text visual question answering
Wang, X., Liu, Y., Shen, C., Ng, C.C., Luo, C., Jin, L., Chan, C.S., van den Hengel, A., Wang, L., 2020c · 2002
Earlier work this paper cites.
Object relational graph with teacher-recommended learning for video captioning
Zhang, Z., Shi, Y., Yuan, C., Li, B., Wang, P., Hu, W., Zha, Z.J., 2020 · 2002
Earlier work this paper cites.
Counterfactual samples synthesizing for robust visual question answering
Chen, L., Yan, X., Xiao, J., Zhang, H., Pu, S., Zhuang, Y., 2020c · 2003
Earlier work this paper cites.
Say as you wish: Fine-grained control of image caption generation with abstract scene graphs
Chen, S., Jin, Q., Wang, P., Wu, Q., 2020d · 2003
Earlier work this paper cites.
Normalized and geometry-aware self-attention network for image captioning
Guo, L., Liu, J., Zhu, X., Yao, P., Lu, S., Lu, H.Q., 2020b · 2003
Earlier work this paper cites.
Spatio-temporal graph for video captioning with knowledge distillation
Pan, B., Cai, H., Huang, D.A., Lee, K.H., Gaidon, A., Adeli, E., Niebles, J.C., 2020a · 2003
Earlier work this paper cites.
X-linear attention networks for image captioning
Pan, Y., Yao, T., Li, Y., Mei, T., 2020b · 2003
Earlier work this paper cites.
Show, edit and tell: A framework for editing image captions
Sammani, F., Melas-Kyriazi, L., 2020 · 2003
Earlier work this paper cites.
Environment-agnostic multitask learning for natural language grounded navigation
Wang, X., Jain, V., Ie, E., Wang, W.Y., Kozareva, Z., Ravi, S., 2020b · 2003
Earlier work this paper cites.
Multi-view learning for vision-and-language navigation
Xia, Q., Li, X., Li, C., Bisk, Y., Sui, Z., Choi, Y., Smith, N.A., 2020 · 2003
Earlier work this paper cites.
Sub-instruction aware vision-and-language navigation
Hong, Y., Rodriguez-Opazo, C., Wu, Q., Gould, S., 2020 · 2004
Earlier work this paper cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
Krantz, J., Wijmans, E., Majumdar, A., Batra, D., Lee, S., 2020 · 2004
Earlier work this paper cites.
Mcqa: Multimodal co-attention based network for question answering
Kumar, A., Mittal, T., Manocha, D., 2020 · 2004
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X., Yin, X., Li, C., Hu, X., Zhang, P., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., Gao, J., 2020c · 2004
Earlier work this paper cites.
Context-aware group captioning via self-attention and contrastive features
Li, Z., Tran, Q.H., Mai, L., Lin, Z., Yuille, A.L., 2020d · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries, in: ACL 2004
Lin, C.Y., 2004 · 2004
Earlier work this paper cites.
lambert: Language and action learning using multimodal bert
Miyazawa, K., Aoki, T., Horii, T., Nagai, T., 2020 · 2004
Earlier work this paper cites.
Audio-visual based emotion recognition - a new approach
Song, M., Bu, J., Chen, C., Li, N., 2004 · 2004
Earlier work this paper cites.
Transform and tell: Entity-aware news image captioning
Tran, A., Mathews, A.P., Xie, L., 2020 · 2004
Earlier work this paper cites.
Vd-bert: A unified vision and dialog transformer with bert
Wang, Y., Joty, S.R., Lyu, M.R., King, I., Xiong, C., Hoi, S.C.H., 2020d · 2004
Earlier work this paper cites.
More grounded image captioning by distilling image-text matching model
Zhou, Y., Wang, M., Liu, D., Hu, Z., Zhang, H., 2020b · 2004
Earlier work this paper cites.
Clue: Cross-modal coherence modeling for caption generation
Alikhani, M., Sharma, P., Li, S.J., Soricut, R., Stone, M.B., 2020 · 2005
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments, in: IEEvaluation@ACL
Banerjee, S., Lavie, A., 2005 · 2005
Earlier work this paper cites.
Fashionbert: Text and image matching with adaptive loss for cross-modal retrieval
Gao, D., Jin, L., Chen, B., Qiu, M., Wei, Y., Hu, Y., Wang, H.M., 2020 · 2005
Earlier work this paper cites.
Non-autoregressive image captioning with counterfactuals-critical multi-agent learning
Guo, L., Liu, J., Zhu, X., He, X., Jiang, J., Lu, H.Q., 2020a · 2005
Earlier work this paper cites.
Misa: Modality-invariant and-specific representations for multimodal sentiment analysis
Hazarika, D., Zimmermann, R., Poria, S., 2020 · 2005
Earlier work this paper cites.
A better use of audio-visual cues: Dense video captioning with bi-modal transformer
Iashin, V., Rahtu, E., 2020 · 2005
Earlier work this paper cites.
Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning
Lei, J., Wang, L., Shen, Y., Yu, D., Berg, T.L., Bansal, M., 2020 · 2005
Earlier work this paper cites.
Explainable deep learning models in medical image analysis
Singh, A., Sengupta, S., Lakshminarayanan, V., 2020 · 2005
Earlier work this paper cites.
Cobra: Contrastive bi-modal representation algorithm
Udandarao, V., Maiti, A., Srivatsav, D., Vyalla, S.R., Yin, Y., Shah, R.R., 2020 · 2005
Earlier work this paper cites.
C3vqg: Category consistent cyclic visual question generation
Uppal, S., Madan, A., Bhagat, S., Yu, Y., Shah, R.R., 2020 · 2005
Earlier work this paper cites.
Cross-modality relevance for reasoning on language and vision
Zheng, C., Guo, Q., Kordjamshidi, P., 2020a · 2005
Earlier work this paper cites.
Large-scale adversarial training for vision-and-language representation learning
Gan, Z., Chen, Y.C., Li, L., Zhu, C., Cheng, Y., jing Liu, J., 2020 · 2006
Earlier work this paper cites.
Contrastive learning for weakly supervised phrase grounding
Gupta, T., Vahdat, A., Chechik, G., Yang, X., Kautz, J., Hoiem, D., 2020 · 2006
Earlier work this paper cites.
Counterfactual vqa: A cause-effect look at language bias
Niu, Y., Tang, K., Zhang, H., Lu, Z., Hua, X.S., Wen, J.R., 2020 · 2006
Earlier work this paper cites.
Improving image captioning with better use of captions
Shi, Z., Zhou, X., Qiu, X., Zhu, X., 2020b · 2006
Earlier work this paper cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Yu, F., Tang, J., Yin, W., Sun, Y., Tian, H., Wu, H., Wang, H., 2020b · 2006
Earlier work this paper cites.
Training combination strategy of multi-stream fused hidden markov model for audio-visual affect recognition, in: in MM
Zeng, Z., Hu, Y., Liu, M., Fu, Y., Huang, T.S., 2006 · 2006
Earlier work this paper cites.
Visual question answering as a multi-task problem
Pollard, A.E., Shapiro, J.L., 2020 · 2007
Earlier work this paper cites.
Contrastive visual-linguistic pretraining
Shi, L., Shuang, K., Geng, S., Su, P., Jiang, Z., Gao, P., Fu, Z., de Melo, G., Su, S., 2020a · 2007
Earlier work this paper cites.
Semantic equivalent adversarial data augmentation for visual question answering
Tang, R., Ma, C., Zhang, W., Wu, Q., Yang, X., 2020 · 2007
Earlier work this paper cites.
Nvae: A deep hierarchical variational autoencoder
Vahdat, A., Kautz, J., 2020 · 2007
Earlier work this paper cites.
A multilevel fusion approach for audiovisual emotion recognition, in: AVSP
Chetty, G., Wagner, M., 2008 · 2008
Earlier work this paper cites.
X-lxmert: Paint, caption and answer questions with multi-modal transformers
Cho, J., Lu, J., Schwenk, D., Hajishirzi, H., Kembhavi, A., 2020 · 2009
Earlier work this paper cites.
R-Precision. Springer US
Craswell, N., 2009 · 2009
Earlier work this paper cites.
Multimodal fusion for multimedia analysis: A survey
Atrey, P., Hossain, M., El Saddik, A., Kankanhalli, M., 2010 · 2010
Earlier work this paper cites.
Classification of affects using head movement, skin color features and physiological signals, in: SMC, pp. 2664–2669
Monkaresi, H., Hussain, M.S., Calvo, R.A., 2012 · 2012
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines, in: NIPS, pp. 2222–2230
Srivastava, N., Salakhutdinov, R.R., 2012 · 2012
Earlier work this paper cites.
Interactive search in image retrieval: a survey
Thomee, B., Lew, M.S., 2012 · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model, in: NIPS
Frome, A., Corrado, G.S., Shlens, J., Bengio, S., Dean, J., Ranzato, M., Mikolov, T., 2013 · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A., 2013 · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics (extended abstract)
Hodosh, M., Young, P., Hockenmaier, J., 2013 · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J., 2013 · 2013
Cited alongside, same era.
Utterance-level multimodal sentiment analysis, in: ACL
Pérez-Rosas, V., Mihalcea, R., Morency, L.P., 2013 · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation, in: EMNLP
Cho, K., van Merrienboer, B., Çaglar Gülçehre, Bahdanau, D., Bougares, F., Schwenk, H., Bengio, Y., 2014 · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R.B., Donahue, J., Darrell, T., Malik, J., 2014 · 2014
Cited alongside, same era.
Flipdial: A generative model for two-way visual dialogue
Massiceti, D., Siddharth, N., Dokania, P.K., Torr, P.H.S., 2018 · 2018
Later among the works it cites.
Semstyle: Learning to generate stylised image captions using unaligned text
Mathews, A.P., Xie, L., He, X., 2018 · 2018
Later among the works it cites.
Visual text correction, in: ECCV
Mazaheri, A., Shah, M., 2018 · 2018
Later among the works it cites.
Did the model understand the question?
Mudrakarta, P.K., Taly, A., Sundararajan, M., Dhamdhere, K., 2018 · 2018
Later among the works it cites.
Nagarajan, T., Grauman, K., 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving image-sentence embeddings using large weakly annotated photo collections, in: ECCV
Gong, Y., Wang, L., Hodosh, M., Hockenmaier, J., Lazebnik, S., 2014 · 2014
Cited alongside, same era.
Explaining and harnessing adversarial examples
Goodfellow, I.J., Shlens, J., Szegedy, C., 2014 · 2014
Cited alongside, same era.
Graves, A., Wayne, G., Danihelka, I., 2014 · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Kiros, R., Salakhutdinov, R., Zemel, R., 2014 · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Lin, T.Y., Maire, M., Belongie, S.J., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L., 2014 · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation, in: EMNLP, pp. 1532–1543
Pennington, J., Socher, R., Manning, C., 2014 · 2014
Cited alongside, same era.
Feature analysis for computational personality recognition using youtube personality data set, in: WCPR, p. 11–14
Sarkar, C., Bhatia, S., Agarwal, A., Li, J., 2014 · 2014
Cited alongside, same era.
Nguyen, D.K., Okatani, T., 2018 · 2018
Later among the works it cites.
Learning conditioned graph structures for interpretable visual question answering, in: NeurIPS
Norcliffe-Brown, W., Vafeias, E., Parisot, S., 2018 · 2018
Later among the works it cites.
Attention u-net: Learning where to look for the pancreas
Oktay, O., Schlemper, J., Folgoc, L.L., Lee, M.C.H., Heinrich, M.P., Misawa, K., Mori, K., McDonagh, S.G., Hammerla, N.Y., Kainz, B., Glocker, B., Rueckert, D., 2018 · 2018
Later among the works it cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
Park, D.H., Hendricks, L.A., Akata, Z., Rohrbach, A., Schiele, B., Darrell, T., Rohrbach, M., 2018 · 2018
Later among the works it cites.
Multimodal differential network for visual question generation, in: EMNLP
Patro, B.N., Kumar, S., Kurmi, V.K., Namboodiri, V., 2018 · 2018
Later among the works it cites.
Multimodal sentiment analysis, in: Socio-Affective Computing
Poria, S., Hussain, A., Cambria, E., 2018 · 2018
Later among the works it cites.
A unified framework for multimodal domain adaptation, in: ACM MM, p. 429–437
Qi, F., Yang, X., Xu, C., 2018 · 2018
Later among the works it cites.
Overcoming language priors in visual question answering with adversarial regularization
Ramakrishnan, S., Agrawal, A., Lee, S., 2018 · 2018
Later among the works it cites.
Nneval: Neural network based evaluation metric for image captioning, in: ECCV
Sharif, N., White, L., Bennamoun, M., Shah, S.A.A., 2018 · 2018
Later among the works it cites.
Text2scene: Generating abstract scenes from textual descriptions
Tan, F., Feng, S., Ordonez, V., 2018 · 2018
Later among the works it cites.
Object ordering with bidirectional matchings for visual reasoning
Tan, H., Bansal, M., 2018 · 2018
Later among the works it cites.
Hermitian co-attention networks for text matching in asymmetrical domains, in: in IJCAI
Tay, Y., Luu, A.T., Hui, S.C., 2018 · 2018
Later among the works it cites.
Interpretable counting for visual question answering
Trott, A., Xiong, C., Socher, R., 2018 · 2018
Later among the works it cites.
Recent advances in autoencoder-based representation learning
Tschannen, M., Bachem, O., Lucic, M., 2018 · 2018
Later among the works it cites.
Gated hierarchical attention for image captioning, in: ICCV
Wang, Q., Chan, A.B., 2018 · 2018
Later among the works it cites.
Explainable social contextual image recommendation with hierarchical attention
Wu, L., Ge, Y., Liu, Q., Chen, E., Hong, R., Wang, M., Du, J., 2018 · 2018
Later among the works it cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., He, X., 2018 · 2018
Later among the works it cites.
Co-attention based neural network for source-dependent essay scoring, in: Workshop on Innovative Use of NLP for Building Educational Applications
Zhang, H., Litman, D., 2018 · 2018
Later among the works it cites.
A multi-task learning approach for image captioning, in: IJCAI, p. 1205–1211
Zhao, W., Wang, B., Ye, J., Yang, M., Zhao, Z., Luo, R., Qiao, Y., 2018 · 2018
Later among the works it cites.
Compact and efficient multitask learning in vision, language and speech, in: ICCV Workshop, pp. 2933–2942
Al-Rawi, M., Valveny, E., 2019 · 2019
Later among the works it cites.
Fusion of detected objects in text for visual question answering, in: EMNLP/IJCNLP
Alberti, C., Ling, J., Collins, M., Reitter, D., 2019 · 2019
Later among the works it cites.
Sequential latent spaces for modeling the intention during diverse image captioning
Aneja, J., Agrawal, H., Batra, D., Schwing, A.G., 2019 · 2019
Later among the works it cites.
Multimodal machine learning: A survey and taxonomy
Baltrušaitis, T., Ahuja, C., Morency, L., 2019 · 2019
Later among the works it cites.
Modality-based factorization for multimodal fusion, in: RepL4NLP@ACL 2019, pp. 260–269
Barezi, E.J., Fung, P., · 2019
Later among the works it cites.
Why does a visual question have different answers?
Bhattacharya, N., Li, Q., Gurari, D., 2019 · 2019
Later among the works it cites.
Scene text visual question answering
Biten, A.F., Tito, R., Mafla, A., Gómez, L., Rusiñol, M., Valveny, E., Jawahar, C.V., Karatzas, D., 2019 · 2019
Later among the works it cites.
Latent variable model for multi-modal translation, in: ACL, pp. 6392–6405
Calixto, I., Rios, M., Aziz, W., 2019 · 2019
Later among the works it cites.
Image captioning with unseen objects, in: BMVC
Demirel, B., Cinbis, R.G., Ikizler-Cinbis, N., 2019 · 2019
Later among the works it cites.
Fast, diverse and accurate image captioning guided by part-of-speech
Deshpande, A., Aneja, J., Wang, L., Schwing, A.G., Forsyth, D.A., 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, in: in NAACL-HLT
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2019 · 2019
Later among the works it cites.
Compact trilinear interaction for visual question answering, in: ICCV
Do, T., Do, T.T., Tran, H., Tjiputra, E., Tran, Q.D., 2019 · 2019
Later among the works it cites.
Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction
El-Nouby, A., Sharma, S., Schulz, H., Hjelm, D., Asri, L.E., Kahou, S.E., Bengio, Y., W.Taylor, G., 2019 · 2019
Later among the works it cites.
Sodeep: A sorting deep net to learn ranking loss surrogates
Engilberge, M., Chevallier, L., Pérez, P., Cord, M., 2019 · 2019
Later among the works it cites.
Bridging by word: Image grounded vocabulary construction for visual captioning, in: ACL
Fan, Z., Wei, Z., Wang, S., Huang, X., 2019 · 2019
Later among the works it cites.
Modularized textual grounding for counterfactual resilience
Fang, Z., Kong, S., Fowlkes, C.C., Yang, Y., 2019 · 2019
Later among the works it cites.
Unsupervised image captioning
Feng, Y., Ma, L., Liu, W., Luo, J., 2019 · 2019
Later among the works it cites.
Exploring overall contextual information for image captioning in human-like cognitive style
Ge, H., Yan, Z., Zhang, K., Zhao, M., Sun, L., 2019 · 2019
Later among the works it cites.
Unpaired image captioning via scene graph alignments
Gu, J., Joty, S.R., Cai, J., Zhao, H., Yang, X., Wang, G., 2019 · 2019
Later among the works it cites.
Image-question-answer synergistic network for visual dialog
Guo, D., Xu, C., Tao, D., 2019 · 2019
Later among the works it cites.
Visual concept-metaconcept learning, in: NeurIPS
Han, C.W., Mao, J., Gan, C., Tenenbaum, J., jun Wu, J., 2019 · 2019
Later among the works it cites.
UR-FUNNY: A multimodal language dataset for understanding humor, in: EMNLP-IJCNLP
Hasan, M.K., Rahman, W., Bagher Zadeh, A., Zhong, J., Tanveer, M.I., Morency, L.P., Hoque, M.E., 2019 · 2019
Later among the works it cites.
It’s not about the journey; it’s about the destination: Following soft paths under question-guidance for visual reasoning, in: CVPR
Haurilet, M., Roitberg, A., Stiefelhagen, R., 2019 · 2019
Later among the works it cites.
Image captioning: Transforming objects into words, in: NeurIPS, pp. 11137–11147
Herdade, S., Kappeler, A., Boakye, K., Soares, J., 2019 · 2019
Later among the works it cites.
Dense multimodal fusion for hierarchically joint representation, in: ICASSP, pp. 3941–3945
Hu, D., Wang, C., Nie, F., Li, X., 2019 · 2019
Later among the works it cites.
Challenges and prospects in vision and language research
Kafle, K., Shrestha, R., Kanan, C., 2019 · 2019
Later among the works it cites.
Multimodal explanations by predicting counterfactuality in videos
Kanehira, A., Takemoto, K., Inayoshi, S., Harada, T., 2019 · 2019
Later among the works it cites.
Dense relational captioning: Triple-stream networks for relationship-based captioning
Kim, D.J., Choi, J., Oh, T.H., Kweon, I.S., 2019 · 2019
Later among the works it cites.
Improving visual question answering by referring to generated paragraph captions, in: ACL
Kim, H., Bansal, M., 2019 · 2019
Later among the works it cites.
Attention is (not) all you need for commonsense reasoning, in: ACL
Klein, T., Nabi, M., 2019 · 2019
Later among the works it cites.
Towards unsupervised image captioning with shared multimodal embeddings
Laina, I., Rupprecht, C., Navab, N., 2019 · 2019
Later among the works it cites.
End-to-end video captioning with multitask reinforcement learning
Li, L., Gong, B., 2019 · 2019
Later among the works it cites.
Tab-vcr: Tags and attributes based visual commonsense reasoning baselines
Lin, J., Jain, U., Schwing, A.G., 2019 · 2019
Later among the works it cites.
Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing, in: ACL, pp. 481–492
Mai, S., Hu, H., Xing, S., 2019 · 2019
Later among the works it cites.
Explicit bias discovery in visual question answering models
Manjunatha, V., Saini, N., Davis, L., 2019 · 2019
Later among the works it cites.
Mode seeking generative adversarial networks for diverse image synthesis
Mao, Q., Lee, H.Y., Tseng, H.Y., Ma, S., Yang, M.H., 2019 · 2019
Later among the works it cites.
Ocr-vqa: Visual question answering by reading text in images, in: ICDAR, pp. 947–952
Mishra, A., Shekhar, S., Singh, A.K., Chakraborty, A., 2019 · 2019
Later among the works it cites.
Trends in integration of vision and language research: A survey of tasks, datasets, and methods
Mogadala, A., Kalimuthu, M., Klakow, D., 2019 · 2019
Later among the works it cites.
Multi-task learning of hierarchical vision-language representation
Nguyen, D.K., Okatani, T., 2019 · 2019
Later among the works it cites.
Recursive visual attention in visual dialog, in: CVPR
Niu, Y., Zhang, H., Zhang, M., Zhang, J., Lu, Z., Wen, J.R., 2019 · 2019
Later among the works it cites.
Mfas: Multimodal fusion architecture search
Pérez-Rúa, J.M., Vielzeuf, V., Pateux, S., Baccouche, M., Jurie, F., 2019 · 2019
Later among the works it cites.
Found in translation: Learning robust joint representations by cyclic translations between modalities, in: AAAI, pp. 6892–6899
Pham, H., Liang, P.P., Manzini, T., Morency, L., Póczos, B., 2019 · 2019
Later among the works it cites.
Mirrorgan: Learning text-to-image generation by redescription
Qiao, T., Zhang, J., Xu, D., Tao, D., 2019 · 2019
Later among the works it cites.
Look back and predict forward in image captioning
Qin, Y., Du, J., Zhang, Y., Lu, H., 2019 · 2019
Later among the works it cites.
Watch, listen and tell: Multi-modal weakly supervised dense event captioning
Rahman, T., Xu, B., Sigal, L., 2019 · 2019
Later among the works it cites.
Towards explainable artificial intelligence, in: Explainable AI
Samek, W., Müller, K.R., 2019 · 2019
Later among the works it cites.
Learning to caption images through a lifetime by asking questions
Shen, T., Kar, A., Fidler, S., 2019 · 2019
Later among the works it cites.
Engaging image captioning via personality
Shuster, K., Humeau, S., Hu, H., Bordes, A., Weston, J., 2019 · 2019
Later among the works it cites.
Unsupervised multi-modal neural machine translation
Su, Y., Fan, K., Bach, N., Kuo, C.C.J., Huang, F., 2019 · 2019
Later among the works it cites.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., Zhou, S., Zhang, I.D., Bai, H., Artzi, Y., 2019 · 2019
Later among the works it cites.
Learning to compose dynamic tree structures for visual contexts
Tang, K., Zhang, H., Wu, B., Luo, W., Liu, W., 2019 · 2019
Later among the works it cites.
Audio-visual interpretable and controllable video captioning, in: CVPR Workshops
Tian, Y., Guan, C., Goodman, J., Moore, M., Xu, C., 2019 · 2019
Later among the works it cites.
Probabilistic neural-symbolic models for interpretable visual question answering, in: ICML
Vedantam, R., Desai, K., Lee, S., Rohrbach, M., Batra, D., Parikh, D., 2019 · 2019
Later among the works it cites.
Joint optimization for cooperative image captioning
Vered, G., Oren, G., Atzmon, Y., Chechik, G., 2019 · 2019
Later among the works it cites.
Composing text and image for image retrieval - an empirical odyssey
Vo, N.S., Jiang, L., Sun, C., Murphy, K., Li, L.J., Fei-Fei, L., Hays, J., 2019 · 2019
Later among the works it cites.
Multitask learning for cross-domain image captioning
Yang, M., Zhao, W., Xu, W., Feng, Y., Zhao, Z., Chen, X., Lei, K., 2019 · 2019
Later among the works it cites.
Hierarchy parsing for image captioning
Yao, T., Pan, Y., Li, Y., Mei, T., 2019 · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning, in: CVPR
Zellers, R., Bisk, Y., Farhadi, A., Choi, Y., 2019 · 2019
Later among the works it cites.
Spatiotemporal-textual co-attention network for video question answering
Zha, Z.J., Liu, J., Yang, T., Zhang, Y., 2019 · 2019
Later among the works it cites.
Intention oriented image captions with guiding objects
Zheng, Y., Li, Y., Wang, S., 2019 · 2019
Later among the works it cites.
A study on multimodal and interactive explanations for visual question answering, in: SafeAI@AAAI
Alipour, K., Schulze, J.P., Yao, Y., Ziskind, A., Burachas, G., 2020 · 2020
Closest in time.
Disentangling multiple features in video sequences using gaussian processes in variational autoencoders
Bhagat, S., Uppal, S., Yin, V.T., Lim, N., 2020 · 2020
Closest in time.
Leaf-qa: Locate, encode & attend for figure question answering, in: WACV
Chaudhry, R., Shekhar, S., Gupta, U., Maneriker, P., Bansal, P., Joshi, A., 2020 · 2020
Closest in time.
Just ask: An interactive learning framework for vision and language navigation, in: AAAI
Chi, T.C., Eric, M., Kim, S., Shen, M., Hakkani-Tür, D.Z., 2020 · 2020
Closest in time.
Crhasum: extractive text summarization with contextualized-representation hierarchical-attention summarization network
Diao, Y., Lin, H., Yang, L., chao Fan, X., Chu, Y., Wu, D., Zhang, D., Xu, K., 2020 · 2020
Closest in time.
Knowit vqa: Answering knowledge-based questions about videos, in: AAAI
Garcia, N., Otani, M., Chu, C., Nakashima, Y., 2020 · 2020
Closest in time.
Towards learning a generic agent for vision-and-language navigation via pre-training
Hao, W., Li, C., Li, X., Carin, L., Gao, J., 2020 · 2020
Closest in time.
Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!
Hessel, J., Lee, L., 2020 · 2020
Closest in time.
Unsupervised multimodal neural machine translation with pseudo visual pivoting, in: ACL
Huang, P.Y., Hu, J., Chang, X., Hauptmann, A.G., 2020 · 2020
Closest in time.
Answering questions about data visualizations using efficient bimodal fusion, in: WACV
Kafle, K., Shrestha, R., Cohen, S., Price, B., Kanan, C., 2020 · 2020
Closest in time.
Vision and language: from visual perception to content creation
Mei, T., Zhang, W., Yao, T., 2020 · 2020
Closest in time.
Bias in multimodal AI: testbed for fair automatic recruitment, in: CVPR Workshops, pp. 129–137
Peña, A., Serna, I., Morales, A., Fiérrez, J., 2020 · 2020
Closest in time.
Reinforcing an image caption generator using off-line human feedback, in: AAAI
Seo, P.H., Sharma, P., Levinboim, T., Han, B., Soricut, R., 2020 · 2020
Closest in time.
Ha-ccn: Hierarchical attention-based crowd counting network
Sindagi, V., Patel, V.M., 2020 · 2020
Closest in time.
Using image captions and multitask learning for recommending query reformulations
Verma, G., Vinay, V., Bansal, S., Oberoi, S., Sharma, M., Gupta, P., 2020 · 2020
Closest in time.
Gamma: A graph and multi-view memory attention mechanism for top-n heterogeneous recommendation, in: Advances in Knowledge Discovery and Data Mining, pp. 28–40
Vijaikumar, M., Shevade, S., Narasimha Murty, M., 2020 · 2020
Closest in time.
Explainable deep learning: A field guide for the uninitiated
Xie, N., Ras, G., Gerven, M.V., Doran, D., 2020 · 2020
Closest in time.
Fooled by imagination: Adversarial attack to image captioning via perturbation in complex domain, in: ICME, pp. 1–6
Zhang, S., Wang, Z., Xu, X., Guan, X., Yang, Y., 2020 · 2020
Closest in time.
Memcap: Memorizing style knowledge for image captioning, in: AAAI
Zhao, W., Wu, X., Zhang, X., 2020 · 2020
Closest in time.
Factor graph attention
Schwartz, I., Yu, S., Hazan, T., Schwing, A.G., 2019 · 2048
Closest in time.
Learning fragment self-attention embeddings for image-text matching, in: ACM MM, p. 2088–2096
Wu, Y., Wang, S., Song, G., Huang, Q., 2019c · 2096
Closest in time.