Fetching the paper…
Reading the bibliography…
Modeling expressive cross-modal interactions seems crucial in multimodal tasks, such as visual question answering.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le. 2019 · 1905
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Interaction effects of picture and caption on humor ratings of cartoons
James M. Jones, Gary Alan Fine, and Robert G. Brust. 1979 · 1979
Earlier work this paper cites.
Generalized additive models: some applications
Trevor Hastie and Robert Tibshirani. 1987 · 1987
Earlier work this paper cites.
Image-music-text
Roland Barthes. 1988 · 1988
Earlier work this paper cites.
The language of displayed art
Michael O’Toole. 1994 · 1994
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire. 1995 · 1995
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire. 1995 · 1995
Earlier work this paper cites.
Multiplying meaning
Jay Lemke. 1998 · 1998
Earlier work this paper cites.
Random forests
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
Jerome H. Friedman. 2001 · 2001
Earlier work this paper cites.
Global sensitivity indices for nonlinear mathematical models and their Monte Carlo estimates
Ilya M Sobol. 2001 · 2001
Earlier work this paper cites.
A taxonomy of relationships between images and text
Emily E Marsh and Marilyn Domas White. 2003 · 2003
Earlier work this paper cites.
Discovering additive structure in black box functions
Giles Hooker. 2004 · 2004
Earlier work this paper cites.
Multimodal discourse analysis: Systemic functional perspectives
Kay O’Halloran. 2004 · 2004
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. 2020 · 2005
Earlier work this paper cites.
A system for image–text relations in new (and old) media
Radan Martinec and Andrew Salway. 2005 · 2005
Earlier work this paper cites.
Early versus late fusion in semantic video analysis
Cees G.M. Snoek, Marcel Worring, and Arnold W.M. Smeulders. 2005 · 2005
Earlier work this paper cites.
Estimating mean dimensionality of analysis of variance decompositions
Ruixue Liu and Art B. Owen. 2006 · 2006
Earlier work this paper cites.
Predictive learning via rule ensembles
Jerome H. Friedman and Bogdan E. Popescu. 2008 · 2008
Earlier work this paper cites.
An efficient explanation of individual classifications using game theory
Erik Štrumbelj and Igor Kononenko. 2010 · 2010
Earlier work this paper cites.
Large-scale visual sentiment ontology and detectors using adjective noun pairs
Damian Borth, Rongrong Ji, Tao Chen, Thomas Breuel, and Shih-Fu Chang. 2013 · 2013
Cited alongside, same era.
Text and image: A critical introduction to the visual/verbal divide
John Bateman. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Bayesian optimization of text representations
Dani Yogatama and Noah A. Smith. 2015 · 2015
Cited alongside, same era.
Lightning: large-scale linear classification, regression and ranking in Python
Mathieu Blondel and Fabian Pedregosa. 2016 · 2016
TVQA: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L. Berg. 2018 · 2018
Later among the works it cites.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Later among the works it cites.
Knowledge-rich image gist understanding beyond literal meaning
Lydia Weiland, Ioana Hulpuş, Simone Paolo Ponzetto, Wolfgang Effelsberg, and Laura Dietz. 2018 · 2018
Later among the works it cites.
Visual entailment task for visually-grounded language learning
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2018 · 2018
Later among the works it cites.
Advise: Symbolism and external knowledge for decoding advertisements
Keren Ye and Adriana Kovashka. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens Van Der Maaten. 2016 · 2016
Cited alongside, same era.
Interpretable decision sets: A joint framework for description and prediction
Himabindu Lakkaraju, Stephen H. Bach, and Jure Leskovec. 2016 · 2016
Cited alongside, same era.
Sentiment analysis on multi-view social data
Teng Niu, Shiai Zhu, Lei Pang, and Abdulmotaleb El-Saddik. 2016 · 2016
Cited alongside, same era.
Model-agnostic interpretability of machine learning
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Supersparse linear integer models for optimized medical scoring systems
Berk Ustun and Cynthia Rudin. 2016 · 2016
Cited alongside, same era.
Mingda Zhang, Rebecca Hwa, and Adriana Kovashka. 2018 · 2018
Later among the works it cites.
A co-memory network for multimodal sentiment analysis
Nan Xu, Wenji Mao, and Guandan Chen. 2018 · 2018
Later among the works it cites.
CITE: A corpus of image–text discourse relations
Malihe Alikhani, Sreyasi Nag Chowdhury, Gerard de Melo, and Matthew Stone. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Something’s brewing! Early prediction of controversy-causing posts from discussion features
Jack Hessel and Lillian Lee. 2019 · 2019
Later among the works it cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Integrating text and image: Determining multimodal document intent in instagram posts
Julia Kruk, Jonah Lubin, Karan Sikka, Xiao Lin, Dan Jurafsky, and Ajay Divakaran. 2019 · 2019
Later among the works it cites.
Analyzing compositionality of visual question answering
Sanjay Subramanian, Sameer Singh, and Matt Gardner. 2019 · 2019
Later among the works it cites.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Categorizing and inferring the relationship between the text and image of Twitter posts
Alakananda Vempala and Daniel Preoţiuc-Pietro. 2019 · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Categorizing and inferring the relationship between the text and image of Twitter posts
Alakananda Vempala and Daniel Preoţiuc-Pietro. 2019 · 2019
Later among the works it cites.
Multiplicative interactions and where to find them
Siddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz, Jack Rae, Simon Osindero, Yee Whye Teh, Tim Harley, and Razvan Pascanu. 2020 · 2020
Closest in time.