Fetching the paper…
Reading the bibliography…
Visual reasoning tasks such as visual question answering (VQA) require an interplay of visual perception with reasoning about the question semantics grounded in perception.
Lee, K., Palangi, H., Chen, X., Hu, H., and Gao, J · 1909
Earlier work this paper cites.
Probabilities for SV machines. advances in large margin classifiers (pp. 61–74), 2000
Platt, J · 2000
Earlier work this paper cites.
Reasoning with neural tensor networks for knowledge base completion
Socher, R., Chen, D., Manning, C. D., and Ng, A · 2013
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
VQA: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Bidirectional LSTM-CRF models for sequence tagging
Huang, Z., Xu, W., and Yu, K · 2015
Earlier work this paper cites.
Neural programmer: Inducing latent programs with gradient descent
Neelakantan, A., Le, Q. V., and Sutskever, I · 2015
Earlier work this paper cites.
Neural programmer-interpreters
Reed, S. and De Freitas, N · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2015
Earlier work this paper cites.
Simple baseline for visual question answering
Zhou, B., Tian, Y., Sukhbaatar, S., Szlam, A., and Fergus, R · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Agrawal, A., Batra, D., and Parikh, D · 2016
Earlier work this paper cites.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D · 2016
Earlier work this paper cites.
Neural symbolic machines: Learning semantic parsers on freebase with weak supervision
Liang, C., Berant, J., Le, Q., Forbus, K. D., and Lao, N · 2016
Earlier work this paper cites.
Learning executable semantic parsers for natural language understanding
Liang, P · 2016
Earlier work this paper cites.
Learning a natural language interface with neural programmer
Neelakantan, A., Le, Q. V., Abadi, M., McCallum, A., and Amodei, D · 2016
Earlier work this paper cites.
Neuro-symbolic program synthesis
Parisotto, E., Mohamed, A.-r., Singh, R., Li, L., Zhou, D., and Kohli, P · 2016
Earlier work this paper cites.
Question relevance in VQA: identifying non-visual and false-premise questions
Ray, A., Christie, G., Bansal, M., Batra, D., and Parikh, D · 2016
Cited alongside, same era.
Logic tensor networks: Deep learning and logical reasoning from data and knowledge
Serafini, L. and Garcez, A. d · 2016
Cited alongside, same era.
Making neural programming architectures generalize via recursion
Cai, J., Shin, R., and Song, D · 2017
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Cited alongside, same era.
Program synthesis
Gulwani, S., Polozov, O., Singh, R., et al · 2017
Cited alongside, same era.
Pdp: A general neural framework for learning constraint satisfaction solvers
Amizadeh, S., Matusevych, S., and Weimer, M · 2019
Later among the works it cites.
Garcez, A. d., Gori, M., Lamb, L. C., Serafini, L., Spranger, M., and Tran, S. N · 2019
Later among the works it cites.
Unicoder-VL: A universal encoder for vision and language by cross-modal pre-training
Li, G., Duan, N., Fang, Y., Gong, M., Jiang, D., and Zhou, M · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R · 2017
Cited alongside, same era.
Beta calibration: a well-founded and easily implemented improvement on logistic calibration for binary classifiers
Kull, M., Silva Filho, T., and Flach, P · 2017
Cited alongside, same era.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A · 2018
Cited alongside, same era.
Learning to solve circuit-sat: An unsupervised differentiable approach
Amizadeh, S., Matusevych, S., and Weimer, M · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L · 2018
Cited alongside, same era.
Compositional attention networks for machine reasoning
Hudson, D. A. and Manning, C. D · 2018
Cited alongside, same era.
Learning a sat solver from single-bit supervision
Selsam, D., Lamm, M., Bünz, B., Liang, P., de Moura, L., and Dill, D. L · 2018
Cited alongside, same era.
Mao, J., Gan, C., Kohli, P., Tenenbaum, J. B., and Wu, J · 2019
Later among the works it cites.
Learning compositional neural programs with recursive tree search and planning
Pierrot, T., Ligner, G., Reed, S. E., Sigaud, O., Perrin, N., Laterre, A., Kas, D., Beguir, K., and de Freitas, N · 2019
Later among the works it cites.
Complex program induction for querying knowledge bases in the absence of gold programs
Saha, A., Ansari, G. A., Laddha, A., Sankaranarayanan, K., and Chakrabarti, S · 2019
Later among the works it cites.
Explainable and explicit visual reasoning over scene graphs
Shi, J., Zhang, H., and Li, J · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M · 2019
Later among the works it cites.
Probabilistic neural-symbolic models for interpretable visual question answering
Vedantam, R., Desai, K., Lee, S., Rohrbach, M., Batra, D., and Parikh, D · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Later among the works it cites.
Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension
Chen, X., Liang, C., Yu, A. W., Zhou, D., Song, D., and Le, Q · 2020
Closest in time.
SQuINTing at VQA Models: Interrogating VQA Models with Sub-Questions
Selvaraju, R. R., Tendulkar, P., Parikh, D., Horvitz, E., Ribeiro, M., Nushi, B., and Kamar, E · 2020
Closest in time.
Analyzing differentiable fuzzy logic operators
van Krieken, E., Acar, E., and van Harmelen, F · 2020
Closest in time.
Unified vision-language pre-training for image captioning and VQA
Zhou, L., Palangi, H., Zhang, L., Hu, H., Corso, J. J., and Gao, J · 2020
Closest in time.