Fetching the paper…
Reading the bibliography…
We introduce CLEVR-Math, a multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction, represented partly by a textual description and partly by an image illustrating the scenario.
A multi-world approach to question answering about real-world scenes based on uncertain input,
M. Malinowski, M. Fritz, · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context,
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, C. L. Zitnick, · 2014
Earlier work this paper cites.
Deep neural language models for machine translation,
M.-T. Luong, M. Kayser, C. D. Manning, · 2015
Earlier work this paper cites.
Vqa: Visual question answering,
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, D. Parikh, · 2015
Earlier work this paper cites.
Mawps: A math word problem repository,
R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, H. Hajishirzi, · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding,
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, M. Rohrbach, · 2016
Earlier work this paper cites.
Where to look: Focus regions for visual question answering,
K. J. Shih, S. Singh, D. Hoiem, · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, R. Girshick, · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems,
W. Ling, D. Yogatama, C. Dyer, P. Blunsom, · 2017
Earlier work this paper cites.
Mutan: Multimodal tucker fusion for visual question answering,
H. Ben-Younes, R. Cadene, M. Cord, N. Thome, · 2017
Earlier work this paper cites.
Fvqa: Fact-based visual question answering,
P. Wang, Q. Wu, C. Shen, A. Dick, A. Van Den Hengel, · 2017
Cited alongside, same era.
Mask r-cnn. corr abs/1703.06870,
K. He, G. Gkioxari, P. Dollár, R. Girshick, · 2017
Cited alongside, same era.
Out of the box: Reasoning with graph convolution nets for factual visual question answering,
M. Narasimhan, S. Lazebnik, A. Schwing, · 2018
Cited alongside, same era.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding,
K. Yi, J. Wu, C. Gan, A. Torralba, P. Kohli, J. Tenenbaum, · 2018
Cited alongside, same era.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks,
B. Lake, M. Baroni, · 2018
Cited alongside, same era.
Graph-to-tree learning for solving math word problems,
J. Zhang, L. Wang, R. K.-W. Lee, Y. Bin, Y. Wang, J. Shao, E.-P. Lim, · 2020
Later among the works it cites.
Learning 3d semantic scene graphs from 3d indoor reconstructions,
J. Wald, H. Dhamo, N. Navab, F. Tombari, · 2020
Later among the works it cites.
Action genome: Actions as compositions of spatio-temporal scene graphs,
J. Ji, R. Krishna, L. Fei-Fei, J. C. Niebles, · 2020
Later among the works it cites.
Are nlp models really able to solve simple math word problems?,
A. Patel, S. Bhattamishra, N. Goyal, · 2021
Later among the works it cites.
Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations,
W. Stammer, P. Schramowski, K. Kersting, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Xie, S. Sun, · 2019
Cited alongside, same era.
Clevrer: Collision events for video representation and reasoning,
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, J. B. Tenenbaum, · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering,
D. A. Hudson, C. D. Manning, · 2019
Cited alongside, same era.
Kandinsky patterns as iq-test for machine learning,
A. Holzinger, M. Kickmeier-Rust, H. Müller, · 2019
Cited alongside, same era.
J. Mao, C. Gan, P. Kohli, J. B. Tenenbaum, J. Wu, · 2019
Cited alongside, same era.
S. K. Sampat, A. Kumar, Y. Yang, C. Baral, · 2021
Later among the works it cites.
Graph learning: A survey,
F. Xia, K. Sun, S. Yu, A. Aziz, L. Wan, S. Pan, H. Liu, · 2021
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graphs,
F. Yu, J. Tang, W. Yin, Y. Sun, H. Tian, H. Wu, H. Wang, · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., · 2021
Later among the works it cites.
Winoground: Probing vision and language models for visio-linguistic compositionality,
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, C. Ross, · 2022
Closest in time.