Fetching the paper…
Reading the bibliography…
NLVR2 (Suhr et al., 2019) was designed to be robust for language bias through a data collection process that resulted in each natural language sentence appearing with both true and false labels.
VisualBERT: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W. (2019) · 1908
Earlier work this paper cites.
A corpus of natural language for visual reasoning
Suhr, A., Lewis, M., Yeh, J., and Artzi, Y. (2017) · 2017
Earlier work this paper cites.
Weakly supervised semantic parsing with abstract examples
Goldman, O., Latcinnik, V., Nave, E., Globerson, A., and Berant, J. (2018) · 2018
Cited alongside, same era.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., and Artzi, Y. (2019) · 2019
Cited alongside, same era.
LXMERT: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M. (2019) · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…