Fetching the paper…
Reading the bibliography…
Today's scene graph generation (SGG) task is still far from practical, mainly due to the severe training bias, e.g., collapsing diverse "human walk on / sit on / lay on beach" into "human on beach".
Edge and curve detection for visual scene analysis
A. Rosenfeld and M. Thurston · 1971
Earlier work this paper cites.
Bounded rationality
H. A. Simon · 1990
Earlier work this paper cites.
Identifiability and exchangeability for direct and indirect effects
J. M. Robins and S. Greenland · 1992
Earlier work this paper cites.
Counterfactual thinking
N. J. Roese · 1997
Earlier work this paper cites.
Causality: models, reasoning and inference
J. Pearl · 2000
Earlier work this paper cites.
Direct and indirect effects
J. Pearl · 2001
Earlier work this paper cites.
Mediation analysis
D. P. MacKinnon, A. J. Fairchild, and M. S. Fritz · 2007
Earlier work this paper cites.
A political mediation model of corporate response to social movement activism
B. G. King · 2008
Earlier work this paper cites.
Learning from imbalanced data
H. He and E. A. Garcia · 2009
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba, A. A. Efros, et al · 2011
Earlier work this paper cites.
Mediation analysis in epidemiology: methods, interpretation and bias
L. Richiardi, R. Bellocco, and D. Zugna · 2013
Earlier work this paper cites.
A three-way decomposition of a total effect into direct, indirect, and interactive effects
T. J. VanderWeele · 2013
Earlier work this paper cites.
Learning fair representations
R. Zemel, Y. Wu, K. Swersky, T. Pitassi, and C. Dwork · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Influence of resampling on accuracy of imbalanced classification
E. Burnaev, P. Erofeev, and A. Papanov · 2015
Earlier work this paper cites.
Evaluation and validation of social and psychological markers in randomised trials of complex interventions in mental health: a methodological research programme
G. Dunn, R. Emsley, H. Liu, S. Landau, J. Green, I. White, and A. Pickles · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
The statistics of causal inference: A view from political methodology
L. Keele · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
S. Schuster, R. Krishna, A. Chang, L. Fei-Fei, and C. D. Manning · 2015
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
K. S. Tai, R. Socher, and C. D. Manning · 2015
Earlier work this paper cites.
Cognitive neuroscience of human counterfactual reasoning
N. Van Hoeck, P. D. Watson, and A. K. Barbey · 2015
Earlier work this paper cites.
Explanation in causal inference: methods for mediation and interaction
T. VanderWeele · 2015
Earlier work this paper cites.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
Seeing through the human reporting bias: Visual classifiers from noisy human-centric labels
I. Misra, C. Lawrence Zitnick, M. Mitchell, and R. Girshick · 2016
Cited alongside, same era.
Causal inference in statistics: A primer
J. Pearl, M. Glymour, and N. P. Jewell · 2016
Cited alongside, same era.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Cited alongside, same era.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Cited alongside, same era.
Counterfactual fairness
M. J. Kusner, J. Loftus, C. Russell, and R. Silva · 2017
Cited alongside, same era.
Exploring visual relationship for image captioning
T. Yao, Y. Pan, Y. Li, and T. Mei · 2018
Later among the works it cites.
Zoom-net: Mining deep feature interactions for visual relationship recognition
G. Yin, L. Sheng, B. Liu, N. Yu, X. Wang, J. Shao, and C. Change Loy · 2018
Later among the works it cites.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Later among the works it cites.
Rubi: Reducing unimodal biases in visual question answering
R. Cadene, C. Dancette, H. Ben-younes, M. Cord, and D. Parikh · 2019
Later among the works it cites.
Counterfactual critic multi-agent training for scene graph generation
L. Chen, H. Zhang, J. Xiao, X. He, S. Pu, and S.-F. Chang · 2019
Later among the works it cites.
Knowledge-embedded routing network for scene graph generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scene graph generation from objects, phrases and caption regions
Y. Li, W. Ouyang, B. Zhou, K. Wang, and X. Wang · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Cited alongside, same era.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Cited alongside, same era.
Graph-structured representations for visual question answering
D. Teney, L. Liu, and A. van den Hengel · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Cited alongside, same era.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Cited alongside, same era.
T. Chen, W. Yu, R. Chen, and L. Lin · 2019
Later among the works it cites.
Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel · 2019
Later among the works it cites.
Scene graph generation with external knowledge and image reconstruction
J. Gu, H. Zhao, Z. Lin, S. Li, J. Cai, and M. Ling · 2019
Later among the works it cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Later among the works it cites.
Learning by abstraction: The neural state machine
D. A. Hudson and C. D. Manning · 2019
Later among the works it cites.
Repair: Removing representation bias by dataset resampling
Y. Li and N. Vasconcelos · 2019
Later among the works it cites.
Vrr-vg: Refocusing visually-relevant relationships
Y. Liang, Y. Bai, W. Zhang, X. Qian, L. Zhu, and T. Mei · 2019
Later among the works it cites.
Explicit bias discovery in visual question answering models
V. Manjunatha, N. Saini, and L. S. Davis · 2019
Later among the works it cites.
Causal induction from visual observations for goal directed tasks
S. Nair, Y. Zhu, S. Savarese, and L. Fei-Fei · 2019
Later among the works it cites.
Two causal principles for improving visual dialog
J. Qi, Y. Niu, J. Huang, and H. Zhang · 2019
Later among the works it cites.
Attentive relational networks for mapping images to scene graphs
M. Qi, W. Li, Z. Yang, Y. Wang, and J. Luo · 2019
Later among the works it cites.
Explainable and explicit visual reasoning over scene graphs
J. Shi, H. Zhang, and J. Li · 2019
Later among the works it cites.
Learning to compose dynamic tree structures for visual contexts
K. Tang, H. Zhang, B. Wu, W. Luo, and W. Liu · 2019
Later among the works it cites.
Exploring context and visual pattern of relationship for scene graph generation
W. Wang, R. Wang, S. Shan, and X. Chen · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
X. Yang, K. Tang, H. Zhang, and J. Cai · 2019
Later among the works it cites.
Graphical contrastive losses for scene graph parsing
J. Zhang, K. J. Shih, A. Elgammal, A. Tao, and B. Catanzaro · 2019
Later among the works it cites.
Learning filter pruning criteria for deep convolutional neural networks acceleration
Y. He, Y. Ding, P. Liu, L. Zhu, H. Zhang, and Y. Yang · 2020
Closest in time.
Counterfactual vqa: A cause-effect look at language bias
Y. Niu, K. Tang, H. Zhang, Z. Lu, X. Hua, and J.-R. Wen · 2020
Closest in time.
Visual commonsense r-cnn
W. Tan, H. Jianqiang, Z. Hanwang, and S. Qianru · 2020
Closest in time.
Deconfounded image captioning: A causal retrospect, 2020
X. Yang, H. Zhang, and J. Cai · 2020
Closest in time.