Fetching the paper…
Reading the bibliography…
The task of dynamic scene graph generation (SGG) from videos is complicated and challenging due to the inherent dynamics of a scene, temporal fluctuation of model predictions, and the long-tailed distribution of the visual relationships in addition to the already existing challenges in image-based SGG.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
Armen Der Kiureghian and Ove Ditlevsen · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D. Manning · 2015
Earlier work this paper cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Meta-learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Deep bayesian active learning with image data
Yarin Gal, Riashat Islam, and Zoubin Ghahramani · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Learning to remember rare events
Łukasz Kaiser, Ofir Nachum, Aurko Roy, and Samy Bengio · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Scene graph generation from objects, phrases and region captions
Yikang Li, Wanli Ouyang, Bolei Zhou, Kun Wang, and Xiaogang Wang · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
The perception of causality
Albert Michotte · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2017
Cited alongside, same era.
Object level visual reasoning in videos
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille, and Greg Mori · 2018
Cited alongside, same era.
Uncertainty-aware learning from demonstration using mixture density networks with sampling-free variance modeling
Sungjoon Choi, Kyungjae Lee, Sungbin Lim, and Songhwai Oh · 2018
Inflated episodic memory with region self-attention for long-tailed visual recognition
Linchao Zhu and Yi Yang · 2020
Later among the works it cites.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Later among the works it cites.
Estimating and exploiting the aleatoric uncertainty in surface normal estimation
Gwangbin Bae, Ignas Budvytis, and Roberto Cipolla · 2021
Later among the works it cites.
Active learning for deep object detection via probabilistic modeling
Jiwoong Choi, Ismail Elezi, Hyuk-Jae Lee, Clement Farabet, and Jose M Alvarez · 2021
Later among the works it cites.
Spatial-Temporal Transformer for Dynamic Scene Graph Generation
Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, and Michael Ying Yang · 2021
Later among the works it cites.
Learning of visual relations: The devil is in the tails
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Image captioning with scene-graph based semantic concepts
Lizhao Gao, Bo Wang, and Wenmin Wang · 2018
Cited alongside, same era.
Dynamic few-shot visual learning without forgetting
Spyros Gidaris and Nikos Komodakis · 2018
Cited alongside, same era.
Knowledge-embedded routing network for scene graph generation
Tianshui Chen, Weihao Yu, Riquan Chen, and Liang Lin · 2019
Cited alongside, same era.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Single-model uncertainties for deep learning
Natasa Tagasovska and David Lopez-Paz · 2019
Cited alongside, same era.
Alakh Desai, Tz-Ying Wu, Subarna Tripathi, and Nuno Vasconcelos · 2021
Later among the works it cites.
A simple baseline for weakly-supervised human-centric relation detection
Raghav Goyal12 and Leonid Sigal123 · 2021
Later among the works it cites.
Detecting human-object relationships in videos
Jingwei Ji, Rishi Desai, and Juan Carlos Niebles · 2021
Later among the works it cites.
Bipartite graph network with adaptive message passing for unbiased scene graph generation
Rongjie Li, Songyang Zhang, Bo Wan, and Xuming He · 2021
Later among the works it cites.
Activity graph transformer for temporal action localization
Megha Nawhal and Greg Mori · 2021
Later among the works it cites.
Target adaptive context aggregation for video scene graph generation
Yao Teng, Limin Wang, Zhifeng Li, and Gangshan Wu · 2021
Later among the works it cites.
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al · 2022
Later among the works it cites.
Uncertainty-aware learning against label noise on imbalanced datasets
Yingsong Huang, Bing Bai, Shengwei Zhao, Kun Bai, and Fei Wang · 2022
Later among the works it cites.
Prompting visual-language models for efficient video understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie · 2022
Later among the works it cites.
Iterative scene graph generation
Siddhesh Khandelwal and Leonid Sigal · 2022
Later among the works it cites.
Sgtr: End-to-end scene graph generation with transformer
Rongjie Li, Songyang Zhang, and Xuming He · 2022
Later among the works it cites.
Ppdl: Predicate probability distribution based loss for unbiased scene graph generation
Wei Li, Haiwei Zhang, Qijie Bai, Guoqing Zhao, Ning Jiang, and Xiaojie Yuan · 2022
Later among the works it cites.
Rethinking the evaluation of unbiased scene graph generation
Xingchen Li, Long Chen, Jian Shao, Shaoning Xiao, Songyang Zhang, and Jun Xiao · 2022
Later among the works it cites.
Dynamic scene graph generation via anticipatory pre-training
Yiming Li, Xiaoshan Yang, and Changsheng Xu · 2022
Later among the works it cites.
Long-tail recognition via compositional knowledge transfer
Sarah Parisot, Pedro M Esperança, Steven McDonagh, Tamas J Madarasz, Yongxin Yang, and Zhenguo Li · 2022
Later among the works it cites.
Dynamic scene graph generation via temporal prior inference
Shuang Wang, Lianli Gao, Xinyu Lyu, Yuyu Guo, Pengpeng Zeng, and Jingkuan Song · 2022
Later among the works it cites.