Fetching the paper…
Reading the bibliography…
Creating artificial social intelligence - algorithms that can understand the nuances of multi-person interactions - is an exciting and emerging challenge in processing facial expressions and gestures from multimodal videos.
Heterogeneous Graph Transformer
Z. Hu, Y. Dong, K. Wang, and Y. Sun · 2003
Earlier work this paper cites.
A new model for learning in graph domains
M. Gori, G. Monfardini, and F. Scarselli · 2005
Earlier work this paper cites.
The graph neural network model
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini · 2008
Earlier work this paper cites.
Action recognition by hierarchical sequence summarization
Y. Song, L.-P. Morency, and R. Davis · 2013
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Earlier work this paper cites.
Hierarchical multimodal lstm for dense visual-semantic embedding
Z. Niu, M. Zhou, L. Wang, X. Gao, and G. Hua · 2017
Earlier work this paper cites.
Tensor fusion network for multimodal sentiment analysis
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L.-P. Morency · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
J. Lei, L. Yu, M. Bansal, and T. L. Berg · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Learning factorized multimodal representations
Y.-H. H. Tsai, P. P. Liang, A. Zadeh, L.-P. Morency, and R. Salakhutdinov · 2018
Earlier work this paper cites.
P. Veličković, W. Fedus, W. L. Hamilton, P. Liò, Y. Bengio, and R. D. Hjelm · 2018
Earlier work this paper cites.
Memory fusion network for multi-view sequential learning
A. Zadeh, P. P. Liang, N. Mazumder, S. Poria, E. Cambria, and L.-P. Morency · 2018
Earlier work this paper cites.
Language, gesture, and emotional communication: An embodied view of social interaction
E. De Stefani and D. De Marco · 2019
Earlier work this paper cites.
Fast graph representation learning with pytorch geometric
M. Fey and J. E. Lenssen · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Cited alongside, same era.
Nemo: a toolkit for building ai applications using neural modules
O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V. Lavrukhin, J. Cook, et al · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2019
Cited alongside, same era.
Social-iq: A question answering benchmark for artificial social intelligence
A. Zadeh, M. Chan, P. P. Liang, E. Tong, and L.-P. Morency · 2019
Mtag: Modal-temporal attention graph for unaligned human multimodal language sequences
J. Yang, Y. Wang, R. Yi, Y. Zhu, A. Rehman, A. Zadeh, S. Poria, and L.-P. Morency · 2020
Later among the works it cites.
Graph contrastive learning with augmentations
Y. You, T. Chen, Y. Sui, T. Chen, Z. Wang, and Y. Shen · 2020
Later among the works it cites.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2021
Later among the works it cites.
A survey on contrastive self-supervised learning
A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, and F. Makedon · 2021
Later among the works it cites.
Video contrastive learning with global context
H. Kuang, Y. Zhu, Z. Zhang, X. Li, J. Tighe, S. Schwertfeger, C. Stachniss, and M. Li · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Factorized multimodal transformer for multimodal sequential learning
A. Zadeh, C. Mao, K. Shi, Y. Zhang, P. P. Liang, S. Poria, and L.-P. Morency · 2019
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
Big self-supervised models are strong semi-supervised learners
T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Cited alongside, same era.
Iterative context-aware graph inference for visual dialog
D. Guo, H. Wang, H. Zhang, Z.-J. Zha, and M. Wang · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
Mcqa: Multimodal co-attention based network for question answering
A. Kumar, T. Mittal, and D. Manocha · 2020
Cited alongside, same era.
Contrastive learning of image representations with cross-video cycle-consistency
H. Wu and X. Wang · 2021
Later among the works it cites.
Rethinking self-supervised correspondence learning: A video frame-level similarity perspective
J. Xu and X. Wang · 2021
Later among the works it cites.
Mtag: Modal-temporal attention graph for unaligned human multimodal language sequences
J. Yang, Y. Wang, R. Yi, Y. Zhu, A. Rehman, A. Zadeh, S. Poria, and L.-P. Morency · 2021
Later among the works it cites.
Contrastive learning of global and local video representations
Z. Zeng, D. McDuff, Y. Song, et al · 2021
Later among the works it cites.
How Attentive are Graph Attention Networks?
S. Brody, U. Alon, and E. Yahav · 2022
Closest in time.
Scaling vision transformers to gigapixel images via hierarchical self-supervised learning
R. J. Chen, C. Chen, Y. Li, T. Y. Chen, A. D. Trister, R. G. Krishnan, and F. Mahmood · 2022
Closest in time.
Motion-aware contrastive video representation learning via foreground-background merging
S. Ding, M. Li, T. Yang, R. Qian, H. Xu, Q. Chen, J. Wang, and H. Xiong · 2022
Closest in time.
Contextualized spatio-temporal contrastive learning with self-supervision
L. Yuan, R. Qian, Y. Cui, B. Gong, F. Schroff, M.-H. Yang, H. Adam, and T. Liu · 2022
Closest in time.
MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound
R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi · 2022
Closest in time.