Fetching the paper…
Reading the bibliography…
The underlying correlation between audio and visual modalities can be utilized to learn supervised information for unlabeled videos.
“Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,”
Michael Gutmann and Aapo Hyvärinen, · 2010
Earlier work this paper cites.
“Two-stream convolutional networks for action recognition in videos,”
Karen Simonyan and Andrew Zisserman, · 2014
Earlier work this paper cites.
“Learning spatiotemporal features with 3d convolutional networks,”
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri, · 2015
Earlier work this paper cites.
“Spatiotemporal residual networks for video action recognition,”
R Christoph and Feichtenhofer Axel Pinz, · 2016
Earlier work this paper cites.
“Unsupervised representation learning by sorting sequences,”
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang, · 2017
Earlier work this paper cites.
“Look, listen and learn,”
Relja Arandjelovic and Andrew Zisserman, · 2017
Earlier work this paper cites.
“The kinetics human action video dataset,”
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al., · 2017
Earlier work this paper cites.
“R-c3d: Region convolutional 3d network for temporal activity detection,”
Huijuan Xu, Abir Das, and Kate Saenko, · 2017
Earlier work this paper cites.
“Transferable feature representation for visible-to-infrared cross-dataset human action recognition,”
Yang Liu, Zhaoyang Lu, Jing Li, Chao Yao, and Yanzi Deng, · 2018
Earlier work this paper cites.
“Global temporal representation based cnns for infrared action recognition,”
Yang Liu, Zhaoyang Lu, Jing Li, Tao Yang, and Chao Yao, · 2018
Earlier work this paper cites.
“Hierarchically learned view-invariant representations for cross-view action recognition,”
Yang Liu, Zhaoyang Lu, Jing Li, and Tao Yang, · 2018
Earlier work this paper cites.
“Objects that sound,”
Relja Arandjelovic and Andrew Zisserman, · 2018
Earlier work this paper cites.
“A closer look at spatiotemporal convolutions for action recognition,”
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri, · 2018
Cited alongside, same era.
“Deep image-to-video adaptation and fusion networks for action recognition,”
Yang Liu, Zhaoyang Lu, Jing Li, Tao Yang, and Chao Yao, · 2019
Cited alongside, same era.
“Self-supervised spatiotemporal learning via video clip order prediction,”
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang, · 2019
Cited alongside, same era.
“Epic-fusion: Audio-visual temporal binding for egocentric action recognition,”
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen, · 2019
Cited alongside, same era.
“Momentum contrast for unsupervised visual representation learning,”
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick, · 2020
Cited alongside, same era.
“A simple framework for contrastive learning of visual representations,”
“Bootstrap your own latent-a new approach to self-supervised learning,”
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al., · 2020
Later among the works it cites.
“Semantics-aware adaptive knowledge distillation for sensor-to-vision action recognition,”
Yang Liu, Keze Wang, Guanbin Li, and Liang Lin, · 2021
Later among the works it cites.
“Temporal contrastive graph learning for video action recognition and retrieval,”
Yang Liu, Keze Wang, Haoyuan Lan, and Liang Lin, · 2021
Later among the works it cites.
“Exploring simple siamese representation learning,”
Xinlei Chen and Kaiming He, · 2021
Later among the works it cites.
“Audio-visual instance discrimination with cross-modal agreement,”
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, · 2020
Cited alongside, same era.
“Pre-training audio representations with self-supervision,”
Marco Tagliasacchi, Beat Gfeller, Felix de Chaumont Quitry, and Dominik Roblek, · 2020
Cited alongside, same era.
“Listen to look: Action recognition by previewing audio,”
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani, · 2020
Cited alongside, same era.
“Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning,”
Ying Cheng, Ruize Wang, Zhihao Pan, Rui Feng, and Yuejie Zhang, · 2020
Cited alongside, same era.
“Learning representations from audio-visual spatial alignment,”
Pedro Morgado, Yi Li, and Nuno Nvasconcelos, · 2020
Cited alongside, same era.
“Video playback rate perception for self-supervised spatio-temporal representation learning,”
Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou, and Qixiang Ye, · 2020
Cited alongside, same era.
“Self-supervised video representation learning using inter-intra contrastive framework,”
Li Tao, Xueting Wang, and Toshihiko Yamasaki, · 2020
Cited alongside, same era.
“Barlow twins: Self-supervised learning via redundancy reduction,”
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny, · 2021
Later among the works it cites.
“Hybrid-order representation learning for electricity theft detection,”
Yuying Zhu, Yang Zhang, Lingbo Liu, Yang Liu, Guanbin Li, Mingzhi Mao, and Liang Lin, · 2022
Closest in time.
“Causal reasoning meets visual representation learning: A prospective study,”
Yang Liu, Yu-Shen Wei, Hong Yan, Guan-Bin Li, and Liang Lin, · 2022
Closest in time.
“Cross-modal causal relational reasoning for event-level visual question answering,”
Yang Liu, Guanbin Li, and Liang Lin, · 2022
Closest in time.
“Tcgl: Temporal contrastive graph for self-supervised video representation learning,”
Yang Liu, Keze Wang, Lingbo Liu, Haoyuan Lan, and Liang Lin, · 2022
Closest in time.
“Urban regional function guided traffic flow prediction,”
Kuo Wang, Lingbo Liu, Yang Liu, Guanbin Li, Fan Zhou, and Liang Lin, · 2023
Closest in time.
“Implicit visual-linguistic deconfounding for radiology report generation,”
Weixing Chen, Yang Liu, Ce Wang, Guanbin Li, Jiarui Zhu, and Liang Lin, · 2023
Closest in time.