Fetching the paper…
Reading the bibliography…
Self-supervised representation learning is a critical problem in computer vision, as it provides a way to pretrain feature extractors on large unlabeled datasets that can be used as an initialization for more efficient and effective training on downstream tasks.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Indoor semantic segmentation using depth information
Camille Couprie, Clément Farabet, Laurent Najman, and Yann LeCun · 2013
Earlier work this paper cites.
Learning rich features from rgb-d images for object detection and segmentation
Saurabh Gupta, Ross Girshick, Pablo Arbeláez, and Jitendra Malik · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Majority vote of diverse classifiers for late fusion
Emilie Morvant, Amaury Habrard, and Stéphane Ayache · 2014
Earlier work this paper cites.
Multi-modal unsupervised feature learning for rgb-d scene labeling
Anran Wang, Jiwen Lu, Gang Wang, Jianfei Cai, and Tat-Jen Cham · 2014
Earlier work this paper cites.
A review and meta-analysis of multimodal affect detection systems
Sidney K D’mello and Jacqueline Kory · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Cross modal distillation for supervision transfer
Saurabh Gupta, Judy Hoffman, and Jitendra Malik · 2016
Earlier work this paper cites.
Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture
Caner Hazirbas, Lingni Ma, Csaba Domokos, and Daniel Cremers · 2016
Earlier work this paper cites.
Black holes and white rabbits: Metaphor identification with visual features
Ekaterina Shutova, Douwe Kiela, and Jean Maillard · 2016
Earlier work this paper cites.
Deep sliding shapes for amodal 3d object detection in rgb-d images
Shuran Song and Jianxiong Xiao · 2016
Earlier work this paper cites.
Unsupervised learning by predicting noise
Piotr Bojanowski and Armand Joulin · 2017
Earlier work this paper cites.
Locality-sensitive deconvolution networks with gated fusion for rgb-d indoor semantic segmentation
Yanhua Cheng, Rui Cai, Zhiwei Li, Xin Zhao, and Kaiqi Huang · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Cascaded feature network for semantic segmentation of rgb-d images
Di Lin, Guangyong Chen, Daniel Cohen-Or, Pheng-Ann Heng, and Hui Huang · 2017
Earlier work this paper cites.
Rdfnet: Rgb-d multi-level residual feature fusion for indoor semantic segmentation
Seong-Jin Park, Ki-Sang Hong, and Seungyong Lee · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Multi-modal deep feature learning for rgb-d object detection
Xiangyang Xu, Yuncheng Li, Gangshan Wu, and Jiebo Luo · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Progressively complementarity-aware fusion network for rgb-d salient object detection
Hao Chen and Youfu Li · 2018
Cited alongside, same era.
3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation
Angela Dai and Matthias Nießner · 2018
Cited alongside, same era.
Deep continuous fusion for multi-sensor 3d object detection
Ming Liang, Bin Yang, Shenlong Wang, and Raquel Urtasun · 2018
Cited alongside, same era.
Multimodal local-global ranking fusion for emotion recognition
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Rgb-d joint modelling with scene geometric information for indoor semantic segmentation
Hong Liu, Wenshan Wu, Xiangdong Wang, and Yueliang Qian · 2018
Densefusion: 6d object pose estimation by iterative dense fusion
Chen Wang, Danfei Xu, Yuke Zhu, Roberto Martín-Martín, Cewu Lu, Li Fei-Fei, and Silvio Savarese · 2019
Later among the works it cites.
Self-labelling via simultaneous clustering and representation learning
Yuki M. Asano, Christian Rupprecht, and Andrea Vedaldi · 2020
Closest in time.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Big self-supervised models are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cross pixel optical-flow similarity for self-supervised learning
Aravindh Mahendran, James Thewlis, and Andrea Vedaldi · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Cross and learn: Cross-modal self-supervision
Nawid Sayed, Biagio Brattoli, and Björn Ommer · 2018
Cited alongside, same era.
Unsupervised feature learning via non-parametric instance discrimination
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin · 2018
Cited alongside, same era.
Learning representations by maximizing mutual information across views
Philip Bachman, R Devon Hjelm, and William Buchwalter · 2019
Cited alongside, same era.
4d spatio-temporal convnets: Minkowski convolutional neural networks
Christopher Choy, JunYoung Gwak, and Silvio Savarese · 2019
Cited alongside, same era.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Closest in time.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Closest in time.
Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning
Ying Cheng, Ruize Wang, Zhihao Pan, Rui Feng, and Yuejie Zhang · 2020
Closest in time.
Metric-guided prototype learning
Vivien Sainte Fare Garnot and Loic Landrieu · 2020
Closest in time.
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al · 2020
Closest in time.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Closest in time.
Self-supervised contrastive video-speech representation learning for ultrasound
Jianbo Jiao, Yifan Cai, Mohammad Alsharid, Lior Drukker, Aris T Papageorghiou, and J Alison Noble · 2020
Closest in time.
Self-supervised modal and view invariant feature learning
Longlong Jing, Yucheng Chen, Ling Zhang, Mingyi He, and Yingli Tian · 2020
Closest in time.
Virtual multi-view fusion for 3d semantic segmentation
Abhijit Kundu, Xiaoqi Yin, Alireza Fathi, David Ross, Brian Brewington, Thomas Funkhouser, and Caroline Pantofaru · 2020
Closest in time.
Prototypical contrastive learning of unsupervised representations
Junnan Li, Pan Zhou, Caiming Xiong, Richard Socher, and Steven CH Hoi · 2020
Closest in time.
Improving unimodal object recognition with multimodal contrastive learning
Johannes Meyer, Andreas Eitel, Thomas Brox, and Wolfram Burgard · 2020
Closest in time.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten · 2020
Closest in time.
Imvotenet: Boosting 3d object detection in point clouds with image votes
Charles R Qi, Xinlei Chen, Or Litany, and Leonidas J Guibas · 2020
Closest in time.
Contrastive visual-linguistic pretraining
Lei Shi, Kai Shuang, Shijie Geng, Peng Su, Zhengkai Jiang, Peng Gao, Zuohui Fu, Gerard de Melo, and Sen Su · 2020
Closest in time.
What should not be contrastive in contrastive learning
Tete Xiao, Xiaolong Wang, Alexei A Efros, and Trevor Darrell · 2020
Closest in time.
Pointcontrast: Unsupervised pre-training for 3d point cloud understanding
Saining Xie, Jiatao Gu, Demi Guo, Charles R Qi, Leonidas J Guibas, and Or Litany · 2020
Closest in time.