Fetching the paper…
Reading the bibliography…
Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
Comparing partitions
Lawrence Hubert and Phipps Arabie · 1985
Earlier work this paper cites.
Cluster ensembles—a knowledge reuse framework for combining multiple partitions
Alexander Strehl and Joydeep Ghosh · 2002
Earlier work this paper cites.
k-means++: The advantages of careful seeding
David Arthur and Sergei Vassilvitskii · 2006
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
Piotr Bojanowski, Rémi Lajugie, Francis Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, and Josef Sivic · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
See, hear, and read: Deep aligned representations
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Lukasz Kaiser, Aidan N Gomez, Noam Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Jointly discovering visual objects and spoken words from raw sensory input
David Harwath, Adria Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass · 2018
Cited alongside, same era.
Supervised autoencoders: Improving generalization performance with unsupervised regularizers
Lei Le et al · 2018
Cited alongside, same era.
How2: a large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze · 2018
Cited alongside, same era.
Towards automatic learning of procedures from web instructional videos
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Later among the works it cites.
Intriguing properties of contrastive losses
Ting Chen and Lala Li · 2020
Later among the works it cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luowei Zhou, Xu Chenliang, and Jason J. Corso · 2018
Cited alongside, same era.
Grounding spoken words in unlabeled video
Angie Boggust, Kartik Audhkhasi, Dhiraj Joshi, David Harwath, Samuel Thomas, Rogerio Feris, Dan Gutfreund, Yang Zhang, Antonio Torralba, Michael Picheny, et al · 2019
Cited alongside, same era.
Large-scale representation learning from visually grounded untranscribed speech
Gabriel Ilharco, Yuan Zhang, and Jason Baldridge · 2019
Cited alongside, same era.
Mining youtube-a dataset for learning fine-grained action concepts from webly supervised video data
Hilde Kuehne, Ahsan Iqbal, Alexander Richard, and Juergen Gall · 2019
Cited alongside, same era.
Use what you have: Video retrieval using representations from collaborative experts
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Xilin Chen, and Ming Zhou · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Later among the works it cites.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten · 2020
Later among the works it cites.
Evolving losses for unsupervised video representation learning
AJ Piergiovanni, Anelia Angelova, and Michael S Ryoo · 2020
Later among the works it cites.
Scan: Learning to classify images without labels
Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, Marc Proesmans, and Luc Van Gool · 2020
Later among the works it cites.
Clusterfit: Improving generalization of visual representations
Xueting Yan, Ishan Misra, Abhinav Gupta, Deepti Ghadiyaram, and Dhruv Mahajan · 2020
Later among the works it cites.
Actbert: Learning global-local video-text representations
Linchao Zhu and Yi Yang · 2020
Later among the works it cites.
Noise estimation using density estimation for self-supervised multimodal learning
Elad Amrani, Rami Ben-Ari, Daniel Rotman, and Alex Bronstein · 2021
Closest in time.
Dual encoding for video retrieval by text
Jianfeng Dong, Xirong Li, Chaoxi Xu, Xun Yang, Gang Yang, Xun Wang, and Meng Wang · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Closest in time.
Prototypical contrastive learning of unsupervised representations
Junnan Li, Pan Zhou, Caiming Xiong, Richard Socher, and Steven CH Hoi · 2021
Closest in time.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, João Henriques, and Andrea Vedaldi · 2021
Closest in time.
Avlnet: Learning audio-visual language representations from instructional videos
Andrew Rouditchenko, Angie Boggust, David Harwath, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, et al · 2021
Closest in time.