Fetching the paper…
Reading the bibliography…
Contrastive learning allows us to flexibly define powerful losses by contrasting positive pairs from sets of negative samples.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Learning fine-grained image similarity with deep ranking
Jiang Wang, Yang Song, Thomas Leung, Chuck Rosenberg, Jingbin Wang, James Philbin, Bo Chen, and Ying Wu · 2014
Earlier work this paper cites.
Deep metric learning using triplet network
Elad Hoffer and Nir Ailon · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf · 2015
Earlier work this paper cites.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks, 2016
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2016
Earlier work this paper cites.
End-to-end concept word detection for video captioning, retrieval, and question answering, 2017
Youngjae Yu, Hyungjin Ko, Jongwook Choi, and Gunhee Kim · 2017
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Learning a text-video embedding from incomplete and heterogeneous data
Antoine Miech, Ivan Laptev, and Josef Sivic · 2018
Earlier work this paper cites.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
Niluthpol C Mithun, Juncheng Li, Florian Metze, and Amit K Roy-Chowdhury · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance discrimination
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin · 2018
Earlier work this paper cites.
Move forward and tell: A progressive generator of video descriptions
Yilei Xiong, Bo Dai, and Dahua Lin · 2018
Earlier work this paper cites.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Cited alongside, same era.
End-to-end dense video captioning with masked transformer, 2018
Luowei Zhou, Yingbo Zhou, Jason J. Corso, Richard Socher, and Caiming Xiong · 2018
Cited alongside, same era.
A theoretical analysis of contrastive unsupervised representation learning, 2019
Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak, Orestis Plevrakis, and Nikunj Saunshi · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context, 2019
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Dual encoding for zero-example video retrieval
Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, Yuan He, Gang Yang, and Xun Wang · 2019
Cited alongside, same era.
A case study on combining ASR and visual features for generating instructional video captions
Multi-modal transformer for video retrieval, 2020
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Later among the works it cites.
Coot: Cooperative hierarchical transformer for video-text representation learning
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox · 2020
Later among the works it cites.
Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko · 2020
Later among the works it cites.
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Later among the works it cites.
Momentum Contrast for Unsupervised Visual Representation Learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jack Hessel, Bo Pang, Zhenhai Zhu, and Radu Soricut · 2019
Cited alongside, same era.
Use what you have: Video retrieval using representations from collaborative experts
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman · 2019
Cited alongside, same era.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Stochastic class-based hard example mining for deep metric learning
Y. Suh, B. Han, W. Kim, and K. M. Lee · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Video classification with channel-separated convolutional networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Cited alongside, same era.
Hard Negative Mixing for Contrastive Learning
Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus · 2020
Later among the works it cites.
Supervised contrastive learning
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan · 2020
Later among the works it cites.
Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning
Jie Lei, Liwei Wang, Yelong Shen, Dong Yu, Tamara L Berg, and Mohit Bansal · 2020
Later among the works it cites.
Learning audio-visual representations with active contrastive coding, 2020
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song · 2020
Later among the works it cites.
End-to-End Learning of Visual Representations from Uncurated Instructional Videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Later among the works it cites.
Learning representations from audio-visual spatial alignment, 2020
Pedro Morgado, Yi Li, and Nuno Vasconcelos · 2020
Later among the works it cites.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge J. Belongie, and Yin Cui · 2020
Later among the works it cites.
Self-supervised representation learning via adaptive hard-positive mining
Xiaokang Yang Shaofeng Zhang, Junchi Yan · 2020
Later among the works it cites.
Learning and evaluating representations for deep one-class classification, 2020
Kihyuk Sohn, Chun-Liang Li, Jinsung Yoon, Minho Jin, and Tomas Pfister · 2020
Later among the works it cites.
Resnest: Split-attention networks
Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Zhi Zhang, Haibin Lin, Yue Sun, Tong He, Jonas Mueller, R Manmatha, et al · 2020
Later among the works it cites.
Actbert: Learning global-local video-text representations
L. Zhu and Y. Yang · 2020
Later among the works it cites.
A straightforward framework for video retrieval using clip, 2021
Jesús Andrés Portillo-Quintero, José Carlos Ortiz-Bayliss, and Hugo Terashima-Marín · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
Contrastive Learning with Hard Negative Samples
Joshua Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka · 2021
Closest in time.