Fetching the paper…
Reading the bibliography…
Learning an object detector or retrieval requires a large data set with manual annotations.
Selective search for object recognition
J. R. Uijlings, K. E. Sande, T. Gevers, and A. W. Smeulders · 2013
Earlier work this paper cites.
Grounded language learning from video described with sentences
H. Yu and J. M. Siskind · 2013
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, A. P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Earlier work this paper cites.
Unsupervised learning of spoken language with visual context
D. Harwath, A. Torralba, and J. R. Glass · 2016
Earlier work this paper cites.
Unsupervised alignment of actions in video with text descriptions
Y. C. Song, I. Naim, A. Al Mamun, K. Kulkarni, P. Singla, J. Luo, D. Gildea, and H. Kautz · 2016
Earlier work this paper cites.
Look, listen, and decode: Multimodal speech recognition with images
F. Sun, D. Harwath, and J. Glass · 2016
Earlier work this paper cites.
Unsupervised deep embedding for clustering analysis
J. Xie, R. Girshick, and A. Farhadi · 2016
Earlier work this paper cites.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Cited alongside, same era.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Cited alongside, same era.
Learning word-like units from joint audio-visual analysis
D. Harwath and J. Glass · 2017
Cited alongside, same era.
Deep self-taught learning for weakly supervised object localization
Z. Jie, Y. Wei, X. Jin, J. Feng, and W. Liu · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi · 2017
Cited alongside, same era.
Learning to separate object sounds by watching unlabeled video
R. Gao, R. Feris, and K. Grauman · 2018
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Later among the works it cites.
How2: a large-scale dataset for multimodal language understanding
R. Sanabria, O. Caglayan, S. Palaskar, D. Elliott, L. Barrault, L. Specia, and F. Metze · 2018
Later among the works it cites.
Cross-modal embeddings for video and audio retrieval
D. Suris, A. Duarte, A. Salvador, J. Torres, and X. G. i Nieto · 2018
Later among the works it cites.
PCL: Proposal cluster learning for weakly supervised object detection
P. Tang, X. Wang, S. Bai, W. Shen, X. Bai, W. Liu, and A. L. Yuille · 2018
Later among the works it cites.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Later among the works it cites.
Scaling and benchmarking self-supervised visual representation learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cooperative learning of audio and video models from self-supervised synchronization
B. Korbar, D. Tran, and L. Torresani · 2018
Cited alongside, same era.
P. Goyal, D. Mahajan, A. Gupta, and I. Misra · 2019
Closest in time.